0.00% on the Benchmark. A Working Exploit in the Field. Both Are True.
Anthropic published a third-party result of zero successful prompt injections across 720 attempts against Claude Code auto mode. Six weeks later Johann Rehberger achieved remote code execution in 3 to 4 runs out of 5. Neither number is wrong, and the gap between them is a controls lesson every finance team already knows.