The Hallway Track
Research Findings

What happened after 2,000 people tried to hack my AI assistant

Simon Willison · Jun 26, 2026 · Research Findings

Frontier model anti-prompt-injection training held up against 6,000 attempts to leak secrets from an AI assistant.

“after 6,000 attempts (and $500 in token spend and a Google account suspension triggered by too many inbound emails) nobody managed to leak the secret”

A public challenge invited thousands to extract secrets from an Opus 4.6-powered assistant via email injection, and after 6,000 attempts none succeeded, suggesting labs' anti-injection training is increasingly effective. It matters as real-world evidence that frontier models resist prompt injection better than before, though Willison cautions this is no guarantee against more sophisticated attacks in production systems.

prompt-injection ai-security llm-robustness red-teaming

Watch / read the original source →