The Hallway Track
Governance & Policy

Third-party cyber evaluations involving OpenAI models

Simon Willison · Aug 05, 2026 · Governance & Policy

Misconfigured AI safety evaluations caused OpenAI and Anthropic models to attack real websites accidentally

“In one test, the name of the fictional target for the CTF challenge unintentionally coincided with a real domain. Because the testing environment was mistakenly connected to the internet, the model exploited a real website, mistaking it to be part of the simulated environment.”

A misconfigured evaluation environment by third-party cybersecurity partner Irregular gave both OpenAI and Anthropic models unintended live internet access during isolated Capture-the-Flag tests, resulting in models attacking real websites they mistook for simulated targets. This incident—part of a growing pattern Simon Willison is now tracking with a dedicated tag—implicates the same testing partner in incidents involving both major frontier AI labs. It raises significant concerns about the rigor of AI cybersecurity evaluation infrastructure and the potential for real-world harm during safety testing itself.

security openai anthropic cybersecurity-evals accidental-cyberattacks ai-safety

Watch / read the original source →