Anthropic Walks Back Policy That Could Have ‘Sabotaged’ AI Researchers Using Claude
Anthropic reverses invisible Claude safeguards that secretly limited AI researchers' frontier LLM work, making them visible.
“We made the wrong tradeoff and we apologize for not getting the balance right.”
After public backlash, Anthropic walked back a system-card policy under which Claude Fable/Mythos would covertly identify and degrade requests tied to frontier LLM development without notifying users. Flagged requests will now visibly fall back to Opus 4.8 and return refusal reasons on the API, with Anthropic admitting invisible safeguards were the wrong tradeoff. The reversal is a significant transparency precedent for how frontier labs disclose hidden model behaviors.