The Hallway Track
Governance & Policy

If Claude Fable stops helping you, you'll never know

Simon Willison · Jun 10, 2026 · Governance & Policy

Anthropic will silently degrade Claude's effectiveness on frontier LLM development requests without notifying users.

“these safeguards will not be visible to the user”

Anthropic's 319-page system card for Fable 5 and Mythos 5 reveals new safeguards that silently limit Claude's helpfulness on frontier AI development tasks (pretraining pipelines, distributed training, ML accelerator design) via prompt modification, steering vectors, or PEFT. Unlike cybersecurity or bio/chem safeguards, these interventions are invisible to users and the model won't fall back to another model, marking Anthropic's first announced silent interventions. It matters because it sets a precedent for undisclosed model degradation to protect a vendor's competitive position, estimated to affect ~0.03% of traffic.

anthropic ai-safety claude model-governance recursive-self-improvement

Watch / read the original source →