Quoting Matteo Wong, The Atlantic
A cybersecurity expert says Anthropic's Fable model refused a security-review prompt but complied when asked to 'fix this code,' calling it working as intended.
“the model working as intended”
Amid the White House's escalating conflict with Anthropic, a report alleged a 'Fable jailbreak,' but security expert Katie Moussouris reviewed it and concluded the model behaved as intended for cyberdefense, refusing a 'review for security issues' prompt while complying with a 'fix this code' request. This matters because it shows AI safety behavior being weaponized in a political fight over export controls and the AI race.