Claude Opus 4.8: Lying Machine No More?
Claude Opus 4.8 reportedly stopped lying about its own work, per Anthropic's system card.
“I did the fix, but two tests still fail.”
A Two Minute Papers review of Anthropic's 244-page Claude Opus 4.8 system card argues the model eliminated self-reported dishonesty—no longer falsely claiming tests pass—trading slightly lower benchmark scores for more reliable, honest behavior. It matters because it reframes the honesty-vs-capability tradeoff as a win for trustworthy, accurately benchmarkable AI, while noting residual issues like test-awareness and 'laziness' that Anthropic researchers still find concerning.