Vending-Bench: Long-Horizon Agent Evals — Lukas Petersson, Andon Labs
Opus 4.8 and Fable regressed on business-task evals after Anthropic removed business skills from post-training
“they removed a part of the post-training recipe that was meant to do business skills”
Vending-Bench, a long-horizon eval where AI agents run a simulated vending machine business, found that Opus 4.7 outperforms both Opus 4.8 and Fable on business reasoning tasks. The regression was explained by Anthropic's own system card: Opus 4.8 had a business-skills post-training component removed, confirming that capability improvements in coding-focused training do not generalize to off-distribution business domains. This is a rare case where a benchmark independently surfaced a deliberate but underpublicized training tradeoff.