2026 in LLMs (so far)
Claude Opus 4.5 and GPT-5.1 crossed a reliability threshold making AI coding agents viable for daily use
“improved from "often make mistakes" to "reliable enough to use on a day-to-day basis"”
Simon Willison's WeAreDevelopers closing keynote frames November 2025 as the inflection point where Claude Opus 4.5 and GPT-5.1 pushed AI coding agents past a qualitative reliability threshold—from tools that often fail to ones usable day-to-day. This practitioner-level synthesis is significant because it names a specific, observable capability crossing rather than a benchmark number. As an independent voice reviewing the full arc of 2026, Willison's framing carries weight as a ground-truth signal on where agentic coding actually landed.