The Hallway Track

computer-use

12 tracked signals on computer-use.

Introducing GPT-6.1 Sol

OpenAI · OpenAI Blog · Sep 29, 2026

OpenAI launches GPT-6.1 Sol at one-fifth the price of its flagship Astra model

“near-Astra intelligence for coding, computer use, and professional work at one-fifth of Astra's standard API input and output token prices”
Computer Use at the Edge of the Statistical Precipice — Pierluca D'Oro, Programma Labs

AI Engineer · Aug 14, 2026

Standard computer use benchmarks are gameable by blind replay scripts, invalidating frontier model comparisons

“if you try to evaluate this kind of agent on standard benchmarks such as OSWorld or MobileWorld, you will see that the success rate of this agent compared to the frontier model from which the agent was extracted is actually the same or even better”
Perception Agents — Antje Barth, Amazon AGI Lab

AI Engineer · Jul 23, 2026

Agent reliability, not capability, is the critical unsolved problem for enterprise automation trust.

“if your agent one in four times deletes a database, you will never touch that agent again”
An opinionated guide to which AI to use to do stuff

Simon Willison · Jul 27, 2026

AI practitioner guides have shifted focus from chat interfaces to agentic systems doing hours of real work

“The most powerful way to use AI is to give it access to your computer. You do that by downloading the ChatGPT or Claude apps and picking a mode to use.”