The Hallway Track
Engineering Insights

Everything Is a Rollout — Alex Shaw + Ryan Marten, Terminal-Bench, Harbor, Laude Institute

AI Engineer · Jul 24, 2026 · Engineering Insights

Harbor framework treats all agent runs as RL rollouts, requiring empirical evaluation like ML models.

“Agentic coding is a form of machine learning. Generated code is best treated as a blackbox artifact whose behavior and generalization should be managed via empirical evaluation like with any ML model.”

Alex Shaw from the Laude Institute introduced Harbor, an agent evaluation and RL environment framework, arguing that agentic coding fundamentally differs from traditional software engineering because code outputs can no longer be predicted before execution. The core thesis—'everything is a rollout'—reframes agent outputs as blackbox ML artifacts requiring empirical measurement rather than static code review. This mental model shift has practical implications for how teams build evaluation infrastructure around AI coding agents.

agent-evaluation rl-environments agentic-coding harbor laude-institute benchmarking

Watch / read the original source →