The Hallway Track

cost-optimization

27 tracked signals on cost-optimization.

How to Build a Model Router in the Harness

LangChain · Oct 06, 2026

LangChain achieved 64% cost reduction in their coding agent via model routing with no quality loss

“we were able to see a 64% reduction in median cost per thread with no measurable change in quality”
Better prompt caching for GPT-6

OpenAI · OpenAI Blog · Sep 22, 2026

GPT-6 improves prompt caching with higher hit rates, explicit breakpoints, and new latency/cost controls.

Quoting Drew Breunig

Simon Willison · Aug 23, 2026

Fable's high cost ends the era of relying on new models to cheaply solve engineering problems

“Prior to Fable, it felt silly to waste too much time improving your coding harness or context strategies. A new model would arrive at the same price (or cheaper!) and paper over most of your problems.”
Frontier Model Cost and Open-Weights Popularity is Driving Demand for Model Routing

Latent Space Blog · Aug 18, 2026

Model routing is now a critical AI deployment strategy driven by frontier cost and open-weights competition

“A big goal of Glean is to avoid using LLMs for tasks where we don't need them. Sometimes you'll see queries in Glean where people are adding two numbers or multiplying two numbers. They could have used a calculator to do that.”
Agentic AI needs more than one model.

NVIDIA GTC · Aug 13, 2026

Agentic AI systems should route tasks across frontier and open models for cost efficiency.

“being able to build these agents that are able to use a system of models to address the problem at hand and do it most cost-effectively is going to benefit us”
Notion's Token Town — Sarah Sachs, Notion

AI Engineer · Jul 23, 2026

Notion is repositioning as a human-agent collaboration platform managing token costs sustainably.

“Today that collaboration happens between humans and agents. Humans and humans, agents and agents.”
Pair Nova 2 Lite with Claude for cost-optimized document processing

AWS Machine Learning Blog · Jun 29, 2026

Pairing Amazon Nova 2 Lite with Claude Sonnet 4.6 in a two-model Bedrock pipeline cuts per-page document-digitization cost by about two-thirds.

“This two-model approach costs about two-thirds less per page than a single-model alternative that sends the entire task to one vision-language model.”
Yes, Jev Is Insane, But There's A Catch

Two Minute Papers · Sep 22, 2026

A new AI called Jev makes decisions instead of generating text, running up to 200x faster than chatbots.

“It pops out instantly, up to about 200 times faster than our chatbots today.”
Run AI SREs without burning token budgets | ODSP928

Microsoft Developer (Build) · Jun 03, 2026

Running AI SRE agents on all enterprise alerts naively costs ~$730,000/year, driven mostly by input-token costs.

“these are all driven by the cost on input tokens”
Bedrock Cost Allocation | The Keys to AWS Optimization | S17 E7

AWS (re:Invent) · Jun 05, 2026

AWS webinar previews new Bedrock cost allocation features and CUR columns for tracking generative AI spend.

“We are excited to talk about some new features that have come out. How to look at your spend using the cur and these new columns”