The Hallway Track

latency

13 tracked signals on latency.

Better prompt caching for GPT-6

OpenAI · OpenAI Blog · Sep 22, 2026

GPT-6 improves prompt caching with higher hit rates, explicit breakpoints, and new latency/cost controls.

Voice Agents That Handle Interrupts - Chintan Agrawal and Daniel Wirjo, AWS

AI Engineer · Jul 20, 2026

Voice agent quality is an audio engineering problem, not an LLM problem, requiring sub-200ms turn detection

“These are all audio engineering problems. They are not LLM problems because you can have the perfect model, perfect track but the experience still might feel broken if the turn taking is off.”
Frontier results, on device - RL Nabors, Arize

AI Engineer · Jun 29, 2026

Local on-device models can replace frontier models like GPT-5 and Claude to cut inference costs, latency, and security risks.

“Every time you reach for foundation models like GPT-5 or Claude, it's costing you, your users, and the environment.”
Cerebras Explains | What Is Time to First Token?

Cerebras · Oct 02, 2026

Cerebras achieves fast TTFT and over 1000 tokens per second simultaneously

“Fast first token and over 1000 tokens per second after that. Responsive start and instant end.”
Cerebras Explains | What Are Tokens per Second?

Cerebras · Oct 01, 2026

Cerebras claims 1,000+ tokens/sec creates instant responses, fundamentally changing the AI interaction experience

“This is the difference between waiting for AI and working with it.”