The Hallway Track

inference-speed

9 tracked signals on inference-speed.

DiffusionGemma

Simon Willison · Jun 10, 2026

Google released DiffusionGemma, an open-weight Apache 2 diffusion-based text generation model.

“That research has returned in the best possible way: as a new open weight (Apache 2 licensed) Gemma model”
Are Open Source Models Actually Ready for Production? | Spill The Tea

LangChain · Jun 13, 2026

Open source models still trail closed models on general agentic tasks but can outperform them once fine-tuned for a specific domain.

“because they're open and because you own the weights, you can actually fine-tune or RL them on your specific domain, which can actually make them perform better than those same closed models”
Cerebras Explains | What Is Time to First Token?

Cerebras · Oct 02, 2026

Cerebras achieves fast TTFT and over 1000 tokens per second simultaneously

“Fast first token and over 1000 tokens per second after that. Responsive start and instant end.”
Cerebras Explains | What Are Tokens per Second?

Cerebras · Oct 01, 2026

Cerebras claims 1,000+ tokens/sec creates instant responses, fundamentally changing the AI interaction experience

“This is the difference between waiting for AI and working with it.”
Yes, Jev Is Insane, But There's A Catch

Two Minute Papers · Sep 22, 2026

A new AI called Jev makes decisions instead of generating text, running up to 200x faster than chatbots.

“It pops out instantly, up to about 200 times faster than our chatbots today.”