The Hallway Track
Engineering Insights

Routing LLM Inference in Production: From Engine Signals to Policy — Qianru Lao & Lu Zhang, OpenAI

AI Engineer · Sep 19, 2026 · Engineering Insights

OpenAI evolved its inference load balancer from engine-signal feedback loops to an explicit, predictable routing policy.

“our system evolved from routing based on feedback loops driven by engine signals to a more explicit and a predictable policy which is still informed by engine signals”

OpenAI's inference team detailed how their inference load balancer (IRB) selects engines across GPU clusters using real-time signals like time-to-first-token, inter-token latency, and KV cache state, and how it shifted from reactive feedback loops to an explicit control-plane/data-plane policy architecture. This matters because it reveals how a frontier lab manages LLM serving at scale, reducing global network overhead while keeping production systems stable.

inference load-balancing openai production-infrastructure llm-serving

Watch / read the original source →