Routing LLM Inference in Production: From Engine Signals to Policy — Qianru Lao & Lu Zhang, OpenAI
OpenAI evolved its inference load balancer from engine-signal feedback loops to an explicit, predictable routing policy.
“our system evolved from routing based on feedback loops driven by engine signals to a more explicit and a predictable policy which is still informed by engine signals”
OpenAI's inference team detailed how their inference load balancer (IRB) selects engines across GPU clusters using real-time signals like time-to-first-token, inter-token latency, and KV cache state, and how it shifted from reactive feedback loops to an explicit control-plane/data-plane policy architecture. This matters because it reveals how a frontier lab manages LLM serving at scale, reducing global network overhead while keeping production systems stable.