The Hallway Track
Engineering Insights

How Open Source Became AI's Backbone | Inferact with a16z

a16z · Aug 06, 2026 · Engineering Insights

vLLM open-source inference engine now runs on half a million GPUs simultaneously

“VLM is a inference engine. It is kind of like databases and operating system other critical software to power AGI.”

Simon Mo, co-founder of Inferact and lead maintainer of vLLM, discusses how the open-source inference engine has scaled to half a million concurrent GPUs and is now considered critical AI infrastructure comparable to databases and operating systems. He argues the capability gap between open-weight and proprietary frontier models is already minimal, and that open source becomes the default for trusted use cases where users need control over guardrails. The conversation frames open-source inference as the foundational layer of the emerging AI stack.

open-source inference vLLM infrastructure GPU

Watch / read the original source →