How Open Source Became AI's Backbone | Inferact with a16z
vLLM open-source inference engine now runs on half a million GPUs simultaneously
“VLM is a inference engine. It is kind of like databases and operating system other critical software to power AGI.”
Simon Mo, co-founder of Inferact and lead maintainer of vLLM, discusses how the open-source inference engine has scaled to half a million concurrent GPUs and is now considered critical AI infrastructure comparable to databases and operating systems. He argues the capability gap between open-weight and proprietary frontier models is already minimal, and that open source becomes the default for trusted use cases where users need control over guardrails. The conversation frames open-source inference as the foundational layer of the emerging AI stack.