The Hallway Track
Engineering Insights

Vertical Mobility: Inference from MVP to Trillion-Parameter Workloads — Sitanshu Gupta, CoreWeave

AI Engineer · Sep 19, 2026 · Engineering Insights

CoreWeave is building an inference platform offering serverless and dedicated consumption models to serve small to trillion-parameter workloads.

“another feature that we have on the serverless side is what we calling provisioned throughput”

CoreWeave's inference lead Sitanshu Gupta described the company's inference platform, which offers serverless (pay-per-token, managed) and dedicated (customer-controlled hardware) consumption models, plus provisioned throughput to address noisy-neighbor capacity problems. It matters as a look at how a major GPU cloud provider is architecting inference infrastructure to span MVP to trillion-parameter workloads, though the talk is high-level and lacks concrete metrics or announcements.

inference coreweave serverless gpu-infrastructure llm-serving

Watch / read the original source →