Vertical Mobility: Inference from MVP to Trillion-Parameter Workloads — Sitanshu Gupta, CoreWeave
CoreWeave is building an inference platform offering serverless and dedicated consumption models to serve small to trillion-parameter workloads.
“another feature that we have on the serverless side is what we calling provisioned throughput”
CoreWeave's inference lead Sitanshu Gupta described the company's inference platform, which offers serverless (pay-per-token, managed) and dedicated (customer-controlled hardware) consumption models, plus provisioned throughput to address noisy-neighbor capacity problems. It matters as a look at how a major GPU cloud provider is architecting inference infrastructure to span MVP to trillion-parameter workloads, though the talk is high-level and lacks concrete metrics or announcements.