Speculative decoding cuts LLM inference cost but only when GPU memory has spare capacity
vLLM
8 tracked signals on vLLM.
AI21 engineers traced rare 'one-in-a-thousand gibberish' output bugs unique to vLLM during Jamba RL training.
“there is no crash there's no warning no error and there's high confidence that's not a quality issue”
Red Hat demonstrates KV cache-aware routing and prefill/decode disaggregation for agentic LLM workloads on Kubernetes.
“benchmarks actually don't show you is the chaotic reality of multi-turn interactions, massive context fluctuations which are very typical of agentic workloads”
AWS tiered KV cache on SageMaker HyperPod delivers 2.7x TTFT improvement and 100% cross-Pod cache hit rate
“With this architecture, workloads that previously required P5 instances can run on lower-cost G6e instances, reducing per-endpoint cost.”
vLLM open-source inference engine now runs on half a million GPUs simultaneously
“VLM is a inference engine. It is kind of like databases and operating system other critical software to power AGI.”
Inference engine performance bottlenecks live in the scheduler and GPU execution layer
AWS enables streaming TTS on SageMaker via vLLM-Omni DLC for real-time voice apps
DeepLearning.AI and Red Hat launch a course on efficient open-source LLM inference using vLLM.
“The techniques you learn in this course are what power efficient LM serving in production today.”