The Hallway Track

vLLM

8 tracked signals on vLLM.

Tiered KV cache for large LLMs on Amazon SageMaker HyperPod with Curvine

AWS Machine Learning Blog · Aug 12, 2026

AWS tiered KV cache on SageMaker HyperPod delivers 2.7x TTFT improvement and 100% cross-Pod cache hit rate

“With this architecture, workloads that previously required P5 instances can run on lower-cost G6e instances, reducing per-endpoint cost.”