The Hallway Track

llm-inference

10 tracked signals on llm-inference.

Introducing Amazon SageMaker HyperPod Inference Gateway

AWS Machine Learning Blog · Sep 18, 2026

AWS launches SageMaker HyperPod Inference Gateway, GPU-aware Kubernetes routing that cuts first-token latency up to 82%.

“A chatbot user waiting 4.4 seconds for the first token now sees it in under 800 ms.”
Benchmarking LLM Inference at Scale with AIPerf

NVIDIA Developer Blog · Sep 18, 2026

NVIDIA introduces AIPerf, a tool for benchmarking LLM inference performance at scale.

“All of these paths have the same problem: single-process performance limits, Python's GIL capping concurrency”
Parallelize speculative decoding with P-EAGLE on Amazon SageMaker AI

AWS Machine Learning Blog · Jun 16, 2026

AWS invented and open-sourced P-EAGLE, parallelizing speculative decoding for up to 1.69x faster LLM inference on SageMaker.

“P-EAGLE instead fills positions 2–4 with learnable placeholders and predicts all four tokens at once”
What Lies Beneath the API — Benjamin Cowen, Modal

AI Engineer · Jun 02, 2026

As AI products mature, more companies turn to fine-tuning over frontier APIs for performance and cost gains.

“If you tell your LLM to speak like a caveman, you can reduce your tokens by like a lot.”
DynoSim: Simulating the Pareto Frontier

NVIDIA Developer Blog · May 29, 2026

NVIDIA's DynoSim simulates the Pareto frontier of interacting LLM serving configuration choices to ease deployment tuning.

Right-size generative AI endpoints with concurrency sweeps on Amazon SageMaker AI

AWS Machine Learning Blog · Sep 22, 2026

AWS details concurrency sweeps in SageMaker AI Inference Recommendations to right-size generative AI endpoints.

“A concurrency sweep is a systematic benchmarking approach that sends controlled, increasing levels of concurrent traffic to your Amazon SageMaker AI endpoint and analyzes its performance.”