The Hallway Track
Engineering Insights

Reduce ASR inference costs by 75% with NVIDIA MPS on Amazon EC2

AWS Machine Learning Blog · Aug 27, 2026 · Engineering Insights

NVIDIA MPS cuts ASR GPU infrastructure 75% while maintaining sub-second latency at 92 RPS.

“CUDA MPS partitions the GPU into 4 concurrent instances at 25% SM each, achieving 92.1 RPS with only 4 GPUs”

AWS, NVIDIA, and Heidi Health demonstrate that NVIDIA CUDA Multi-Process Service (MPS) with Triton Inference Server reduces GPU instance count from 16 to 4 for ASR workloads by enabling true concurrent SM execution instead of sequential time-slicing. The technique exploits the fact that a single ASR request uses only 15-20% of GPU compute capacity, leaving 80% idle under default CUDA behavior. This is a practical cost-reduction playbook for latency-sensitive inference at scale, particularly relevant for speech and audio AI products.

inference-cost gpu-optimization asr nvidia aws mlops

Watch / read the original source →