AI workloads are shifting towards smaller transfers where latency is crucial.
“Homa...can reduce tail latency by an order of magnitude or more.”
13 tracked signals on latency.
AI workloads are shifting towards smaller transfers where latency is crucial.
“Homa...can reduce tail latency by an order of magnitude or more.”
GPT-6 improves prompt caching with higher hit rates, explicit breakpoints, and new latency/cost controls.
Voice agents need sub-950ms response; small models beat frontier models on latency
“A frontier model that think for a full second has already lost the room, no matter how good the answer is.”
Speculative decoding cuts LLM inference cost but only when GPU memory has spare capacity
NVIDIA DOCA GPUNetIO enables GPUs to control networking directly, eliminating CPU bottlenecks
“GPU applications increasingly need networking and data movement to behave like first-class GPU-controlled operations rather than host-driven services.”
Google DeepMind launched NanoBanana 2 Light, a faster cheaper image model approaching frontier quality
“getting kind of that like 3 second latency just unlocks a whole bunch of”
Voice agent quality is an audio engineering problem, not an LLM problem, requiring sub-200ms turn detection
“These are all audio engineering problems. They are not LLM problems because you can have the perfect model, perfect track but the experience still might feel broken if the turn taking is off.”
Local on-device models can replace frontier models like GPT-5 and Claude to cut inference costs, latency, and security risks.
“Every time you reach for foundation models like GPT-5 or Claude, it's costing you, your users, and the environment.”
AWS launches container image caching for SageMaker AI inference, cutting scale-out latency up to 2x for generative AI models.
“This speeds up end-to-end latency by up to 2x for generative AI models during scale-out events.”
Cerebras achieves fast TTFT and over 1000 tokens per second simultaneously
“Fast first token and over 1000 tokens per second after that. Responsive start and instant end.”
Cerebras claims 1,000+ tokens/sec creates instant responses, fundamentally changing the AI interaction experience
“This is the difference between waiting for AI and working with it.”
Natera's healthcare voice agent hit 100% tool-calling accuracy at under $0.01 per call using Bedrock AgentCore
Databricks Feature Store enables sub-second feature freshness for real-time ML models