AWS launches SageMaker HyperPod Inference Gateway, GPU-aware Kubernetes routing that cuts first-token latency up to 82%.
“A chatbot user waiting 4.4 seconds for the first token now sees it in under 800 ms.”
11 tracked signals on kubernetes.
AWS launches SageMaker HyperPod Inference Gateway, GPU-aware Kubernetes routing that cuts first-token latency up to 82%.
“A chatbot user waiting 4.4 seconds for the first token now sees it in under 800 ms.”
Crusoe built 'autoclusters' to automatically detect and replace failing GPU nodes at hyperscale
“GPU failures are inevitable. Therefore, at such a scale, manual troubleshooting is completely unviable.”
Anyscale on Azure brings Ray-powered distributed AI compute to AKS for owning your AI stack at scale
“It has over 12 million downloads per week.”
NVIDIA Dynamo Snapshot reduces Kubernetes inference cold-start from minutes to seconds
NVIDIA proposes isolated tenant Kubernetes clusters to share GPU infrastructure across teams
A practitioner walks through what agentic AI workloads actually require to run on Kubernetes.
“This session is for people already running Kubernetes who are trying to figure out what AI workloads actually need on top of Kubernetes.”
NVIDIA offers real-time GPU utilization visibility tools for Kubernetes AI workloads
NVIDIA introduces NodeWright to manage Kubernetes node fleets for GPU workloads.
“Kubernetes manages what runs on your nodes. Managing the nodes themselves is the challenge”
AWS SageMaker AI Spaces add-on cuts GPU IDE setup on EKS from days to minutes
“Standing up a standalone JupyterHub environment with GPU access, storage, and authentication typically takes a platform team 3–5 days. With the add-on, a data scientist launches a fully configured Space in about 5 minutes.”
CloudNativePG's new failoverQuorum feature brings quorum-based consistency to synchronous Postgres replication on Kubernetes.
“if you're running Postgres on Kubernetes, then I do think everybody should use synchronous replication”
Microsoft's Azure Linux is a Fedora-upstream distro adding enterprise reliability, compliance, and predictable patch cadence for Azure.
“So you get a lot of that upstream open source innovation and that upstream open source speed, but with the programmatic reliability that customers want for Azure.”