OpenAI's custom Jalapeño chip delivers industry-leading AI inference speed and efficiency.
inference
66 tracked signals on inference.
AMD acquires Taalas, signaling Lisa Su's conviction in custom ASICs for AI inference
OpenAI and Broadcom unveiled Jalapeño, a custom AI chip built for LLM inference.
Specialized GPU kernel generation can significantly enhance inference efficiency.
Apple launches Core AI framework for on-device model inference across iOS and macOS.
OpenAI GPT-5.6 Sol, Terra, and Luna are now generally available on Amazon Bedrock.
NVIDIA Rubin GPU architecture is purpose-built for agentic AI workloads at scale
“What began as discrete AI model training and human-facing chat interfaces has evolved into always-on AI factories dedicated to producing intelligence at scale.”
Clay runs over 350 million go-to-market AI agents monthly, processing trillions of tokens per week.
“We run this over 350 million times a month. It processes trillions of tokens every week.”
Fireworks and Baseten reach decacorn status as AI inference infrastructure sees explosive growth
“if you are gonna do multimodel inference, you are gonna need a router”
A startup is betting on diffusion-based LLMs because they are inherently more parallel at inference time than autoregressive transformers.
“the bitter lesson is that the more parallel solution is the one that is eventually going to win”
OpenAI evolved its inference load balancer from engine-signal feedback loops to an explicit, predictable routing policy.
“our system evolved from routing based on feedback loops driven by engine signals to a more explicit and a predictable policy which is still informed by engine signals”
Moonshot AI's Kimi K3, a 2.8-trillion-parameter open-weight model, is now available on Amazon Bedrock.
“Kimi K3 is its most capable model and the first open model to reach 2.8 trillion parameters.”
Inference is no longer the bottleneck in AI development.
“Just do what I want.”
NVIDIA Vera Rubin and Blackwell set new performance-per-watt benchmarks for agentic AI workloads
“across 100 trillion tokens of real-world usage, OpenRouter's State of AI report found that average prompt tokens per request grew roughly fourfold”
NVIDIA Groq 3 LPX accelerator enables ultrafast interactive inference on Vera Rubin NVL72 at long context
“NVIDIA Vera Rubin NVL72, the most versatile machine ever built, delivering high throughput and interactivity across the widest range of AI workloads—from small to large models, both open and closed.”
NVIDIA reframes AI factory efficiency around performance per watt, not GPU count
“The question is no longer how many GPUs fit in a data center, but how much AI output each available megawatt can deliver.”
NVIDIA's Rubin GPU architecture kills megakernels, ending a major inference optimization research direction
“the GPU is designed in such a way that it kills mega kernels. So it seems like that entire research field won't be continued.”
Baseten raised $13B Series F as inference engineering emerges as a critical AI discipline
“How do you turn those weights from training into a product that is fast, reliable, and affordable at scale?”
Cerebras and Flex are scaling CS3 wafer-scale AI supercomputer production to thousands of units monthly in Silicon Valley.
“We have been producing with flex at hundreds of systems a month and now preparing to build thousands of systems a month going into next year.”
Baseten says inference with a software layer is sticky—top 30 customers never churned at 400% NDR.
“none of our top 30 customers have ever churned. You know, we're talking like 400% annual NDR.”
NVIDIA introduces Green Contexts for fine-grained GPU resource sharing between concurrent workloads
“Controlling how GPU resources are shared between them remains difficult.”
Nebius, backed by $2B Nvidia investment, offers full-stack open LLM inference from silicon to service
“Typically, most AI teams are stuck between choosing two bad options. Closed APIs are very easy to get started with, but you often hit a ceiling very quickly.”
Transformer architecture standardization will enable extreme chip specialization and massive inference performance gains
“I think we're gonna have crazy levels of specialization because there's gonna be so much inference demand. And that's gonna unlock just like super exciting levels of performance.”
Fractile is building inference chips targeting 3-6 month architectural leads over competitors
“if you find a way to structurally secure a 3-6 month lead, you will win in all these implementations.”
NVIDIA open-sourced TensorRT Model Connect, built from the ground up for coding agents
Two lines of algebra make transformer RMS norm cheaper and faster via weight folding and deferred division.
“the GPUs are not slow or bad at math, but they're bad at everything else around the actual math”
AI21 engineers traced rare 'one-in-a-thousand gibberish' output bugs unique to vLLM during Jamba RL training.
“there is no crash there's no warning no error and there's high confidence that's not a quality issue”
Google engineers argue production-scale LLM inference benchmarks need purpose-built tooling like inference-perf.
“LLMD is a distributed inference framework um that makes production scale inference possible.”
SageMaker AI shipped 13 new inference capabilities in 2026 across managed endpoints and HyperPod.
“Generative AI inference is uniquely hard: models are tens to hundreds of gigabytes, latency requirements are measured in tokens per second, cold starts can span multiple minutes as containers and weights transfer, GPU capacity is constrained, and traditional monitoring tools expose none of the token-level signals that matter in production.”
Cerebras claims it has eliminated the speed-throughput tradeoff in AI inference
“In AI, speed is productivity.”
Databricks Unity AI Gateway offers smart routing to cut costs 30%+ while matching frontier model quality
Post-training is the primary moat for AI application companies over black-box APIs
“right now, with one person, a few weeks, uh without understanding how to write a single line of code, you can do that”
vLLM open-source inference engine now runs on half a million GPUs simultaneously
“VLM is a inference engine. It is kind of like databases and operating system other critical software to power AGI.”
GPU-aware attention architecture co-design is now the key lever for long-context inference performance
“Shaping model architecture around how GPUs execute it is the premise of AI model co-design.”
MiniMax open-sourced its strongest model M3, with Together AI capturing lion's share of inference traffic.
“We do believe that the open source community as a whole is very strong and powerful.”
Hugging Face enables running a vLLM inference server on HF Jobs with a single command.
AWS launches container image caching for SageMaker AI inference, cutting scale-out latency up to 2x for generative AI models.
“This speeds up end-to-end latency by up to 2x for generative AI models during scale-out events.”
NVIDIA leads agentic coding performance on AA-AgentPerf, the industry's first multi-vendor agentic AI benchmark.
“Artificial Analysis AgentPerf (AA-AgentPerf) offers the industry's first multi-vendor open benchmarks profiling trajectories that are representative of real-world AI agent coding tasks.”
Google DeepMind's DiffusionGemma brings diffusion-based, high-throughput text generation optimized for NVIDIA platforms.
NVIDIA introduces the 'AI grid' reference design, positioning telcos to become intelligence providers in the token economy.
“There has truly never been a better time to be in Telecom.”
Fireworks AI brings day-zero open-source model inference to Microsoft Foundry for production AI workflows.
“We serve around 13 trillion tokens per day and 180,000 requests per second.”
NVIDIA's Vera Rubin platform is in full production, delivering 10x faster agent inference.
“Useful AI has arrived.”
StepFun's Step 3.7 Flash 198B multimodal model is now available on NVIDIA-accelerated infrastructure
“AI applications are moving beyond text generation to multimodal systems that can perceive, search, and reason across images, documents, video, and language in real time”
NVIDIA Dynamo Snapshot reduces Kubernetes inference cold-start from minutes to seconds
AWS combines NVIDIA NIM, Bedrock AgentCore, and Strands Agents for production-grade multi-agent systems
Google reveals TPU V8 splits into training (8T) and inference (8I) specialized variants
“a lot of the intelligence is actually coming from inference”
Cerebras inference API delivers 1000+ tokens/sec via one-line OpenAI migration
“changing one line of your base URL, and suddenly your app is running at over 1000 tokens per second”
NVIDIA Dynamo-Triton enables deployment of HSTU generative recommender systems at scale
AWS vLLM-Omni DLC enables unified text-to-image-to-video pipelines on SageMaker AI
Cerebras runs models at over 1000 tokens per second for near-instantaneous inference
“We run models at over 1000 tokens per second, so thinking seems instantaneous.”
CoreWeave is building an inference platform offering serverless and dedicated consumption models to serve small to trillion-parameter workloads.
“another feature that we have on the serverless side is what we calling provisioned throughput”
NVIDIA TensorRT Model Connect deploys open models to inference in two commands
“Open AI models are evolving faster than ever, but bringing them into native applications can still require model-specific conversion, preprocessing, post-processing, and runtime code.”
Salesforce achieved Multi-AZ HA for Agentforce using SageMaker's new SchedulingConfig placement API
Baseten joins Hugging Face as an inference provider partner.
Fireworks AI's open-weight model inference is now available inside Microsoft's Azure Foundry.
“We process more than 30 trillion tokens per day.”
Fireworks AI's managed inference and training service is now available on Microsoft Azure.
“Fireworks is a managed service where we do both AI inference performance as well as AI training all as a managed service that you can now use on Microsoft Azure.”
Databricks has built a proprietary inference platform serving frontier AI models at scale
“At Databricks, we've built a unique inference platform that serves every frontier”
LLM inference is bottlenecked by memory bandwidth, not compute speed
“LLM inference is actually limited by memory bandwidth, not computational speed.”
Cerebras positions inference speed and cost as the critical battleground for AI infrastructure
“output is where speed and cost directly impact your users”
NVIDIA offers guidance on sizing GPUs for AI inference to optimize TCO