The Hallway Track

inference

66 tracked signals on inference.

[AINews] AMD buys Taalas

Lisa Su · Latent Space Blog · Aug 07, 2026

AMD acquires Taalas, signaling Lisa Su's conviction in custom ASICs for AI inference

Introducing Core AI

Apple Developer (WWDC) · Aug 17, 2026

Apple launches Core AI framework for on-device model inference across iOS and macOS.

Inside NVIDIA Rubin GPU Architecture: Powering the Era of Agentic AI

NVIDIA Developer Blog · Jul 21, 2026

NVIDIA Rubin GPU architecture is purpose-built for agentic AI workloads at scale

“What began as discrete AI model training and human-facing chat interfaces has evolved into always-on AI factories dedicated to producing intelligence at scale.”
Betting on Diffusion

No Priors · Sep 20, 2026

A startup is betting on diffusion-based LLMs because they are inherently more parallel at inference time than autoregressive transformers.

“the bitter lesson is that the more parallel solution is the one that is eventually going to win”
Introducing Kimi K3 on Amazon Bedrock

AWS Machine Learning Blog · Sep 18, 2026

Moonshot AI's Kimi K3, a 2.8-trillion-parameter open-weight model, is now available on Amazon Bedrock.

“Kimi K3 is its most capable model and the first open model to reach 2.8 trillion parameters.”
[AINews] Megakernels are so dead and so back

Latent Space Blog · Aug 05, 2026

NVIDIA's Rubin GPU architecture kills megakernels, ending a major inference optimization research direction

“the GPU is designed in such a way that it kills mega kernels. So it seems like that entire research field won't be continued.”
Cerebras and Flex Expand U.S. AI Supercomputer Manufacturing

Cerebras · Jul 23, 2026

Cerebras and Flex are scaling CS3 wafer-scale AI supercomputer production to thousands of units monthly in Silicon Valley.

“We have been producing with flex at hundreds of systems a month and now preparing to build thousands of systems a month going into next year.”
Baseten: "We've Never Lost Our Top Customers”

No Priors · Jun 01, 2026

Baseten says inference with a software layer is sticky—top 30 customers never churned at 400% NDR.

“none of our top 30 customers have ever churned. You know, we're talking like 400% annual NDR.”
Control How Your GPU Shares Work with Green Contexts

NVIDIA Developer Blog · Oct 06, 2026

NVIDIA introduces Green Contexts for fine-grained GPU resource sharing between concurrent workloads

“Controlling how GPU resources are shared between them remains difficult.”
What Makes Open Models Fast in Production — Sujee Maniyam, Nebius

AI Engineer · Oct 03, 2026

Nebius, backed by $2B Nvidia investment, offers full-stack open LLM inference from silicon to service

“Typically, most AI teams are stuck between choosing two bad options. Closed APIs are very easy to get started with, but you often hit a ceiling very quickly.”
What building Tesla's autopilot chips taught Cognition's President about AI

LangChain · Oct 02, 2026

Transformer architecture standardization will enable extreme chip specialization and massive inference performance gains

“I think we're gonna have crazy levels of specialization because there's gonna be so much inference demand. And that's gonna unlock just like super exciting levels of performance.”
Amazon SageMaker Inference: 2026 year-to-date launches in review

AWS Machine Learning Blog · Sep 18, 2026

SageMaker AI shipped 13 new inference capabilities in 2026 across managed endpoints and HyperPod.

“Generative AI inference is uniquely hard: models are tens to hundreds of gigabytes, latency requirements are measured in tokens per second, cold starts can span multiple minutes as containers and weights transfer, GPU capacity is constrained, and traditional monitoring tools expose none of the token-level signals that matter in production.”
Cerebras Supernova 2026

Cerebras · Aug 21, 2026

Cerebras claims it has eliminated the speed-throughput tradeoff in AI inference

“In AI, speed is productivity.”
NVIDIA Achieves Leading Agentic Coding Performance on First Agentic AI Benchmark

NVIDIA Developer Blog · Jun 12, 2026

NVIDIA leads agentic coding performance on AA-AgentPerf, the industry's first multi-vendor agentic AI benchmark.

“Artificial Analysis AgentPerf (AA-AgentPerf) offers the industry's first multi-vendor open benchmarks profiling trajectories that are representative of real-world AI agent coding tasks.”
AI Grid 101: Top 5 Things You Need to Know

NVIDIA GTC · Jun 09, 2026

NVIDIA introduces the 'AI grid' reference design, positioning telcos to become intelligence providers in the token economy.

“There has truly never been a better time to be in Telecom.”
Run Step 3.7 Flash on NVIDIA GPUs with Enterprise-Ready Multimodal AI

NVIDIA Developer Blog · May 29, 2026

StepFun's Step 3.7 Flash 198B multimodal model is now available on NVIDIA-accelerated infrastructure

“AI applications are moving beyond text generation to multimodal systems that can perceive, search, and reason across images, documents, video, and language in real time”
Scale AI with Google's TPU software stack

Google Developers (Google I/O) · May 21, 2026

Google reveals TPU V8 splits into training (8T) and inference (8I) specialized variants

“a lot of the intelligence is actually coming from inference”
Cerebras Explains | What Is an Inference API?

Cerebras · Oct 01, 2026

Cerebras inference API delivers 1000+ tokens/sec via one-line OpenAI migration

“changing one line of your base URL, and suddenly your app is running at over 1000 tokens per second”
Cerebras Explains | What is Fast AI Inference?

Cerebras · Sep 25, 2026

Cerebras runs models at over 1000 tokens per second for near-instantaneous inference

“We run models at over 1000 tokens per second, so thinking seems instantaneous.”
Shipping custom models at scale from fine-tuning to inference | BRK234

Microsoft Developer (Build) · Jun 04, 2026

Fireworks AI's managed inference and training service is now available on Microsoft Azure.

“Fireworks is a managed service where we do both AI inference performance as well as AI training all as a managed service that you can now use on Microsoft Azure.”
Reliable LLM Inference at Scale

Databricks Blog · May 27, 2026

Databricks has built a proprietary inference platform serving frontier AI models at scale

“At Databricks, we've built a unique inference platform that serves every frontier”
Cerebras Explains | What Is LLM Inference?

Cerebras · Sep 30, 2026

LLM inference is bottlenecked by memory bandwidth, not compute speed

“LLM inference is actually limited by memory bandwidth, not computational speed.”
Cerebras Explains | What Is AI Inference?

Cerebras · Sep 30, 2026

Cerebras positions inference speed and cost as the critical battleground for AI infrastructure

“output is where speed and cost directly impact your users”