Memory prices up 500% in 12 months as hyperscalers lock in all 2027 DRAM production capacity
“Some are calling it the RAMpocalypse; I prefer "RAMageddon."”
Memory prices up 500% in 12 months as hyperscalers lock in all 2027 DRAM production capacity
“Some are calling it the RAMpocalypse; I prefer "RAMageddon."”
NVIDIA BlueField-4 DPU enables dedicated networking for agentic AI factory infrastructure
“Agentic AI factories connect diverse users, agents, applications, data sources, and storage systems to massively accelerated compute at multi-terabit bandwidth per server, making dedicated DPU processing essential for line-rate networking, storage, and security.”
Google reveals TPU V8 splits into training (8T) and inference (8I) specialized variants
“a lot of the intelligence is actually coming from inference”
Railway is repositioning as agent-native cloud infrastructure with 70% margins and 3M users.
“the activation energy to ship something to production should be near zero.”
At 10 million server scale, one-in-a-million daily failures happen 10 times per day
“if it's a one in million chance of happening in a day, that means it's happened 10 times today.”
NVIDIA and OpenAI are building the largest data center ever at gigawatt scale.
“When it's fully built out, it will be the largest data center ever built.”
OpenAI's Jalapeño chip claims 1.5–1.9x better performance per watt than NVIDIA GB200/GB300.
OpenAI's custom Jalapeño chip delivers industry-leading AI inference speed and efficiency.
Anthropic signed $1.25B/month compute deal with SpaceX's Colossus clusters through May 2029
“Cloud Services Agreements with Anthropic PBC...with respect to access to compute capacity across COLOSSUS and COLOSSUS II...the customer has agreed to pay us $1.25 billion per month through May 2029”
California high-speed rail project estimated to cost $236 billion without completed segments.
“Governor Nuomo in a private video... doesn't believe that this project will ever be done in our lifetime.”
NVIDIA is shaping infrastructure for agentic AI with a new platform called Vera Rubin.
“The value gets generated in the computing and the transformation of the actual data into a calculation.”
Meta introduces ZGateway, a proxy for ZippyDB traffic management.
A new industrial revolution is unfolding powered by AI technologies.
“No company can build this infrastructure alone.”
NVIDIA Vera Rubin delivers 10x measured performance per watt versus Blackwell on CoreWeave
“Not projections, not estimates, but real silicon measurements.”
NVIDIA Rubin GPU architecture is purpose-built for agentic AI workloads at scale
“What began as discrete AI model training and human-facing chat interfaces has evolved into always-on AI factories dedicated to producing intelligence at scale.”
Apple sued OpenAI for trade secret theft tied to its $6.5B hardware push
“OpenAI believes the product's defining feature will be its personality and ability to connect on a human-like level with users.”
Databricks launched Omnigent, an open-source 'meta-harness' layer on top of the agentic stack to make agents effective at scale.
“we call it a meta harness of harnesses, if you know what an agent harness is.”
NVIDIA's Jensen Huang frames the AI factory as the largest infrastructure buildout in history and the best enterprise investment of the next decade.
“you can't think if you don't generate words”
AWS launches AgentCore Payments enabling AI agents to autonomously execute microtransactions via stablecoins
“Amazon Bedrock AgentCore payments is the first managed service within Amazon Bedrock AgentCore that helps AI agents autonomously execute microtransaction payments for paid APIs, MCPs, and content with a few lines of code.”
AWS AgentCore Runtime Instances enable GPU-backed, 14-day multi-agent sessions on persistent EC2
“serverless sessions that cap at a few hours don't cut it”
Hyperscalers are investing all short-term operating cash flow into AI capacity as demand outpaces supply
“Hyperscalers are investing all of their short-term operating cash flow into building this capacity to meet demand that continues to outpace supply in almost every case we see”
Databricks launches Lakebase Search with full-text and vector search natively in Postgres for AI agents
NVIDIA releases open reference platform for continuous hardware-level AI agent safety monitoring
Google DeepMind adds private, server-side memory to Private AI Compute for personal AI.
“Introducing private, server-side memory to Private AI Compute for personal AI.”
Stripe engineer demos autonomous shopping agent as agentic commerce infrastructure matures
“the infrastructure for agentic transactions has been laid down by companies like Google, OpenAI, and Stripe”
a16z launches Machine Age Fund targeting AI infrastructure from chips to copper mines
“the model is no longer the bottleneck and in fact using AI these models are getting better faster and faster and faster. Now the bottleneck is all what I call south of the model”
NVIDIA NVLink Fusion enables NVHBM memory for custom XPU accelerators at hyperscale
Agents will use the web 1,000x more than humans, requiring reinvented search infrastructure
“We started parallel with the bet that agents would do it a thousandx more than humans ever have.”
OpenAI CFO frames full-stack chip-to-product compounding as path to cheaper, scalable intelligence
Meta open-sourced MetaRoCE, a new RDMA transport protocol built for million-GPU AI clusters on Ethernet.
“The fabric sees packets, but the NIC sees intent.”
NVIDIA Vera Rubin and Blackwell set new performance-per-watt benchmarks for agentic AI workloads
“across 100 trillion tokens of real-world usage, OpenRouter's State of AI report found that average prompt tokens per request grew roughly fourfold”
NVIDIA Groq 3 LPX accelerator enables ultrafast interactive inference on Vera Rubin NVL72 at long context
“NVIDIA Vera Rubin NVL72, the most versatile machine ever built, delivering high throughput and interactivity across the widest range of AI workloads—from small to large models, both open and closed.”
Agent infrastructure is now commoditized by cloud platforms, making context the new competitive frontier
“They're all taxes one has to pay in order to get an agent out there to play the game.”
Modal proposes decoupling RL rollout workers from trainer clusters to use distributed GPU capacity across datacenters.
“IO wants all four of these at the same time. Enough GPU, same region, fast fabric, and available now. Any of these like is manageable, but all four of them that are pretty hard to get at the same time.”
NVIDIA RTX Spark is a 1-petaflop superchip for running advanced AI agents locally on laptops
“The reinvention of the personal computer has begun and it starts with the RTX Spark”
Baseten raised $13B Series F as inference engineering emerges as a critical AI discipline
“How do you turn those weights from training into a product that is fast, reliable, and affordable at scale?”
Meta doubled GEM ads model training efficiency to 20-25% MFU while scaling FLOPs 4x in 12 months
Y Combinator is seeking founders to build offshore AI compute flotillas on the ocean
“It sounds crazy, but we think part of the answer may be to move compute offshore.”
NVIDIA's Vera CPU Olympus cores are optimized for agentic AI workloads requiring high single-thread performance
“Agentic AI shifts more of the critical execution path onto the CPU.”
NVIDIA GB300 NVL72 sets world record pre-training DeepSeek-V3 671B at 1,648 TFLOPs per GPU
“As compute per token falls, communication increasingly determines how efficiently models scale across thousands of GPUs.”