OpenAI's Jalapeño chip claims 1.5–1.9x better performance per watt than NVIDIA GB200/GB300.
hardware
44 tracked signals on hardware.
OpenAI's custom Jalapeño chip delivers industry-leading AI inference speed and efficiency.
NVIDIA Vera Rubin delivers 10x measured performance per watt versus Blackwell on CoreWeave
“Not projections, not estimates, but real silicon measurements.”
Memory prices up 500% in 12 months as hyperscalers lock in all 2027 DRAM production capacity
“Some are calling it the RAMpocalypse; I prefer "RAMageddon."”
NVIDIA Rubin GPU architecture is purpose-built for agentic AI workloads at scale
“What began as discrete AI model training and human-facing chat interfaces has evolved into always-on AI factories dedicated to producing intelligence at scale.”
Apple sued OpenAI for trade secret theft tied to its $6.5B hardware push
“OpenAI believes the product's defining feature will be its personality and ability to connect on a human-like level with users.”
a16z launches Machine Age Fund targeting AI infrastructure from chips to copper mines
“the model is no longer the bottleneck and in fact using AI these models are getting better faster and faster and faster. Now the bottleneck is all what I call south of the model”
NVIDIA Vera Rubin and Blackwell set new performance-per-watt benchmarks for agentic AI workloads
“across 100 trillion tokens of real-world usage, OpenRouter's State of AI report found that average prompt tokens per request grew roughly fourfold”
NVIDIA Groq 3 LPX accelerator enables ultrafast interactive inference on Vera Rubin NVL72 at long context
“NVIDIA Vera Rubin NVL72, the most versatile machine ever built, delivering high throughput and interactivity across the widest range of AI workloads—from small to large models, both open and closed.”
NVIDIA RTX Spark is a 1-petaflop superchip for running advanced AI agents locally on laptops
“The reinvention of the personal computer has begun and it starts with the RTX Spark”
NVIDIA's Vera CPU Olympus cores are optimized for agentic AI workloads requiring high single-thread performance
“Agentic AI shifts more of the critical execution path onto the CPU.”
NVIDIA GB300 NVL72 sets world record pre-training DeepSeek-V3 671B at 1,648 TFLOPs per GPU
“As compute per token falls, communication increasingly determines how efficiently models scale across thousands of GPUs.”
NVIDIA's Vera CPU targets agentic AI and reinforcement learning workloads in AI factories.
“Each wave of AI has created a new scaling law.”
AI data center demand for HBM memory is pricing out cheap consumer smartphones globally.
“a single gigabyte of HBM consumes more than three times the wafer capacity that a gigabyte of DDR or LPDDR does”
Physical experiments are becoming the bottleneck as AI accelerates chip design iteration speed
“as we accelerate intelligence and shorten the time between these experiments, it becomes increasingly important to accelerate the tests themselves”
Fractile is building inference chips targeting 3-6 month architectural leads over competitors
“if you find a way to structurally secure a 3-6 month lead, you will win in all these implementations.”
NVIDIA delivers Vera Rubin and Vera C servers to AWS in Seattle
NVIDIA BlueField-4 DPU enables dedicated networking for agentic AI factory infrastructure
“Agentic AI factories connect diverse users, agents, applications, data sources, and storage systems to massively accelerated compute at multi-terabit bandwidth per server, making dedicated DPU processing essential for line-rate networking, storage, and security.”
Cerebras claims it has eliminated the speed-throughput tradeoff in AI inference
“In AI, speed is productivity.”
US AI supply chain faces extreme concentration risk beyond chips, especially in robotics components dominated by China
“Robotics is an incredibly promising industry and the supply chain is right now completely dominated by China.”
NVIDIA Nemotron 3 Ultra tops open models for agentic RTL chip design coding.
Google reveals TPU V8 splits into training (8T) and inference (8I) specialized variants
“a lot of the intelligence is actually coming from inference”
GPU design shifted from compute efficiency to memory bandwidth after transformers emerged
“attention is equal to n-squared”
Cerebras claims 1,000+ tokens/sec creates instant responses, fundamentally changing the AI interaction experience
“This is the difference between waiting for AI and working with it.”
Cerebras runs models at over 1000 tokens per second for near-instantaneous inference
“We run models at over 1000 tokens per second, so thinking seems instantaneous.”
NVIDIA previews RTX Spark laptops with 128GB memory for 12K video editing and real-time DLSS 4.5 rendering.
“we can make decisions as creators because we see how the final content will look like”
Meta engineered 7mm-wide steel-can batteries to power all-day AI features in smart glasses.
“Smart glasses need a battery that can claim every micron of space – something rigid, precise, and shaped to the product rather than the other way around.”
Google's quantum AI team, founded 2012, employs Nobel Prize-winning hardware scientists.
“Michelle Devet is actually our chief hardware scientist on the quantum computing team at Google.”
LLM inference is bottlenecked by memory bandwidth, not compute speed
“LLM inference is actually limited by memory bandwidth, not computational speed.”
At 10 million server scale, one-in-a-million daily failures happen 10 times per day
“if it's a one in million chance of happening in a day, that means it's happened 10 times today.”
AI inference speed is an infrastructure problem; Cerebras claims 1,000+ tokens per second vs GPUs
“run a great model on infrastructure built for speed. On Cerebrus, that's over 1,000 tokens per second—a speed that GPUs have a hard time approaching.”
Valar built a nuclear reactor protection system for $400K vs $5M vendor quote in far less time.
“we're going to have to build our own RPS. And we spent about $400,000 on it.”
Cerebras shows CS-3 wafer-scale chip assembly with custom liquid cooling manifold
“It's very very critical that we assemble it so there's no leak”
Cerebras CS-3 supercomputers undergo 4-5 days of liquid cooling stress tests before shipment
“it takes us through multiple days of rigorous testing and eventually after cooling we pass and we ship it”
Cerebras CS-3 AI supercomputer uses wafer-scale engine requiring 26 kW of vertical power delivery
“A total of about 26 kW total into the chip.”
An engineer built a physical, AI-native terminal device with dual displays to interact with an LLM.
“I just wanted to build a device which is physical and AI native like the device which comes from the future.”
Applied Electrodynamics built a radar camera that images through walls in 3D using GPUs
“the whole GPU revolution allows you to process them "on the fly" and generate these images”
Cerebras and Flex began manufacturing partnership in 2024, scaling CS-2 production in Silicon Valley.
“our partnership started in 2024 with our first shipment out the door in October of 2024”
Microsoft's new Surface Laptop Ultra is powered by the NVIDIA RTX Spark chip in an 18mm serviceable aluminum body.
“This laptop uses the largest thermal solution ever in a Surface Laptop so the RTX Spark can sustain high performance during longer sessions.”
A .NET conference talk covers building apps that communicate with hardware using MAUI and Uno across platforms.
“they've got it to where it can run net packages or net frameworks that can run on very limited hardware it it runs everywhere”
ARM launches Performix, a performance analysis toolkit built with Microsoft to tune workloads on Cobalt.
“Performance analysis can be thought of like a crime investigation.”
Hover Arrow raised $1.37B in letters of intent for rocket-speed cargo delivery.
“we signed commercial letters of intent worth $1.37 billion”
The Motorola 68000, launched in the late 1970s, is considered the first modern processor.
“It was capable of multitasking, virtual memory, and running multiple OSs many years before anyone else, even Intel.”
Hart Aerospace's electric aircraft is the world's largest by a factor of two.
“It's surreal. Yeah. I came to YC with that 3D printed plane uh that was like this this size and now it's you know 100 foot wingspan hurdling down the runway doing taxi testing.”