Scale AI with Google's TPU software stack
Google reveals TPU V8 splits into training (8T) and inference (8I) specialized variants
“a lot of the intelligence is actually coming from inference”
Google detailed its TPU V8 hardware stack at Google I/O, revealing a bifurcated architecture with TPU 8T optimized for large-scale training throughput and TPU 8I optimized for low-latency inference. The talk highlighted a shift where inference compute is increasingly the source of model intelligence, driven by thinking models consuming large token counts during reasoning. This is a foundational infrastructure signal for teams building or scaling on Google Cloud TPUs.