Google released DiffusionGemma, an open-weight Apache 2 diffusion-based text generation model.
“That research has returned in the best possible way: as a new open weight (Apache 2 licensed) Gemma model”
9 tracked signals on inference-speed.
Google released DiffusionGemma, an open-weight Apache 2 diffusion-based text generation model.
“That research has returned in the best possible way: as a new open weight (Apache 2 licensed) Gemma model”
NVIDIA's Nemotron 3.5 Lightning NVFP4 delivers 4x faster throughput compressed from 66GB to 22GB
“preserves accuracy while unlocking up to 4x faster throughput”
Open source models still trail closed models on general agentic tasks but can outperform them once fine-tuned for a specific domain.
“because they're open and because you own the weights, you can actually fine-tune or RL them on your specific domain, which can actually make them perform better than those same closed models”
Nemotron-Labs diffusion language models promise near-instant text generation speeds
Cerebras achieves fast TTFT and over 1000 tokens per second simultaneously
“Fast first token and over 1000 tokens per second after that. Responsive start and instant end.”
Cerebras claims 1,000+ tokens/sec creates instant responses, fundamentally changing the AI interaction experience
“This is the difference between waiting for AI and working with it.”
A new AI called Jev makes decisions instead of generating text, running up to 200x faster than chatbots.
“It pops out instantly, up to about 200 times faster than our chatbots today.”
LFM2.5-DSpark achieves up to 3.2x faster inference speeds over baseline
A simple HTML tool lets users visualize LLM token output speeds from 5 to 800 tokens/second