Google DeepMind launches Gemini 4 Argon with industry-first 1M output tokens and SOTA benchmarks
long-context
9 tracked signals on long-context.
AI genomic language models like Evo 2 can now generate functional virus genomes, making bio-security an AI arms race.
“cyber-warfare defensive capabilities need to be open and to keep pace with frontier models' attack capabilities”
Alibaba previews Qwen4 architecture via 176B MoE model with 1M token context
Memory harnesses for local models can solve context rot in long-horizon research agents.
“that makes this issue of dealing with context rot a priority”
RAG remains essential in high-stakes domains like tax law where every answer needs verifiable legal citations.
GPU-aware attention architecture co-design is now the key lever for long-context inference performance
“Shaping model architecture around how GPUs execute it is the premise of AI model co-design.”
Hugging Face releases LFM2.5-Encoders optimized for fast long-context inference on CPU hardware
MiniMax M3 launches on NVIDIA Blackwell infrastructure as a single multimodal system for long-context reasoning and agentic workflows.
“MiniMax M3—available on NVIDIA accelerated infrastructure including NVIDIA Blackwell—changes this by enabling a single multimodal system capable of long-context reasoning”
Together AI details research scaling long-context training to 5 million token sequence lengths via context parallelism.