Xiaomi released MiMo-V2.6-Pro, a natively omnimodal 1T-A42B open-weights model trained for $3M.
“MiMo-V2.6-Pro is our most capable model to date, while MiMo-V2.6-Flash strikes the best balance between intelligence, efficiency, and cost.”
38 tracked signals on reinforcement-learning.
Xiaomi released MiMo-V2.6-Pro, a natively omnimodal 1T-A42B open-weights model trained for $3M.
“MiMo-V2.6-Pro is our most capable model to date, while MiMo-V2.6-Flash strikes the best balance between intelligence, efficiency, and cost.”
GLM-5.3 proves post-training RL on long-horizon tasks beats parameter scaling for reasoning
“Parameter count is only meaningful alongside three others — how much data you have, where you intend to spend your compute, and who will run the model, under what conditions.”
Microsoft Foundry assembles agents, tuning, evals, and OpenEnv into an owned reinforcement-learning loop that improves over time.
“the durable asset is not the model you rent, it is the learning loop you own”
Microsoft launches Frontier Tuning, letting enterprises reinforcement-fine-tune AI models on their own M365 data and workflows.
“With Frontier Tuning, we're making it possible for you to create your own enterprise AI.”
GEPA proposes reflective optimization in text space to overcome RL's sample inefficiency
“instead of using only a zero or one reward signal, we can make a language model or agent analyze the entire execution process to understand what worked and what didn't”
Modal proposes decoupling RL rollout workers from trainer clusters to use distributed GPU capacity across datacenters.
“IO wants all four of these at the same time. Enough GPU, same region, fast fabric, and available now. Any of these like is manageable, but all four of them that are pretty hard to get at the same time.”
Kimi K3 is trained to be comparable to Claude Opus 4.8 using novel 3x3 domain expert distillation.
“they've trained a model that is probably comparable to Opus 4.8”
Former OpenAI VP and Gemini pre-training lead co-founded Core Automation to build an automated AGI lab.
“The first step to replacing transformers is appreciating deeply how far they were able to carry us.”
A new benchmark, SocioHack, shows RL-trained LLMs can discover loopholes that game society's rule systems while staying formally compliant.
“an RL-trained model discovers strategies that remain formally compliant, yet undermine the intended purpose of those systems”
Microsoft Foundry adds reinforcement learning and a low-level training API to turn production agents into cheaper, faster models.
“Think of it as PyTorch as a service.”
Microsoft launches Frontier Tuning: RL-based AI customization within enterprise compliance boundaries
“a new approach to making AI work the way your business does by applying reinforcement learning inside your compliance boundary with your own data, processes, and conventions”
Numerical mismatch in inference log-probabilities is a hidden bug plaguing large-scale asynchronous RL runs.
“Hopefully next composer versions are going to be our own model instead of basing it off an open source base.”
Cursor ships RL model updates as compressed weight deltas ~20x smaller than the full 1TB model.
“My delta might be like 20 times smaller than was shipping the full model with and this makes it practical”
TypeSafe's Jev model uses RLCD to deliver reliable machine-native intelligence for software automation
AWS achieves 40% throughput gain for MoE reinforcement learning using EKS, EFA, and DeepEP
Rich Sutton argues continual learning is just learning — the field's framing is the anomaly
“I'm not weird. The field is weird.”
Google DeepMind marks 15 years of games AI research with new studio partnerships
Amazon Nova Forge enables custom reward functions for multi-turn reinforcement learning via BYOO capability
“A subtly wrong reward can quietly teach the wrong thing while every training curve looks healthy.”
On-policy distillation uses RL policy shift signals to train student models without full RL cost
“the student might be stronger than the teacher in this case”
AI architectural research fails because it tests at insufficient compute scale
“to get to any interesting results you need certain level of compute to even see the capabilities in the model”
David Brumley is using RL environments to teach AI to find vulnerabilities at machine speed
“we're pushing out programs faster than ever and so we need to be able to check them at machine speeds in scale”
The traditional base-model paradigm of web-scale pre-training is being displaced by post-training
“RL was mostly just a cherry on top, shaping the flavor of the interactions more than conferring extra knowledge or quality onto the base model itself.”
Prime Intellect is extending RL to real-world tasks lacking verifiable reward signals.
An RL agent autonomously diagnoses and remediates ETL pipeline failures with safety bounds, cutting manual recovery from ~2.5 days.
“The central question is simply whether an agent can act, but whether it can act usefully, explainably, and within the boundaries that an operation would actually trust.”
Low-quality RL training environments and broken harnesses actively degrade models and ruin training runs.
“researchers don’t want your broken RL environments because they will make our models worse”
Post-training open-source reasoning models with reinforcement learning keeps agent costs flat while quality improves in Foundry.
“deploying a model into production is just the beginning”
Cursor uses online (real-time) RL only to polish already-shipped models, not build them from scratch.
“That's kind of the paradox of online RL or how we like to call it real time is that, you know, we can't use this to really create the model from scratch because users need to be using the model.”
Multi-turn RL on SageMaker lets small models match frontier reliability for search agents
“Fine-tuning offers a third path: you teach a small model your tools and environment directly. The result is a small model's speed and cost with the reliability that would otherwise require a frontier model.”
SkyRL on SageMaker HyperPod trains vision-language models to 95% maze-solving accuracy via GRPO
Amazon AGI Lab researcher details failure modes of RL-trained agents in real-world deployment
“That's why we've been able to train really compelling coding agents using RL.”
NVIDIA combines imitation learning and goal-based RL to train parkour AI on just 30 seconds of data
“Systems that copy us are beautiful, but brittle, repetitive. Systems that chase goals adapt, but often stop moving like us.”
General Reasoning is building RL systems to scale agents toward long-horizon tasks
Hugging Face launches a recurring live tutorial series on fine-tuning coding agents for continual learning.
“I think you could come from from zero here.”
AlphaZero-style self-play, unbiased by human data, is the likely path to far more intelligent systems and possibly AGI.
“alpha zero unbiased by um humans meandering is uh the way we'll get to much more intelligent systems, maybe even dare say agi”
AWS shows how to train Unitree H1 humanoid robot policies using NVIDIA Isaac Lab on SageMaker HyperPod and Training Jobs.
“GPU-accelerated simulation can compress months of learning into hours”
Hugging Face is teaching GRPO reinforcement learning for agent training in a live stream series
“it's pretty straightforward to learn from the available options and you can apply it on most use cases”
NVIDIA gave Unsloth a DGX Station to accelerate its model quantization and RL research.
“We're going to be use utilizing this to do all of the models. And just provide more and more for the community.”
AWS releases bootloader letting developers install custom OS on DeepRacer RL devices