From RL to IRL — Gaurav Mishra, Amazon AGI Lab
Amazon AGI Lab researcher details failure modes of RL-trained agents in real-world deployment
“That's why we've been able to train really compelling coding agents using RL.”
Gaurav Mishra from Amazon AGI Lab (ex-Google DeepMind, 10+ years) outlines when RL outperforms supervised fine-tuning: tasks with verifiable outcomes, multiple solution paths, and reasoning-heavy domains—noting coding fits this paradigm especially well. The talk promises to cover what breaks when RL-trained agents move from training to real-world deployment, but the transcript is cut off before reaching that core thesis. The setup is solid practitioner framing, but the most actionable signal—the failure modes—is missing from this excerpt.