The Hallway Track
Research Findings

Beating RL With Reflection: GEPA and Optimize Anything — Lakshya A. Agrawal, GEPA

AI Engineer · Sep 26, 2026 · Research Findings

GEPA proposes reflective optimization in text space to overcome RL's sample inefficiency

“instead of using only a zero or one reward signal, we can make a language model or agent analyze the entire execution process to understand what worked and what didn't”

GEPA introduces reflective optimization, a text-space alternative to reinforcement learning that mines rich execution traces—chain-of-thought, tool calls, error messages—rather than discarding them in favor of a scalar reward. The core problem it addresses is that most teams lack the trillions of tokens or hundreds of thousands of RL iterations traditional approaches require, especially as agent workloads grow longer and more expensive. By having a language model analyze full execution processes, GEPA aims to extract learning signal that RL-based methods like GRPO systematically waste.

optimization reinforcement-learning sample-efficiency agents reflective-optimization

Watch / read the original source →