The Hallway Track
Industry Trends

5 Papers That Show Where AI Research Is Heading Right Now

Y Combinator · Jun 12, 2026 · Industry Trends

AlphaZero-style self-play, unbiased by human data, is the likely path to far more intelligent systems and possibly AGI.

“alpha zero unbiased by um humans meandering is uh the way we'll get to much more intelligent systems, maybe even dare say agi”

An informal YC research talk intro frames self-play (AlphaZero-style) and memory as the frontier directions for LLMs, arguing that training only on human-generated solutions can't feasibly reach the full solution space even with unlimited test-time compute or recursive self-improvement. It positions 'intelligence per sample' and 'intelligence per watt' as the two key remaining problems. The framing reflects a live research debate but is a casual session preamble rather than a concrete finding or announcement.

self-play reinforcement-learning AGI test-time-compute memory

Watch / read the original source →