5 Papers That Show Where AI Research Is Heading Right Now
AlphaZero-style self-play, unbiased by human data, is the likely path to far more intelligent systems and possibly AGI.
“alpha zero unbiased by um humans meandering is uh the way we'll get to much more intelligent systems, maybe even dare say agi”
An informal YC research talk intro frames self-play (AlphaZero-style) and memory as the frontier directions for LLMs, arguing that training only on human-generated solutions can't feasibly reach the full solution space even with unlimited test-time compute or recursive self-improvement. It positions 'intelligence per sample' and 'intelligence per watt' as the two key remaining problems. The framing reflects a live research debate but is a casual session preamble rather than a concrete finding or announcement.