Beyond Static Intelligence: Evaluating Continual Learning — Parth Asawa, UC Berkeley
Current LLM benchmarks fail to measure learning ability across sequential tasks over time.
“imagine that every time you do something, you completely forget your memory... That's the premise under which we're evaluating language models today.”
UC Berkeley PhD student Parth Asawa argues that standard LLM benchmarks evaluate tasks in isolation, effectively assuming models have no memory between evaluations—a flawed premise for assessing real-world learning ability. He proposes continual learning evaluation, measuring whether models improve performance as a function of prior experience rather than treating each task as stateless. The critique is methodologically interesting but the talk appears to be early-stage academic framing without concrete benchmark releases or empirical results shown.