Why Video Agent models are next — Ethan He, xAI Grok Imagine
Video models get their intelligence from LLMs, so the next Sora won't be a better video model but a video agent.
“In the near term, the next Sora won't be a better video model, but a video agent.”
xAI's Ethan He, who built Grok Imagine in three months, argues that video models derive their intelligence primarily from LLMs rather than video training data, and that generative media will follow AI coding's path from one-shot output to agentic systems that plan, generate, edit, and iterate. This reframes the video-gen roadmap around orchestration and language models over diffusion improvements, signaling 'video agents' as the next industry trend.