Video Has No Memory. Here's How We Built One. — James Le, TwelveLabs
TwelveLabs argues video AI needs a memory layer treating video as spatial-temporal volumes, not frame sequences
“video is not a bag of frames”
James Le of TwelveLabs (Series B) presented on the architectural gap in video AI: most systems treat video as stacks of images plus transcripts, discarding the continuity and spatial-temporal relationships that define video. He argues building a true memory layer requires preserving relationships across visual, speech, motion, OCR, and metadata modalities at scale. The framing is relevant for engineers building video-native AI pipelines, though the talk was cut off before detailing the specific solution.