Voice In, Visuals Out: The Agony and the Ecstasy - Allen Pike, Forestwalk Labs
Audio-in, visuals-out AI experiences are now feasible, matching Karpathy's claim that voice is the preferred input and visuals the preferred output.
“voice is the human preferred input for AIs. But that we prefer visuals as the output.”— Andrej Karpathy
Allen Pike of Forestwalk Labs presents lessons from building 'voice in, visuals out' AI experiences, framed around Andrej Karpathy's argument that voice is humans' preferred AI input while visuals are the preferred output. Recent breakthroughs in models generating rich HTML, tool calls, and images make these multimodal experiences feasible, though voice-as-input remains controversial given past slow and clumsy assistants like Siri.