Stop Rationing Tokens: Let the Harness Pick the Model — Kimchi by Cast AI
Cast AI's Kimchi agent routes tasks to cheap open-source models to eliminate developer token rationing
“Our job is not to prohibit a developer from using an agent for coding. Our job is to make it so they can use it as much as they want, whenever they want, without any restrictions.”
Cast AI co-founders presented Kimchi, a coding agent built on a hybrid model strategy — open source for most tasks, proprietary only for complex cases — to keep token costs low enough that companies can give developers unrestricted access. The talk frames runaway enterprise AI spending (citing an Indian firm spending $500M/month on Anthropic and Uber burning its annual Anthropic budget in four months) as the forcing function. The core engineering insight is that cost-per-task optimization, not cost-per-token, is the right metric for evaluating model routing decisions.