Two Bugs That Hid in Plain Sight: A vLLM Debugging Detective Story — Asaf Gardin & Yuval Belfer
AI21 engineers traced rare 'one-in-a-thousand gibberish' output bugs unique to vLLM during Jamba RL training.
“there is no crash there's no warning no error and there's high confidence that's not a quality issue”
AI21 engineers describe hunting down silent, load-dependent 'imposter request' bugs that produced rare gibberish output only in vLLM while doing RL (GRPO) training on their hybrid Jamba (transformer + Mamba SSM) model. The talk highlights how the hardest production issues are engineering problems—high-confidence but bad output with no crash or error—rather than model quality problems. Useful practitioner insight for teams running inference at scale, though it is a niche debugging story rather than a major industry signal.