Evals in AI: A Deep Dive — Tejas Kumar, IBM
AI evals provide reliability in advance, not just post-hoc validation
“Evals provide reliability, but not just any reliability, but reliability in advance.”
IBM AI engineer Tejas Kumar opens a deep-dive session on AI evaluations at AI Engineer, framing evals as the mechanism for achieving advance reliability in AI systems rather than reactive quality checks. Drawing on real-world experience building assessment and RAG pipelines on IBM's WatsonX team, the talk promises to cover eval architectures and methods grounded in industry research. The content is early-stage framing and the transcript cuts off before substantive technical content is delivered.