The Hallway Track
Engineering Insights

Stop Burning Tokens: Why self-improvement needs domain expertise first - Annabell Schäfer, Langfuse

AI Engineer · Jul 18, 2026 · Engineering Insights

AI self-improvement loops require domain-specific evaluation functions, not just agentic design.

“does the code compile or not”

Langfuse growth engineer Annabelle Schäfer argues that the 'design loops not prompts' trend works in coding because it has a clear binary target function (does code compile?), but most other domains lack this clarity. Teams that invest early in defining domain-specific evaluators are the ones that can ship with confidence and continuously improve their AI applications. The core prescription is: before scaling agentic loops, build the evaluation infrastructure that tells you whether the loop is actually heading in the right direction.

agentic-loops evaluation domain-expertise observability langfuse self-improvement

Watch / read the original source →