The Hallway Track
Engineering Insights

Your LLM App Returned 200 OK. It Was Still Wrong. — Marina Petzel, Datadog

AI Engineer · Oct 03, 2026 · Engineering Insights

Traditional golden signals are insufficient to monitor LLM app quality and safety

“a 200 OK response from your server doesn't necessarily mean that the response was useful or correct for your end user”

Datadog's Marina Petzel argues that classic monitoring metrics (latency, errors, traffic, saturation) fail to capture the unique failure modes of generative AI apps: non-deterministic outputs, variable token costs, new attack vectors like prompt injection and jailbreaking, and subjective output quality. She calls for continuous quality assessment integrated directly into the monitoring stack. This matters because it signals a growing industry need for LLM-native observability tooling beyond HTTP status codes.

observability LLM monitoring prompt injection cost monitoring production AI

Watch / read the original source →