Your LLM App Returned 200 OK. It Was Still Wrong. — Marina Petzel, Datadog
Traditional golden signals are insufficient to monitor LLM app quality and safety
“a 200 OK response from your server doesn't necessarily mean that the response was useful or correct for your end user”
Datadog's Marina Petzel argues that classic monitoring metrics (latency, errors, traffic, saturation) fail to capture the unique failure modes of generative AI apps: non-deterministic outputs, variable token costs, new attack vectors like prompt injection and jailbreaking, and subjective output quality. She calls for continuous quality assessment integrated directly into the monitoring stack. This matters because it signals a growing industry need for LLM-native observability tooling beyond HTTP status codes.