Are LLM Performance Benchmarks Reliable? — Ashok Chandrasekar & Jason Kramberger, Google
Google engineers argue production-scale LLM inference benchmarks need purpose-built tooling like inference-perf.
“LLMD is a distributed inference framework um that makes production scale inference possible.”
Google engineers Ashok Chandrasekar and Jason Kramberger explain why existing LLM benchmark tools (model-server scripts, competitive analysis tools, web load testers) fall short for production-scale inference serving, and present their open-source inference-perf tool and LLMD distributed inference framework as a solution. The talk highlights the gap between simple developer benchmarks and the realistic multi-workload benchmarking needed for production inference stacks.