Ship AI with confidence
Apple introduces a new Evaluations Framework at WWDC26 to measure model outputs, including model-judge evaluators.
“if models are judging other models, who's judging the model judges?”
Apple unveiled an Evaluations Framework at WWDC26 that lets developers measure model responses using metrics, evaluators, and model-judge scoring to ship AI features with more predictable results. It signals Apple's push to give developers tooling for testing on-device AI, but as a promotional teaser it carries limited industry-moving weight.