WWDC26: Meet the Evaluations framework | Apple
Apple's new Evaluations framework lets developers measure quality of generative-AI features powered by on-device models.
“These models break a contract that is fundamental to software testing.”
At WWDC26 Apple introduced the Evaluations framework, a Swift Testing-integrated system for measuring the quality of intelligent features built on its on-device Foundation Models, using metrics, evaluators, and model judges. It matters because generative AI breaks the deterministic input-output contract that unit tests rely on, and Apple is giving developers standardized tooling to verify probabilistic behavior at scale before shipping.