Every model ships its own benchmark and, where a comparable open system exists, measures it on the same data with the same grading — so the headline number is always a like-for-like comparison.