Benchmark
A fixed set of test questions used to compare models. Useful, gamed often, and frequently misread.
Benchmarks make claims comparable: the same exam, the same marking, several systems. They have two standing weaknesses. Test questions leak into training data, so a model may have seen the answers; and a score on a fixed set of questions says little about behavior on the messy problems people actually bring. Read a benchmark result as evidence about that benchmark, and no more than that.