What Is Model Evaluation (Benchmarks and Evals)?
Model evaluation is how a model’s capability and behaviour are measured — standardised benchmarks for knowledge and reasoning, task-specific "evals" for a particular use, and human or model-graded judgement for open-ended quality. The hard part is not running an eval but trusting it: a single score hides where a model is strong or weak, and holistic evaluation argues for measuring many axes — accuracy, robustness, bias, efficiency — rather than one leaderboard number.
Why a benchmark number is not the truth
Two problems undermine naive evaluation. The first is contamination: once a benchmark is public, its answers leak into training data, so a high score may reflect memorisation rather than capability, and benchmarks decay as a signal the moment they are popular. The second is construct validity — a benchmark measures its own narrow task, which may not predict the real one you care about.
The research answer, exemplified by holistic evaluation frameworks, is to measure many models on many scenarios across multiple axes and report the distribution, not a rank. For agents the lesson is sharper still: a capability benchmark says little about whether an agent behaves safely under adversarial inputs, which needs its own evals.
Related standards
Questions
Does a top benchmark score mean the best model for me?
Not necessarily — benchmarks can be contaminated and may not match your task. Build a small eval on your own data.
Why do benchmarks get easier over time?
Their contents leak into training data (contamination), inflating scores without a real capability gain.
Keep reading
By Michael Gord · published 2026-10-09 · part of the Agentic Encyclopedia. Dates are the day of publication; events are cited at their own dates.