What Is Model Evaluation (Benchmarks and Evals)?

Model evaluation is how a model’s capability and behaviour are measured — standardised benchmarks for knowledge and reasoning, task-specific "evals" for a particular use, and human or model-graded judgement for open-ended quality. The hard part is not running an eval but trusting it: a single score hides where a model is strong or weak, and holistic evaluation argues for measuring many axes — accuracy, robustness, bias, efficiency — rather than one leaderboard number.

Forms
Benchmarks, task evals, human/model grading
Risk
One score hides the distribution of behaviour
Benchmark decay
Contamination inflates results over time

Why a benchmark number is not the truth

Two problems undermine naive evaluation. The first is contamination: once a benchmark is public, its answers leak into training data, so a high score may reflect memorisation rather than capability, and benchmarks decay as a signal the moment they are popular. The second is construct validity — a benchmark measures its own narrow task, which may not predict the real one you care about.

The research answer, exemplified by holistic evaluation frameworks, is to measure many models on many scenarios across multiple axes and report the distribution, not a rank. For agents the lesson is sharper still: a capability benchmark says little about whether an agent behaves safely under adversarial inputs, which needs its own evals.

Related standards

Liang et al., 2022 — Holistic Evaluation of Language Models (HELM)

Questions

Does a top benchmark score mean the best model for me?

Not necessarily — benchmarks can be contaminated and may not match your task. Build a small eval on your own data.

Why do benchmarks get easier over time?

Their contents leak into training data (contamination), inflating scores without a real capability gain.

Keep reading

related
What Is an AI Hallucination?
related
What Are AI Guardrails?
related
What Is Chain-of-Thought Prompting?
related
Supervised vs Unsupervised Learning: What’s the Difference?
referenced by
What Is RLHF (Reinforcement Learning from Human Feedback)?
Governance & Accountability
AI Agent Governance and Accountability
Governance & Accountability
The Consent Layer of the Agentic Internet
Governance & Accountability
What Is an Agent Policy Engine (and Why Deny-by-Default)?
Governance & Accountability
Human-in-the-Loop vs On-the-Loop vs Autonomous Agents
Governance & Accountability
When Should an AI Agent Escalate to a Human?
Governance & Accountability
How Do You Audit an Autonomous AI Agent?
Governance & Accountability
Kill Switches and Dead-Man’s Switches for Autonomous Organizations
Governance & Accountability
Agent Compliance and Regulation
Governance & Accountability
Managing the Risk of Autonomous Agents
Governance & Accountability
What Is Prompt Injection?
Governance & Accountability
What Is Capability-Based Security?
Governance & Accountability
What Is the Principle of Least Privilege?

By Michael Gord · published 2026-10-09 · part of the Agentic Encyclopedia. Dates are the day of publication; events are cited at their own dates.