Term
Benchmark (AI)
An AI benchmark is a standardized task suite that makes language models comparable — same prompts, same scoring, transparent results for reasoning, knowledge, or coding.
Benchmark (AI) — explained in detail
AI benchmarks consist of fixed task sets with unambiguous scoring. Models receive identical inputs, and their responses are checked against gold solutions (multiple-choice accuracy, test-suite pass rate, judge score). This makes models comparable and forms the basis for leaderboards. The most important benchmarks in 2026 include MMLU (knowledge), GPQA-Diamond (PhD-level reasoning), SWE-bench Verified (coding agents), MATH and HumanEval (code), as well as the Chatbot Arena with Elo rankings from real user votes.
Example / Practical use
When selecting a model for a project, teams typically look at a mix: GPQA for reasoning, SWE-bench Verified for engineering tasks, Arena Elo for general answer quality. MMLU has lost its discriminating power because top models now score above 90 percent — the benchmark has “saturated”. Newer tests such as HLE (Humanity’s Last Exam) and LiveCodeBench are replacing it as differentiators.
Distinction from related terms
A benchmark is static and curated once; an eval suite (see Evals) is typically project-specific and continuously adjusted. A leaderboard aggregates benchmark results but is not a benchmark itself. Risks: benchmark contamination (training data contains test items) and overfitting to popular suites — hence the move to verified variants and private hold-out sets.
Discover more
AI Workflows by Keyword: How We Make Recurring Routines Enforceable
A typed keyword triggers a fixed AI routine — and every single step must be committed before the next one appears. Why that's the actual trick.
GlossaryGround Truth
A reference taken to be correct, against which model outputs are measured. In AI code analysis it is a piece of code with deliberately injected, known defects, so you can measure how many real problems a model actually finds.
EncyclopediaAgent Evaluation — Task Success Rate, Eval Harnesses, Benchmarks
How to evaluate AI agents: eval harnesses, task success rate, benchmarks like GAIA and Tau-Bench, LLM-as-a-judge — and how it differs from pure model quality.