A benchmark is a defined set of tasks and an evaluation protocol used to compare systems. A leaderboard orders submitted or measured results under a stated scoring method. A result describes a particular model or agent configuration on that test; it is not an unconditional measure of intelligence or production readiness.
The interview skill is to explain what was tested, how it was scored, what may be misleading, and which product evaluation follows. Public results help narrow candidates. Your workload, constraints and verified outcomes determine deployment.
Source review: September 24, 2026. This chapter names current versions where verified and avoids maintaining a table of rapidly changing model winners.
Read six parts of every claim
| Part | Record it | Why it changes the interpretation |
|---|---|---|
| Task/data | Dataset version, split, exclusions and release date | A repaired or harder set changes the denominator |
| System | Model revision, prompt, agent scaffold, memory and tools | An agent score is not a bare-model score |
| Resources | Tokens, attempts, time, compute and external access | More resources can buy a different outcome |
| Grading | Tests, judge, rubric, failure and tie rules | Scoring errors can reject valid work or reward shortcuts |
| Statistics | Case count, repeats, uncertainty and paired changes | A small lead need not be reproducible |
| Provenance | Who ran/checked it, when, and whether artifacts are available | A self-report and an independently reproduced run offer different evidence |
Harness means the machinery that runs tasks and grades outputs. Scaffold means the agent's surrounding orchestration, prompts, tools and control logic. They interact: sandbox resources and timeout behavior can affect an otherwise unchanged agent.
For model isolation, compare models under appropriately matched conditions. For a product decision, comparing complete stacks with different scaffolds can be valid if the workload, acceptance criteria and constraints are comparable. State which question the experiment answers; “never compare different harnesses” is too absolute.
Capability map: knowledge, reasoning and mathematics
| Benchmark family | What it tests | How to use the evidence |
|---|---|---|
| MMLU | Multiple-choice knowledge across academic subjects | Inspect subject mix, prompting and item quality |
| MMLU-Pro | Harder knowledge/reasoning questions with more answer choices | Distinguish it from original MMLU; inspect current headroom |
| MMLU-Redux | Reassessment/correction of MMLU items | Check the particular revision and subset, not a universal error rate |
| HellaSwag; WinoGrande | Commonsense completion; reference resolution | Useful scoped diagnostics, not a full assistant evaluation |
| ARC-Challenge | School science questions | Do not confuse it with ARC-AGI grid tasks |
| GLUE; SuperGLUE | Collections of language-understanding tasks | Historical comparisons and regression diagnostics |
| BIG-Bench Hard; BIG-Bench Extra Hard | Challenging collections of reasoning tasks | Name the task suite and prompt protocol |
| GPQA, including Diamond | Expert-authored science multiple choice | Subset and sample size matter; not all scientific work |
| HLE, HLE-Rolling, HLE-Diamond | Difficult academic questions across domains | Version, tools and grading conditions are essential |
| ARC-AGI series | Novel abstract reasoning tasks; newer versions include interactive environments | Read the version-specific interface, compute budget and verification status |
| CritPt | Research-oriented physics problems | Check challenge versus component scoring and verifier validity |
| GSM8K; MATH/MATH-500 | School word problems; competition mathematics | High scores on public items need robustness and contamination checks |
| AIME; HMMT; Putnam | Competition problems under specified year/selection | Exact answers and complete proofs require different grading |
| FrontierMath | Advanced mathematics, with distinct tiers and research-problem tracks | State the track and whether a result verifies an answer or a proof |
Sources and definitions: MMLU-Pro project, GPQA paper, HLE project, ARC Prize, CritPt dataset, FrontierMath tracks. The LM Evaluation Harness task catalog helps identify exact configurations for many established text benchmarks.
HLE-Diamond was released on September 22, 2026 with 1,000 questions split into knowledge and reasoning partitions. Its release reports tool-free and tool-enabled conditions separately. A score on this subset cannot be substituted for a score on the earlier 2,500-question HLE set. Release and protocol.
AIME has a small annual problem set, so one item can move a score noticeably. Repeated sampling measures model variability but does not create more distinct mathematics problems. A high answer score does not by itself establish proof-writing ability.
Coding: identify the unit of engineering work
| Family | Work performed | Important limitation |
|---|---|---|
| HumanEval / MBPP | Small function synthesis | Narrow specification and public tasks |
| HumanEval+ / MBPP+ | Expanded executable tests for those functions | Stronger tests still do not cover repository work |
| SWE-bench, Lite, Verified | Repair repository issues | Task/test validity, contamination and environment matter |
| SWE-bench Multilingual / Multimodal | Repository work across languages or visually described issues | Different tasks and modalities prevent score substitution |
| SWE-bench Pro | Longer/harder repository changes under its protocol | A newer benchmark also needs validity audits |
| SWE-Bench ProMax | Coordinated multilingual code refactoring | Refactoring correctness differs from isolated bug fixing |
| SWE-rebench / SWE-bench-Live | Refreshed repository issue evaluation | Pin task window and revision; “live” is not proof of no leakage |
| LiveCodeBench | Time-organized competitive programming and related code tasks | Algorithmic skill is not full software delivery |
| Aider Polyglot | Coding/edit tasks in several languages | Agent edit format and allowed attempts affect results |
| SciCode | Scientific programming with problem/subproblem structure | Report which scoring level is used |
| Terminal-Bench | Multi-step terminal tasks in controlled environments | Agent, resources, task revision and cost are part of the result |
The official SWE-bench site distinguishes its subsets and a matched Bash Only view. EvalPlus strengthens function-level tests. LiveCodeBench organizes problems by release time. ProMax research concerns large multilingual refactoring. These are complementary workloads, not successive names for one identical exam.
Current Terminal-Bench: the maintainer documents 4.0. Its changes include environment resources, task repairs and removals, requiring fresh runs for the major version. It reports cost and tokens alongside resolution rate. Do not present a 3.0 result as a 4.0 result or say that public benchmarks never include cost. 4.0 release.
A benchmark audit needs the right denominator
OpenAI's February 2026 SWE-bench Verified audit examined 138 selected difficult cases, not a representative random sample of every task. It found material issues in 59.4% of that audited subset. That does not mean “60% of all 500 tasks are broken.” The same report describes contamination evidence on particular tasks. Audit and selection method.
The subsequent July 2026 SWE-bench Pro audit also reports substantial task-quality issues. Therefore “switch to Pro and validity is solved” is not a defensible conclusion. Inspect the current task release, contested cases and maintainer fixes; reproduce important comparisons on your own audited tickets.
Agents, tools, computers and knowledge work
| Family | What it measures | Follow-up question |
|---|---|---|
| BFCL | Function/tool selection and arguments, with multi-turn/agentic categories | Which version, category and execution/AST checks? |
| τ-bench family | Policy-constrained interaction with tools and simulated users | Which domain, simulator, reset protocol and repeated-trial metric? |
| GAIA | General assistance across reasoning, files and tools | Is the submission one model or a complete multi-model agent? |
| OSWorld family | Tasks in real computer environments, with execution checks | Which OS/task version, interface and action/time budget? |
| WebArena / Online-Mind2Web | Website interaction under their respective environments | Controlled site or live web; what changed between runs? |
| GDPval | Deliverables for selected economically relevant occupations | Expert rubric, task selection and tools; not all work in an occupation |
| METR time horizons | Task difficulty expressed using human completion-time estimates | At what success probability, on which task distribution? |
BFCL V4 provides versioned tool evaluation. The τ-bench repository now describes τ³-bench, including voice and knowledge tasks; its July 2026 grading change affects banking-knowledge comparisons. The repository name still contains tau2-bench, so the URL alone does not identify the evaluated version. Its earlier τ² paper is from 2025, not a Q1 2026 debut.
The OSWorld project points to OSWorld 2.0, released June 26, 2026. Historical OSWorld-Verified scores remain tied to that older protocol. Check current primary documentation rather than using a historical human/model comparison as today's ceiling.
METR's time-horizon explanation is especially important: a 50%-success horizon is based on estimated human task duration for tasks in its suite. It is not the time the agent runs unattended, not a guarantee it can do every task of that duration, and not evidence of 50% reliability in every occupation.
pass@k and pass^k answer different questions
- pass@k: at least one of k attempts succeeds under the stated protocol.
- pass^k: all k repeated attempts succeed under the stated protocol.
For independent identical success probability p = 0.8 and k = 4, these are 1 − (1 − p)^4 = 99.84% and p^4 = 40.96%. Real tasks have different probabilities and attempts may correlate. Compute benchmark estimators from repeated per-task outcomes; do not exponentiate an aggregate observed success rate and call it the measured result. A production user may have only one attempt and no oracle to choose the correct candidate. See agent evaluation.
Long context, multimodal work and factuality
| Family | Test | Interpretation limit |
|---|---|---|
| Needle-in-a-haystack | Find inserted information in long input | One retrievable fact is not complete document understanding |
| RULER | Synthetic retrieval, tracing, aggregation and QA at controlled lengths | Useful stress tests; do not turn historical model results into a universal context fraction |
| MRCR | Resolve requests about repeated/near-duplicate conversational content | Needle count, distractors and length change difficulty |
| Fiction.liveBench | Narrative comprehension over long stories | Story/task selection and sample size limit generalization |
| LongBench v2 / LongBench Pro | Long-context tasks under separate datasets/protocols | They are distinct; check units, language and grading |
| MMMU / MMMU-Pro | Multidisciplinary multimodal understanding | Text-only shortcuts, image resolution and answer options matter |
| MathVista / DocVQA / ChartQA / MMBench | Visual mathematics, documents, charts or broad multimodal tasks | OCR, tolerance rules and task mix influence the score |
| Video-MME / Video-MME-v2 | Video understanding and temporal/multimodal reasoning | Frame sampling, audio/subtitles and aggregation rules matter |
| IFEval | Programmatically checkable instructions | Does not cover every semantic instruction |
| SimpleQA / SimpleQA Verified | Short-form factuality under a defined knowledge protocol | Distinguish correct, incorrect and abstaining outputs |
| TruthfulQA | Responses to questions inviting common misconceptions | Known test content and label assumptions need scrutiny |
| FACTS Grounding | Long-form responses supported by given source material | Grounding is not independent proof that the source is true |
RULER is a useful reminder to measure performance by length and task. LongBench v2 describes lengths in words, while LongBench Pro reports its own token-based range and bilingual tasks. Words and tokens are not interchangeable. A maximum accepted context length is a capacity limit, not a guarantee of uniform answer quality.
Primary descriptions: MMMU-Pro, Video-MME-v2, SimpleQA Verified, FACTS Grounding. Vary length, evidence position, distractors and required reasoning when translating these ideas into a product test.
Preference scores and composite indices
The Text Arena's human-preference board estimates relative preference using a Bradley–Terry model, with uncertainty; calling every current score “Elo” obscures the methodology. Arena-Rank publishes the ranking approach. Style control adjusts for selected formatting/length features; it does not isolate perfect objective truth.
Arena now also documents an optional factuality signal that combines preference with automated claim assessment. State which board/toggle you mean. Ordinary preference votes alone are not factual correctness; a factuality-augmented score also inherits the verifier's limits. New agent-specific boards can use different methods, so inspect their methodology separately.
MT-Bench uses model grading of a compact multi-turn test. Arena-Hard is an automated preference-oriented evaluation, not identical to live human Arena votes. Do not transfer a proxy's measured correlation or bias estimate to every later model and product.
The Leaderboard Illusion raises concerns about selective variant disclosure and other ranking effects. Arena's response disputes several claims and their magnitude. The practical response is to inspect testing/disclosure policy, fresh votes, uncertainty and the actual workload, rather than declaring either all leaderboards fraudulent or every rank authoritative.
| Aggregate | Construction | How to read it |
|---|---|---|
| Artificial Analysis Intelligence Index | Versioned combination of multiple evaluations | Record the version and weights; inspect weak constituent domains |
| Epoch Capabilities Index | Statistical model combining benchmark results of differing difficulty | Read fitted assumptions, data coverage and uncertainty |
| Holistic Agent Leaderboard (HAL) | Standardized agent evaluation infrastructure and multidimensional comparison | Compare models, scaffolds, benchmarks and cost; not a universal one-number truth |
Sources: Artificial Analysis methodology, Epoch methodology explanation, HAL paper. Composite weights encode priorities. They can improve broad comparison but cannot create evidence for a task absent from their components.
Four threats to a useful comparison
| Threat | Meaning | What helps—and what remains |
|---|---|---|
| Saturation | Scores approach the test's useful ceiling for the compared systems | Harder cases or finer analysis; old tests can still expose regressions or weaker models |
| Contamination | Evaluation content or close solutions influenced training/development | Time splits, private cases and provenance audits; none proves perfect cleanliness |
| Harness/measurement error | Resource, task, implementation or grader artifacts change scores | Pinned environments, validity audits and repeated runs |
| Selection/overfitting | Many variants are optimized and only winners reported | Protected holdout, disclosed search budget and fresh evaluation |
A public benchmark is not automatically known to be in a particular model's training data. A claimed cutoff or absence of detected overlap does not establish the opposite either. Canary strings, overlap searches and membership tests have limits; a negative finding from a small audit may be underpowered.
Private cases reduce direct exposure but make independent reproduction harder. Live cases need stable snapshots and maintained graders. More inference effort can help, waste resources or hurt a workflow; measure it rather than assuming monotonic improvement.
Worked comparison: shortlist, test, then decide
Two illustrative systems each attempt 500 private repository tasks once under the same test contract:
| Result | System A | System B |
|---|---|---|
| Accepted patches | 410 | 425 |
| Acceptance rate | 82% | 85% |
| Mean cost per attempted task | $0.20 | $0.50 |
| Total run cost | $100 | $250 |
| Cost per accepted patch | $0.244 | $0.588 |
The paired outcomes are 400 both-pass, ten A-only, 25 B-only and 65 both-fail. B has 15 net additional successes, not 425 wholly different successes. Its additional $150 costs $10 per net additional accepted patch in this run. That may be worthwhile or not, depending on value, human review and severity. These point estimates do not establish generalization or statistical certainty.
For one task in a 500-case test, the score changes by 0.2 percentage points. That is score granularity, not a confidence interval. Analyze paired uncertainty, language/repository slices and unsafe changes. Two separate confidence intervals overlapping does not by itself settle a paired significance question. Repeated runs add evidence about stochastic stability; they do not erase dataset bias.
Read diagram source
flowchart LR
P[Product tasks and constraints] --> B[Relevant public evidence]
B --> S[Small candidate shortlist]
S --> E[Matched private evaluation]
E --> A[Failures, uncertainty, latency and total cost]
A --> G{Meets release criteria}
G -->|No| R[Repair or choose another candidate]
G -->|Yes| C[Bounded canary]
C --> M[Verify real task outcomes]
Interview questions and answer checks
| Question | Strong answer |
|---|---|
| A vendor claims 90% on a coding benchmark. What is missing? | Dataset/split, model and scaffold, tools, budget, grading, repeats, uncertainty and provenance |
| Are different harnesses always incomparable? | They cannot isolate model quality without controlling differences; complete-stack comparisons can answer a product question |
| Is every three-point difference noise? | No. Case count, paired outcomes, dependence and measurement protocol determine uncertainty |
| Does an 85% target on a competition page mean a model achieved it? | No; distinguish qualification/prize thresholds from verified results |
| Is a new/private benchmark automatically clean? | No; review exposure, task validity, graders and access history |
| Why test function generation when selecting a repository agent? | It diagnoses one capability, but repository and workflow cases are also required |
| What does a high pass@8 hide? | The user may not have eight attempts or a correct selector; inspect pass@1 and consistency |
| Does a two-hour METR horizon mean two autonomous work hours? | No; the time describes human-estimated task difficulty under the evaluated distribution |
| Does a preference leaderboard measure truth? | Ordinary votes measure preference; any factuality extension has its own verifier and methodology |
| What should close a model-selection answer? | A scoped candidate decision, own-data evidence, budget, unresolved risks and a monitored rollout |
Final notes
Remember task → system → budget → grader → uncertainty → product fit. A good interview answer does not memorize a winner. It uses public benchmarks for relevant evidence, audits the comparison, and justifies a system with measured outcomes under the actual constraints.