What is a benchmark? Sorting them by who, what and how
Glossary, entry five. The term this time is benchmark.
It started from a reader's perfectly reasonable complaint: a review said
“in developers' own testing, across several benchmarks such as DeepSWE, GDPval-AA, Terminal-bench and OSWorld” — and
it was impossible to tell who measured what, exactly. There are too many benchmarks to make sense of.
So this page stops trying to memorise benchmark names. Instead it hands you three questions —
who runs it, what it measures, and how it measures it (harness and settings) —
that let you sort any new name you meet. It pairs with entry three, the agent harness.
“The same benchmark name does not mean comparable numbers.”
The line this page keeps coming back to. A benchmark's name is only the title of a test. Change who sets the questions, who grades them, what tools the model is handed, or which release you are on, and the score means something else. Concretely: GDPval and GDPval-AA share a dataset, but the operator and the harness differ, so they are not the same number. We walk through it below.
●
Last updated:
| Policy: only what we can confirm in the operator's own repository, site or paper. If it cannot be confirmed, it is not published here
Why does it feel impossible to keep up?
Because benchmarks are not built to compare models. They are built by whoever wants to measure a specific capability.
A company that wants to measure coding, a lab that wants to measure the economic value of real work, a university group that wants to
measure terminal operation — different goals produce different shapes of test. Yet in the news, all of them get flattened into the same
phrase: “the model's benchmark score.” That is where the information is lost.
So the useful skill is not a list of names. It is a procedure that pulls out whose number it is, what it measured, and how. Here is the whole shape of it.
← scroll sideways to see the whole figure →
The procedure for reading a benchmark score. The name is only the entrance; the meaning is set by the operator, the target and the method. The band at the bottom is not a caveat but a precondition — a score is only quotable when all four of operator, release, harness and settings travel with it.
The goal of this page: to reach the point where, on seeing “model A scores 73% on benchmark X,” you can work out
which operator's X, on which release, under which harness that 73% came from — so that your first move is identifying the
source, not deciding whether to believe it.
Six benchmarks, sorted by the three questions
Six that frequently appear side by side, in one format. All of it confirmed against the operator's own repository, site or paper (checked September 23, 2026).
Benchmark
1. Who (operator)
2. What (target)
3. How (harness and grading)
DeepSWE
Datacurve (GitHub org: datacurve-ai)
Software engineering: long-horizon coding work
113 tasks across five languages (TypeScript, Go, Python, JavaScript, Rust). Agents run in isolated environments and are graded by program-based verifiers. Tasks are written from scratch rather than adapted from merged pull requests, as contamination control
GDPval
OpenAI
Economically valuable real-world tasks (44 occupations, top 9 US GDP sectors)
1,320 tasks. Deliverables span documents, slides, diagrams and spreadsheets, shipped with reference files and context. Graded by human industry experts in blind pairwise comparisons
GDPval-AA
Artificial Analysis
GDPval's own dataset — re-run by a different operator with a different method
v2.1. Models get shell access and web browsing in an agentic loop via Stirrup, with Elo ratings derived from blind pairwise comparisons
Terminal-Bench
Harbor + Laude Institute
Working in a terminal / command line
Inside an isolated container the agent drives a real shell; after it stops, a verification suite runs against the final state. A task counts as resolved only if every verification test passes; the score is the fraction resolved. Continuous, with tagged releases
Terminal-Bench-Science
Stanford + Laude Institute (community-run)
Real research workflows across the life, physical, earth, mathematical and engineering sciences
70 expert-curated tasks. Each is graded by verifier tests written by that task's own author. An eight-hour agent phase per task, far beyond a conventional terminal benchmark
OSWorld
XLANG Lab (University of Hong Kong) and collaborators
GUI and OS operation — real office work on a real computer
369 tasks in a real computer environment on a VM. Every task carries a detailed initial-state setup and an execution-based evaluation script
Note: every figure and structure above comes from the operator's primary material, but task counts change with the release — Terminal-Bench 2.0 has 89 tasks; OSWorld has 369 in 1.0 and 108 in 2.0. Quote the release alongside the number.
Taking them one at a time
The same six, retraced through the operators' own material. We keep the verbatim quotes separate from our own interpretation.
DeepSWE — coding measurement with the answer key removed
Operator: Datacurve (GitHub org: datacurve-ai)
The official description, verbatim
"DeepSWE is a benchmark for measuring frontier coding agents on original, long-horizon software engineering tasks drawn from active open-source repositories. The benchmark includes 113 tasks across TypeScript, Go, Python, JavaScript, and Rust, with isolated environments and program-based verifiers."
— github.com/datacurve-ai/deep-swe (checked 2026-09-23)
What it measures
Whether an agent can carry a long-horizon engineering task to completion inside a real open-source repository — not a one-question lookup, but implementation and fixes spanning several files.
How it measures — the important part
Tasks are written from scratch rather than adapted from existing merged pull requests, which reduces the chance the model has already seen the answer during pretraining. Grading does not require matching one reference patch; it uses hand-written verifiers that check behaviour. Datacurve reports that these verifiers disagreed with human graders 1.4% of the time, against 32.4% for a grading approach built on matching hidden test suites (SWE-bench Pro, in the company's own comparison).
An easy name collision
“DeepSWE-Preview” is a different thing. That is a 32B open-weight coding-agent model released by Agentica and Together AI, not the benchmark. Searching for “DeepSWE” surfaces both, so the first thing to establish is whether the conversation is about the benchmark or the model.
GDPval — measuring work people actually get paid for
Operator: OpenAI
What it measures
Not exam-style questions but real deliverables. The paper describes coverage of the top nine US GDP sectors and 44 occupations, spanning the majority of US Bureau of Labor Statistics work activities. Deliverables span documents, slides, diagrams, spreadsheets and multimedia, and come with reference files and context rather than being plain text prompts.
How it measures
1,320 tasks, assessed on quality, speed and cost against human experts, with grading done by expert blind pairwise comparison. A gold subset of 220 tasks is open-sourced alongside a public automated grading service at evals.openai.com.
A caveat from OpenAI itself
The paper states plainly that GDPval is an early step that does not reflect the full nuance of many economic tasks, and that the current version is one-shot, so it cannot capture a model improving across multiple drafts.
The names look like the same thing. They are not: the operator and the method differ; only the dataset is shared. Artificial Analysis describes it as: "GDPval-AA v2.1 is Artificial Analysis' evaluation framework for OpenAI's GDPval dataset."
How it measures
The same page states: "Models are given shell access and web browsing capabilities in an agentic loop via Stirrup to solve tasks, with Elo ratings derived from blind pairwise comparisons." Where OpenAI uses human expert grading, this hands the model tools and ranks it on Elo. Same tasks, a different measurement instrument.
What follows
You cannot place a GDPval score and a GDPval-AA score side by side and read off a winner. Even a shared dataset does not survive a change of operator and harness. That is this page's central point, demonstrated.
Terminal-Bench — did the agent finish the job at the shell?
Operator: Harbor + Laude Institute (project lead: Ryan Marten)
What it measures
Whether an agent can complete work that lives entirely in a command line. The Terminal-Bench 2.0 paper describes "a carefully curated hard benchmark composed of 89 tasks in computer terminal environments inspired by problems from real workflows."
How it measures
Each task starts from a natural-language instruction inside an isolated container. The agent drives a real shell, reads the output and can retry as often as it likes. Once it stops, a verification suite runs against the final container state and a task counts as resolved only when every verification test passes; the score is the fraction resolved. That means the number reflects error recovery, reading documentation and adapting to unexpected output, not one lucky generation.
Mind the release
This is a continuous benchmark with tagged releases, so “x% on Terminal-Bench” is incomplete without the version (2.0, 2.1, 4.0 and so on).
Terminal-Bench-Science — a fork aimed squarely at researchers' work
Operator: Stanford + Laude Institute (community-run). Leads include Steven Dillmann, Sanmi Koyejo and Ludwig Schmidt
What it measures
Not general terminal operation but real scientific research workflows, across five domains: life, physical, earth, mathematical and engineering sciences. The repository states: "Terminal-Bench-Science currently contains 70 expert-curated tasks across five broad scientific domains and is growing toward 100+ tasks."
How it measures
Tasks are authored and reviewed by domain experts, and graded by verifier tests written by the task's own author. A third-party evaluation notes that where Terminal-Bench 2.1 allows 15–60 minutes of agent time per task, this one allows eight hours per task, with up to 4 vCPU and 16GB. Same family name, entirely different difficulty and time scale.
What follows
Use it to see that changing the test changes the ranking: strong general coding ability does not automatically translate into scientific workflows.
OSWorld — can the agent drive a real computer's GUI?
Operator: XLANG Lab (University of Hong Kong), with Salesforce Research, CMU and the University of Waterloo
What it measures
Not text or code but GUI operation on an actual computer. The project site states: "We also create a benchmark of 369 real-world computer tasks in OSWorld with reliable, reproducible setup and evaluation scripts."
How it measures
Tasks run in a real environment on a VM. Each carries a detailed initial-state setup and an execution-based evaluation script, so grading is mechanical, against the state the agent left behind. The paper's published reference points were that humans complete over 72.36% of tasks while the best model managed 12.24%, the failures concentrated in GUI grounding and operational knowledge. (Those are the paper's figures — a 2024 measurement, not today's state of the art.)
Mind the release
The project notes that eight Google Drive tasks may need manual configuration or can be excluded for network reasons (361 tasks). Separately, OSWorld-Verified fixed more than 300 issues, and OSWorld 2.0 is a distinct 108-task long-horizon benchmark across seven professional domains. 1.0's 369 tasks and 2.0's 108 tasks are not the same benchmark.
No memorisation required. Run these in order and the weight of the number changes.
Who operates it — the vendor that built the model, or a third party? A self-reported figure and an independent measurement are not the same claim even at an identical “73%.”
Which release — 1.0, 2.0, Verified, v2.1. Different releases mean different task populations. A score without a version cannot be placed on the same axis as anything else.
What it measures — code, real deliverables, terminal work, or GUI operation. A score says nothing about a capability the test never exercises.
How it was measured (harness and settings) — what the agent was given, how many attempts, what time budget, and whether grading checks behaviour or matches a reference answer. Change these and the same model on the same benchmark moves (see agent harness).
What the number counts — completion rate or Elo? Elo is a relative ranking and a completion rate is an absolute share; they cannot be added or subtracted.
💡 How we handle this on this site
When LLM Data Hub introduces a model, we also record whose harness produced the figure. We do not publish tables that merely stack up benchmark scores. Figures mirrored by aggregator sites, or relayed from search snippets, are not treated as fact. Entry one — how to read “ARC-AGI-3 99.9%, ExploitBench 100%” — is exactly this problem in the wild: 62.7% under a neutral harness, 99.9% under a provider adapter.
Related entries
🧰
What is an agent harness? Why the same model scores differently
Question three on this page, taken apart. A harness is everything that is not the model — system prompt, tools, memory, sandbox, retries, scoring — and it moves the score even when the model and benchmark are fixed. Read together, question three stops being abstract.
GPT-6 Astra: how to read "ARC-AGI-3 99.9%, ExploitBench 100%"
The same benchmark scores 62.7% under ARC Prize's neutral harness. A perfect score means the test can no longer separate models. A worked example of why unsaturated benchmarks carry the most information.
What is a 1M-token context window? What fits, and what it costs
Also about reading numbers. The three big vendors bill long input very differently, to the point where crossing a boundary by a hundred tokens can double the bill.