Who operates the benchmarks — a four-step way to tell
Glossary entry #6. Entry #5 gave you a tool for sorting benchmarks by (1) who operates them, (2) what they measure and (3) how they measure it. This entry digs into part one. Readers told us the operator's name is often missing from the score table, and sometimes from the benchmark's own site. So here is a fixed four-step procedure for getting from a benchmark name to the organisation behind it, plus a table mapping twelve benchmarks we quote on this site.
What defines a benchmark is not the people who publish scores — it is the people who distribute the test. Same test, different distributor, different benchmark. On this page we call the party that authors, distributes and versions the test or evaluation framework the operator; the dataset author, harness owner and score reporter are sometimes different parties.
The procedure — from benchmark name to operator in four steps
- Find the paper or the official repository. Skip the summary articles. Open the original paper (arXiv and similar) or the project repository (GitHub). The repository owner is your first clue to the operator.
Example: DeepSWE → github.com/datacurve-ai/deep-swe → ownerdatacurve-ai= Datacurve - Work out what kind of organisation it is. A model vendor (OpenAI, Anthropic), an independent evaluator (Artificial Analysis, ARC Prize), a research lab (a university or non-profit), or a business-software company (Zapier, Cursor). This tells you how far the operator sits from the models being judged.
- Check who reported the score. The same benchmark can be published by the operator, by a vendor running its own setup, or quoted by a third party. If a vendor's page carries a note like “as reported by OpenAI”, that figure is the vendor's own measurement.
- Check whose harness ran it. A model-vendor harness and a third-party harness are different conditions — see entry #3 on harnesses. Only after this step can you judge whether two numbers are comparable.
💡 Shortcuts that save the five minutes
- Repository owner names (
datacurve-ai,harbor-framework,xlang) are usually abbreviations of the organisation. Searching the owner name lands on the official site. - If the organisation name is inside the benchmark name, you are done: GDPval-AA (AA = Artificial Analysis), AutomationBench (published by Zapier on its own blog).
- The reverse also holds: names with no organisation in them are the ones that need step one.
Figure: who distributes, who runs, who reports
Worked examples — benchmark and operator
Focused on benchmarks that actually appeared in the September 2026 model releases (Claude Opus 5.5, GPT-6 Sol and GPT-6 Luna) and on the ones this site quotes elsewhere.
| Benchmark | Operator | Category | Primary source |
|---|---|---|---|
| DeepSWE / DeepSWE 1.1 | Datacurve | Independent evaluator | Official repository (owner: datacurve-ai) |
| Terminal-Bench 4.0 | Harbor + Laude Institute | Research | Official repository / paper |
| Terminal-Bench-Science 0.1 | Stanford + Laude Institute | Research | Official repository |
| OSWorld / OSWorld 2.0 | XLANG Lab | Research | Official site / paper |
| GDPval | OpenAI | Model vendor | Official page / paper |
| GDPval-AA v2.1 | Artificial Analysis | Independent evaluator | Official page (dataset shared with GDPval, operator is not) |
| AutomationBench 1.0.6 | Zapier | Business-software company | Official page; OpenAI's comparison table also states Zapier ran it |
| FrontierCode 1.1 | Cognition | Developer-tools company | Official page (the link OpenAI cites) |
| CursorBench 4.0 | Cursor | Developer-tools company | Appears in both Anthropic's and OpenAI's comparison tables. The operator's own explainer page is unconfirmed. |
| Agents' Last Exam V1 | agents-last-exam.org (official domain; operating entity not confirmed) | Independent evaluator | Official site |
| ARC-AGI-3 | ARC Prize | Independent evaluator (non-profit) | arcprize.org — the neutral-harness figure we discuss in entry #1 |
| WANDR | Perplexity | AI company | Anthropic notes it ran this under modified conditions and that “scores are not directly comparable” to Perplexity's published setup |
Sources: each benchmark's official site or repository, plus the notes in Anthropic's “Introducing Claude Opus 5.5” and OpenAI's “Introducing GPT-6 Sol and Luna” (verified 2026-09-25). CursorBench is flagged unconfirmed because we could not reach the operator's own explainer.
Three ways the operator changes the number
Case 1: same dataset, different operator, different number
GDPval (OpenAI) vs GDPval-AA (Artificial Analysis)
- What differs
- The dataset is shared, but the operator is not — scoring design and how models are run differ. Anthropic's Opus 5.5 announcement uses GDPval-AA v2.1, reporting Elo scores of 1,846 (Opus 5.5), 1,735 (Fable 5.1) and 1,708 (Opus 5).
- How to read it
- A number you find by searching “GDPval” and a number labelled “GDPval-AA” in an announcement are different numbers. Search with the suffix (-AA, -Verified, v2.1) included.
Case 2: a third-party operator, but vendor-side conditions
AutomationBench (Zapier)
- What differs
- Zapier both operates and runs it. But OpenAI's table specifies that the runs were performed without fallback models, so safeguard interventions counted as failures — a condition that pushes scores down. OpenAI also notes that the Claude Fable 5.1 datapoint understates its real cost, since it omits Opus 5 fallback spend that occurred on ~40% of tasks.
- How to read it
- “A third party operates it” does not by itself mean “fair”. You still need to know who ran it and under what constraints.
Case 3: someone else's benchmark, your own conditions
WANDR (Perplexity)
- What differs
- Anthropic ran Perplexity's benchmark under modified conditions (offline web search and web fetch, programmatic tool calling, code execution, a 980k-token task budget) and states plainly that “this differs from Perplexity's published setup, scores are not directly comparable” — then shows only models measured under identical conditions.
- How to read it
- This is an example of a good caveat. A table that says “not comparable” is more trustworthy than one that is silent. Where conditions are unstated, assume they may not be comparable.
Checklist — when you meet a number in an announcement
- Is an organisation name inside the benchmark name? If not, go back to the paper or repository to find the operator.
- Is the operator a model vendor or a third party? A benchmark owned by a vendor cannot easily measure rivals under identical conditions.
- Is there an “as reported by” note? If so, the figure is that vendor's own measurement, not a re-run by the publisher.
- Whose harness is it? A vendor harness (e.g. Claude Code) can favour that vendor.
- Are the compared generations consistent? OpenAI states outright that it reported Claude Fable 5 scores where Fable 5.1 scores were unavailable. Mixed generations cannot be lined up.
- Does it say “not comparable”? An explicit caveat is a plus, not a minus.
Sources (all verified 2026-09-25)
- Anthropic, “Introducing Claude Opus 5.5” — anthropic.com/claude-opus-5-5 (benchmark footnotes 1–4)
- OpenAI, “Introducing GPT-6 Sol and Luna” — openai.com/index/introducing-gpt-6-sol-and-luna
- DeepSWE — github.com/datacurve-ai/deep-swe (Datacurve)
- Terminal-Bench — github.com/harbor-framework/terminal-bench (Harbor + Laude Institute)
- OSWorld — osworld-v2.xlang.ai (XLANG Lab)
- GDPval — openai.com/index/gdpval (OpenAI) / GDPval-AA — artificialanalysis.ai/evaluations/gdpval-aa (Artificial Analysis)
- AutomationBench — zapier.com/benchmarks (Zapier)
- FrontierCode — cognition.com/frontiercode (Cognition)
- Agents' Last Exam — agents-last-exam.org
- ARC-AGI-3 — arcprize.org (ARC Prize)
Related: all glossary entries · What is a benchmark? (#5) · Harnesses (#3) · Claude Opus 5.5 · GPT-6 Sol / Luna · Release timeline
⚠️ Limits of this entry
- Operators are as confirmed on official sites, papers and repositories as of September 25, 2026; benchmarks get transferred, renamed and re-versioned.
- The category labels (independent evaluator, research, business-software company) are our own grouping, not an industry standard.
- This page does not cover the scores or measurement conditions themselves — see each model's page for numbers.
- This site is not affiliated with any provider.