Who operates the benchmarks — a four-step way to tell

Glossary entry #6. Entry #5 gave you a tool for sorting benchmarks by (1) who operates them, (2) what they measure and (3) how they measure it. This entry digs into part one. Readers told us the operator's name is often missing from the score table, and sometimes from the benchmark's own site. So here is a fixed four-step procedure for getting from a benchmark name to the organisation behind it, plus a table mapping twelve benchmarks we quote on this site.

What defines a benchmark is not the people who publish scores — it is the people who distribute the test. Same test, different distributor, different benchmark. On this page we call the party that authors, distributes and versions the test or evaluation framework the operator; the dataset author, harness owner and score reporter are sometimes different parties.

Concretely: GDPval and GDPval-AA share a dataset but are not the same number, because OpenAI runs one and Artificial Analysis runs the other. The four steps below will get you from an unfamiliar benchmark name to its operator in about five minutes.

● Last updated:  |  Policy: operators are listed only where confirmed on an official site, paper or repository. Anything unconfirmed is flagged as such

The procedure — from benchmark name to operator in four steps

  1. Find the paper or the official repository. Skip the summary articles. Open the original paper (arXiv and similar) or the project repository (GitHub). The repository owner is your first clue to the operator.
    Example: DeepSWE → github.com/datacurve-ai/deep-swe → owner datacurve-ai = Datacurve
  2. Work out what kind of organisation it is. A model vendor (OpenAI, Anthropic), an independent evaluator (Artificial Analysis, ARC Prize), a research lab (a university or non-profit), or a business-software company (Zapier, Cursor). This tells you how far the operator sits from the models being judged.
  3. Check who reported the score. The same benchmark can be published by the operator, by a vendor running its own setup, or quoted by a third party. If a vendor's page carries a note like “as reported by OpenAI”, that figure is the vendor's own measurement.
  4. Check whose harness ran it. A model-vendor harness and a third-party harness are different conditions — see entry #3 on harnesses. Only after this step can you judge whether two numbers are comparable.

💡 Shortcuts that save the five minutes

  • Repository owner names (datacurve-ai, harbor-framework, xlang) are usually abbreviations of the organisation. Searching the owner name lands on the official site.
  • If the organisation name is inside the benchmark name, you are done: GDPval-AA (AA = Artificial Analysis), AutomationBench (published by Zapier on its own blog).
  • The reverse also holds: names with no organisation in them are the ones that need step one.

Figure: who distributes, who runs, who reports

The three roles in a benchmark Diagram showing three roles connected by arrows: the operator who writes and distributes the test, the harness owner who provides the execution environment, and the reporter who publishes the numbers. 1. Operator Writes, distributes and versions the test Datacurve / XLANG Lab Zapier / ARC Prize 2. Harness owner Provides the tooling and environment that runs it Claude Code / vendor setup often vendor-owned 3. Reporter Publishes the resulting numbers change this and you cannot compare “Same benchmark, so comparable” only holds when all three line up AutomationBench: 1 Zapier, 2 Zapier, 3 Zapier — but it also appears in OpenAI's table, under OpenAI's conditions
Created by LLM Data Hub. This is a way of thinking about the roles, not a statement of any benchmark's measured conditions.

Worked examples — benchmark and operator

Focused on benchmarks that actually appeared in the September 2026 model releases (Claude Opus 5.5, GPT-6 Sol and GPT-6 Luna) and on the ones this site quotes elsewhere.

BenchmarkOperatorCategoryPrimary source
DeepSWE / DeepSWE 1.1DatacurveIndependent evaluatorOfficial repository (owner: datacurve-ai)
Terminal-Bench 4.0Harbor + Laude InstituteResearchOfficial repository / paper
Terminal-Bench-Science 0.1Stanford + Laude InstituteResearchOfficial repository
OSWorld / OSWorld 2.0XLANG LabResearchOfficial site / paper
GDPvalOpenAIModel vendorOfficial page / paper
GDPval-AA v2.1Artificial AnalysisIndependent evaluatorOfficial page (dataset shared with GDPval, operator is not)
AutomationBench 1.0.6ZapierBusiness-software companyOfficial page; OpenAI's comparison table also states Zapier ran it
FrontierCode 1.1CognitionDeveloper-tools companyOfficial page (the link OpenAI cites)
CursorBench 4.0CursorDeveloper-tools companyAppears in both Anthropic's and OpenAI's comparison tables. The operator's own explainer page is unconfirmed.
Agents' Last Exam V1agents-last-exam.org (official domain; operating entity not confirmed)Independent evaluatorOfficial site
ARC-AGI-3ARC PrizeIndependent evaluator (non-profit)arcprize.org — the neutral-harness figure we discuss in entry #1
WANDRPerplexityAI companyAnthropic notes it ran this under modified conditions and that “scores are not directly comparable” to Perplexity's published setup

Sources: each benchmark's official site or repository, plus the notes in Anthropic's “Introducing Claude Opus 5.5” and OpenAI's “Introducing GPT-6 Sol and Luna” (verified 2026-09-25). CursorBench is flagged unconfirmed because we could not reach the operator's own explainer.

Three ways the operator changes the number

Case 1: same dataset, different operator, different number

GDPval (OpenAI) vs GDPval-AA (Artificial Analysis)

What differs
The dataset is shared, but the operator is not — scoring design and how models are run differ. Anthropic's Opus 5.5 announcement uses GDPval-AA v2.1, reporting Elo scores of 1,846 (Opus 5.5), 1,735 (Fable 5.1) and 1,708 (Opus 5).
How to read it
A number you find by searching “GDPval” and a number labelled “GDPval-AA” in an announcement are different numbers. Search with the suffix (-AA, -Verified, v2.1) included.

Case 2: a third-party operator, but vendor-side conditions

AutomationBench (Zapier)

What differs
Zapier both operates and runs it. But OpenAI's table specifies that the runs were performed without fallback models, so safeguard interventions counted as failures — a condition that pushes scores down. OpenAI also notes that the Claude Fable 5.1 datapoint understates its real cost, since it omits Opus 5 fallback spend that occurred on ~40% of tasks.
How to read it
“A third party operates it” does not by itself mean “fair”. You still need to know who ran it and under what constraints.

Case 3: someone else's benchmark, your own conditions

WANDR (Perplexity)

What differs
Anthropic ran Perplexity's benchmark under modified conditions (offline web search and web fetch, programmatic tool calling, code execution, a 980k-token task budget) and states plainly that “this differs from Perplexity's published setup, scores are not directly comparable” — then shows only models measured under identical conditions.
How to read it
This is an example of a good caveat. A table that says “not comparable” is more trustworthy than one that is silent. Where conditions are unstated, assume they may not be comparable.

Checklist — when you meet a number in an announcement

  1. Is an organisation name inside the benchmark name? If not, go back to the paper or repository to find the operator.
  2. Is the operator a model vendor or a third party? A benchmark owned by a vendor cannot easily measure rivals under identical conditions.
  3. Is there an “as reported by” note? If so, the figure is that vendor's own measurement, not a re-run by the publisher.
  4. Whose harness is it? A vendor harness (e.g. Claude Code) can favour that vendor.
  5. Are the compared generations consistent? OpenAI states outright that it reported Claude Fable 5 scores where Fable 5.1 scores were unavailable. Mixed generations cannot be lined up.
  6. Does it say “not comparable”? An explicit caveat is a plus, not a minus.

Sources (all verified 2026-09-25)

Related: all glossary entries · What is a benchmark? (#5) · Harnesses (#3) · Claude Opus 5.5 · GPT-6 Sol / Luna · Release timeline

⚠️ Limits of this entry

  • Operators are as confirmed on official sites, papers and repositories as of September 25, 2026; benchmarks get transferred, renamed and re-versioned.
  • The category labels (independent evaluator, research, business-software company) are our own grouping, not an industry standard.
  • This page does not cover the scores or measurement conditions themselves — see each model's page for numbers.
  • This site is not affiliated with any provider.