What is an agent harness? The model is not the whole agent

Glossary entry 3. This one covers the harness. Entry 1 left a line hanging: the same benchmark scores 62.7% under ARC Prize's neutral (Standard) harness. Why would the same model on the same benchmark produce two different numbers? That is what a harness explains. We start from the definition and end with something you can use the next time you read a benchmark score.

“Agent = Model + Harness”

Source: DeepSeek Harness official page (checked 2026-09-19). DeepSeek put this line at the top of the page for its agent platform, DeepSeek Harness. Read it literally: an agent is a model plus a harness. A model on its own is not an agent.

● Last updated:  |  Primary sources: DeepSeek Harness · ARC Prize (GPT-6 Astra on ARC-AGI-3)  |  Checked 2026-09-19

Short version: a harness is everything that is not the model

A harness is the whole set of plumbing you put around a model to make it actually work. The system prompt, which tools it can call, how many attempts it gets, what sandbox it runs in, how earlier turns are carried forward, and how the result is scored. None of that touches the model's weights, and all of it moves the score.

One thing this is not: a pricing story. A harness does not change the per-token rate your vendor charges — it changes what you get for that rate (score, latency, token consumption). In practice, teams talk about swapping models or rewriting prompts. Very often the thing moving the numbers is the design around the model. That is why “harness” is used as a term in its own right: without naming the harness, two people comparing benchmark scores are usually not comparing the same thing at all.

The mental model: the model is the brain, the harness is the body and the kit

DeepSeek's own page puts the relationship in two short sentences.

"The model is the soul of an agent." / "A harness lets an agent understand its environment, use tools, and keep working in real-world settings." The model is the soul of the agent; the harness is what lets that agent understand its environment, use tools, and keep working in real-world settings.
Model
The brain
The reasoning capability itself. Its weights come from training and are not yours to edit
Harness
Body, tools, plan
What it sees, what it can use, how many tries it gets, how results are counted. All designable
So numbers move
Same brain
A different course and a different referee produce a different record

What is inside a harness (from the vendor's own list)

“Plumbing” is too vague to be useful, so here is the parts list. DeepSeek Harness says every agent capability is implemented as a plugin, and names them.

"Plugins provide every agent capability, including models, tools, skills, sessions, sandboxes, storage, loops, scheduling, and the UI." Plugins supply every agent capability — models, tools, skills, sessions, sandboxes, storage, loops, scheduling and the UI.
Harness componentWhat it decidesWhy the number moves
System promptWhat the model is shown, and what role and constraints it getsDifferent instructions, same model, different behaviour
ToolsWhich external capabilities are available (file editing, shell, search, APIs)More tools can solve more tasks; fewer tools change what is being measured
Skills and proceduresWhether a workflow or checklist is handed over up frontA given procedure removes exploration overhead
Sessions and memoryHow prior turns and intermediate results are carried forwardIf carry-over is assumed, long tasks lose less accuracy
SandboxWhere it executes (isolation, permissions, network)A different environment changes which actions are even possible
Loops and retriesHow many times it may reconsider or retryMore attempts raise the success rate; the cap is part of the result
SchedulingHow parallel work and subagents are orderedSequencing changes time and step count
LoggingHow inputs, reasoning and tool calls are recordedWhat you cannot replay, you cannot improve or verify
Scoring and evaluationHow results are counted (pass/fail, partial credit, cut-offs)Change the counting and the same run scores differently

Everything except the quoted list is our plain-language expansion of DeepSeek Harness's own items (models, tools, skills, sessions, sandboxes, storage, loops, scheduling, UI). “Scoring and evaluation” is our addition as a benchmark-side design element; it is not a DeepSeek Harness item.

Why the numbers move: ARC-AGI-3, with the conditions attached

Abstract advice is hard to act on, so here is one measured case. Same model (GPT-6 Astra), same benchmark (ARC-AGI-3), two published results under different harnesses. The model setting also differs — Astra (max) versus Astra (high) — so the gap between 62.7% and 99.9% cannot be attributed to the harness alone.

Harness (measurement condition)Model settingScoreCostWhat it does
Standard harness
(ARC Prize's neutral harness = the same minimal interface for every model)
Astra (max)62.7%$26KThe model carries forward notes it chooses to keep; the same minimal interface for every model
Provider Adapter harness
(OpenAI's)
Astra (high)99.9%$19KPreserves opaque reasoning state between requests, and uses compaction for longer conversations so the model can reuse prior work
Standard harness, verbatim: "Standard harness enables a model to carry forward notes it chooses to keep with it throughout the environment." Provider Adapter harness, verbatim: "The Provider Adapter harness preserves opaque reasoning state between requests and uses compaction for longer conversations, allowing the model to reuse prior work." Both quoted from the ARC Prize blog (checked 2026-09-19).
The difference is what can be carried forward. Under the neutral harness the model carries only notes it chose itself. Under the Provider Adapter it preserves reasoning state across requests and compacts long conversations for reuse. Same weights, different memory. That is the main condition separating 62.7% from 99.9% — though the model setting differs too (max versus high), so not all of the gap is the harness.
  • Both are state of the art under their own conditions. ARC Prize treats both as SOTA. This is not a case of one number being inflated and the other being correct.
  • A like-for-like comparison exists too. ARC Prize also published 7.8% for GPT-5.6 Sol on the same Standard harness (checked 2026-09-15). That is roughly an eightfold step on one playing field. Putting 99.9% next to 7.8% compares two different harnesses.
  • The benchmark operator itself has committed to labelling. See the quote below.
"Going forward, we will report both Standard harness and Provider Adapter harness results on the ARC-AGI leaderboard, with each evaluation condition clearly labeled." Going forward, both harness results will be reported on the ARC-AGI leaderboard with each evaluation condition clearly labelled (ARC Prize blog, checked 2026-09-19). Our reading is that a score with no harness name makes the comparison conditions hard to judge, so it is worth checking which configuration a number came from.

Sources: ARC Prize, “OpenAI's GPT-6 Astra on ARC-AGI-3” (checked 2026-09-19) and the ARC-AGI results page. Scores and costs are ARC Prize's published figures, not our own measurement. Background on the same harness architecture is in entry 1.

What does that look like in practice? DeepSeek Harness

“Harness” is an abstract term, so a real one helps. DeepSeek ships DeepSeek Harness in developer preview, with the source code released at the same time. Here is only what its official page states.

What the page saysWhat it means
Everything is a pluginModels, tools, skills, sessions, sandboxes, storage, loops, scheduling and the UI are all swappable, recomposable parts
Cordis kernelHandles mounting, unmounting and dependencies of plugins; the agent's capabilities live in the plugins
Compose with configurationSelect, swap or extend any capability in configuration, without changing the harness source code
Every run is traceableEverything the model sees is recorded in an append-only session log: system prompts, reasoning, tool calls and results, subagent scheduling, context injections. Resume, fork, search and replay all work on that same event stream

Four runtime modes — and one of them is explicitly for benchmarking

The page lists four runtime modes. All four run the same kind of model; what differs is the kit handed to it. That is a harness difference, shipped as a feature.

Standard mode
Full kit
File editing, shell, file and web search, skills, planning, goals, subagents and workflows
Code mode
Orchestrate in code
Standard capabilities exposed through a Code Mode SDK, so the model can combine multi-step operations in one TypeScript program
Minimal mode
Two tools only
A persistent bash tool and a file editor. The page states this is for benchmarking models in a minimal environment
Creator mode
Build a harness
Standard capabilities plus runtime inspection, plugin experiments and preset-authoring guidance
The interesting part is Minimal mode. DeepSeek describes it as a two-tool agent for benchmarking models in a minimal environment. In other words, a stripped-down harness exists as a first-class feature precisely because the harness changes what you measure. That is not a criticism of anyone's benchmark — it is a design assumption the vendor builds on.

Source: DeepSeek Harness official page (checked 2026-09-19; developer preview). Quoted phrases are the page's own words; the rest is our summary. Repository: github.com/deepseek-ai/deepseek-harness. We are not affiliated with DeepSeek and this is not a product recommendation.

The picture: the harness is what sits outside the model

A harness sits outside the model: system prompt, tools, skills, memory, sandbox and retries surround the model at the centre. Same model, different surroundings, different result (ARC-AGI-3: 62.7% under the Standard harness at model setting max, 99.9% under the Provider Adapter harness at setting high) Harness = everything outside the model Same model (the brain), different surroundings, different result HARNESS (outside the model; this is what you design) System prompt Tools and search Skills, procedures Model (brain) unchanged Sessions, memory Sandbox Loops, retries Scoring and evaluation — how results are counted — belong here too. Example: ARC-AGI-3 — Standard 62.7% (max) / Provider Adapter 99.9% (high), both published by ARC Prize
The model in the middle is fixed; the ring around it is designable. Components follow the list on the DeepSeek Harness page. Scores are ARC Prize figures; the Standard harness was run at model setting max and the Provider Adapter at high.

Everyday version: the same player, a different score

The same player scores differently when the court, the referee and the rulebook change. The player's ability (the model) does not change. But the size of the court, the equipment allowed, what the referee calls a foul, how many substitutions are permitted — change those and the same player's score is a different number. The harness is that whole set: court, referee, rules and equipment. Just as you can raise a score without replacing the player, you can move the numbers without swapping the model.
In sportIn an AI agent
The player's abilityThe model (trained weights; not yours to edit)
Court size, the ballSandbox, available tools
The referee and the rulebookScoring and evaluation (what counts as correct)
Substitutions and timeoutsRetries and loop limits
The coach's game planSystem prompt, skills, procedures
Carrying last game's notes inSessions and memory — the thing that separated 62.7% from 99.9% on ARC-AGI-3

One takeaway from the analogy: when a score is discussed, separate the player's story from the court-and-referee story. That is most of what reading a benchmark requires.

How to use this at work

Three checks when you see a score

  • Whose harness? The vendor's own, or a neutral one run by the benchmark operator. On ARC-AGI-3 that was 62.7% versus 99.9%.
  • Is that harness close to your setup? A score that assumes reasoning state can be carried across requests will not reproduce in a setup that cannot carry it. And vice versa.
  • What are the scoring and the limits? Retry caps, cut-offs, partial credit. Numbers measured under different rules do not make a ranking.
  • When your agent is not good enough, suspect the harness first. Tightening the toolset, fixing how memory is carried, or revisiting retry limits can beat swapping the model. ARC-AGI-3 is the same model under two harnesses producing two very different results.
  • An improvement you cannot replay is not an improvement. DeepSeek makes “everything the model sees is recorded in an append-only session log” a headline feature precisely so harness changes can be verified. If you build your own, do not defer logging.
  • This is separate from pricing. A harness affects performance — score, latency, token consumption — not the per-token rate itself. The pricing tables cover rates; this entry is about what you get for them.

In one sentence

A harness is everything that is not the model

  • The model is the brain. The harness is the body, the tools, the plan and the referee that put it to work.
  • It includes: system prompt, tools, skills, sessions and memory, sandbox, loops and retries, scheduling, logging, scoring and evaluation.
  • That is why the same model on the same benchmark produces different numbers. On ARC-AGI-3, ARC Prize published 62.7% under the Standard harness (Astra max) and 99.9% under the Provider Adapter harness (Astra high). The settings differ as well, so not all of that gap is the harness.
  • When you read a benchmark, check which harness before you check which model.

The term is useful because it moves the discussion from “is this a good model” back to “how is my setup designed”. Same model, different harness, different result. Next time you see a score, look for the harness name first.

Next

The worked example that introduces harnesses is entry 1 (how to read benchmark numbers); cache pricing is entry 2 (prompt caching).

🧪 All glossary entries →

Glossary ・ How to read benchmark numbers ・ Prompt caching ・ Pricing ・ Release timeline

⚠️ Disclaimer

  • Quotes and figures were checked against primary sources (DeepSeek's official page, the ARC Prize blog) on 2026-09-19. ARC-AGI-3 scores and costs are ARC Prize's published figures, not our measurements.
  • Vendor statements are quoted verbatim; the surrounding explanation is ours. Where the two differ, the vendor's wording governs.
  • DeepSeek Harness is in developer preview; features and APIs will change. Check the official page and repository before adopting it.
  • We are not affiliated with any provider, and this is not a recommendation of any product.