What is an agent harness? The model is not the whole agent
Glossary entry 3. This one covers the harness. Entry 1 left a line hanging: the same benchmark scores 62.7% under ARC Prize's neutral (Standard) harness. Why would the same model on the same benchmark produce two different numbers? That is what a harness explains. We start from the definition and end with something you can use the next time you read a benchmark score.
“Agent = Model + Harness”
Short version: a harness is everything that is not the model
One thing this is not: a pricing story. A harness does not change the per-token rate your vendor charges — it changes what you get for that rate (score, latency, token consumption). In practice, teams talk about swapping models or rewriting prompts. Very often the thing moving the numbers is the design around the model. That is why “harness” is used as a term in its own right: without naming the harness, two people comparing benchmark scores are usually not comparing the same thing at all.
The mental model: the model is the brain, the harness is the body and the kit
DeepSeek's own page puts the relationship in two short sentences.
What is inside a harness (from the vendor's own list)
“Plumbing” is too vague to be useful, so here is the parts list. DeepSeek Harness says every agent capability is implemented as a plugin, and names them.
| Harness component | What it decides | Why the number moves |
|---|---|---|
| System prompt | What the model is shown, and what role and constraints it gets | Different instructions, same model, different behaviour |
| Tools | Which external capabilities are available (file editing, shell, search, APIs) | More tools can solve more tasks; fewer tools change what is being measured |
| Skills and procedures | Whether a workflow or checklist is handed over up front | A given procedure removes exploration overhead |
| Sessions and memory | How prior turns and intermediate results are carried forward | If carry-over is assumed, long tasks lose less accuracy |
| Sandbox | Where it executes (isolation, permissions, network) | A different environment changes which actions are even possible |
| Loops and retries | How many times it may reconsider or retry | More attempts raise the success rate; the cap is part of the result |
| Scheduling | How parallel work and subagents are ordered | Sequencing changes time and step count |
| Logging | How inputs, reasoning and tool calls are recorded | What you cannot replay, you cannot improve or verify |
| Scoring and evaluation | How results are counted (pass/fail, partial credit, cut-offs) | Change the counting and the same run scores differently |
Everything except the quoted list is our plain-language expansion of DeepSeek Harness's own items (models, tools, skills, sessions, sandboxes, storage, loops, scheduling, UI). “Scoring and evaluation” is our addition as a benchmark-side design element; it is not a DeepSeek Harness item.
Why the numbers move: ARC-AGI-3, with the conditions attached
Abstract advice is hard to act on, so here is one measured case. Same model (GPT-6 Astra), same benchmark (ARC-AGI-3), two published results under different harnesses. The model setting also differs — Astra (max) versus Astra (high) — so the gap between 62.7% and 99.9% cannot be attributed to the harness alone.
| Harness (measurement condition) | Model setting | Score | Cost | What it does |
|---|---|---|---|---|
| Standard harness (ARC Prize's neutral harness = the same minimal interface for every model) | Astra (max) | 62.7% | $26K | The model carries forward notes it chooses to keep; the same minimal interface for every model |
| Provider Adapter harness (OpenAI's) | Astra (high) | 99.9% | $19K | Preserves opaque reasoning state between requests, and uses compaction for longer conversations so the model can reuse prior work |
- Both are state of the art under their own conditions. ARC Prize treats both as SOTA. This is not a case of one number being inflated and the other being correct.
- A like-for-like comparison exists too. ARC Prize also published 7.8% for GPT-5.6 Sol on the same Standard harness (checked 2026-09-15). That is roughly an eightfold step on one playing field. Putting 99.9% next to 7.8% compares two different harnesses.
- The benchmark operator itself has committed to labelling. See the quote below.
Sources: ARC Prize, “OpenAI's GPT-6 Astra on ARC-AGI-3” (checked 2026-09-19) and the ARC-AGI results page. Scores and costs are ARC Prize's published figures, not our own measurement. Background on the same harness architecture is in entry 1.
What does that look like in practice? DeepSeek Harness
“Harness” is an abstract term, so a real one helps. DeepSeek ships DeepSeek Harness in developer preview, with the source code released at the same time. Here is only what its official page states.
| What the page says | What it means |
|---|---|
| Everything is a plugin | Models, tools, skills, sessions, sandboxes, storage, loops, scheduling and the UI are all swappable, recomposable parts |
| Cordis kernel | Handles mounting, unmounting and dependencies of plugins; the agent's capabilities live in the plugins |
| Compose with configuration | Select, swap or extend any capability in configuration, without changing the harness source code |
| Every run is traceable | Everything the model sees is recorded in an append-only session log: system prompts, reasoning, tool calls and results, subagent scheduling, context injections. Resume, fork, search and replay all work on that same event stream |
Four runtime modes — and one of them is explicitly for benchmarking
The page lists four runtime modes. All four run the same kind of model; what differs is the kit handed to it. That is a harness difference, shipped as a feature.
Source: DeepSeek Harness official page (checked 2026-09-19; developer preview). Quoted phrases are the page's own words; the rest is our summary. Repository: github.com/deepseek-ai/deepseek-harness. We are not affiliated with DeepSeek and this is not a product recommendation.
The picture: the harness is what sits outside the model
Everyday version: the same player, a different score
| In sport | In an AI agent |
|---|---|
| The player's ability | The model (trained weights; not yours to edit) |
| Court size, the ball | Sandbox, available tools |
| The referee and the rulebook | Scoring and evaluation (what counts as correct) |
| Substitutions and timeouts | Retries and loop limits |
| The coach's game plan | System prompt, skills, procedures |
| Carrying last game's notes in | Sessions and memory — the thing that separated 62.7% from 99.9% on ARC-AGI-3 |
One takeaway from the analogy: when a score is discussed, separate the player's story from the court-and-referee story. That is most of what reading a benchmark requires.
How to use this at work
Three checks when you see a score
- Whose harness? The vendor's own, or a neutral one run by the benchmark operator. On ARC-AGI-3 that was 62.7% versus 99.9%.
- Is that harness close to your setup? A score that assumes reasoning state can be carried across requests will not reproduce in a setup that cannot carry it. And vice versa.
- What are the scoring and the limits? Retry caps, cut-offs, partial credit. Numbers measured under different rules do not make a ranking.
- When your agent is not good enough, suspect the harness first. Tightening the toolset, fixing how memory is carried, or revisiting retry limits can beat swapping the model. ARC-AGI-3 is the same model under two harnesses producing two very different results.
- An improvement you cannot replay is not an improvement. DeepSeek makes “everything the model sees is recorded in an append-only session log” a headline feature precisely so harness changes can be verified. If you build your own, do not defer logging.
- This is separate from pricing. A harness affects performance — score, latency, token consumption — not the per-token rate itself. The pricing tables cover rates; this entry is about what you get for them.
In one sentence
A harness is everything that is not the model
- The model is the brain. The harness is the body, the tools, the plan and the referee that put it to work.
- It includes: system prompt, tools, skills, sessions and memory, sandbox, loops and retries, scheduling, logging, scoring and evaluation.
- That is why the same model on the same benchmark produces different numbers. On ARC-AGI-3, ARC Prize published 62.7% under the Standard harness (Astra max) and 99.9% under the Provider Adapter harness (Astra high). The settings differ as well, so not all of that gap is the harness.
- When you read a benchmark, check which harness before you check which model.
The term is useful because it moves the discussion from “is this a good model” back to “how is my setup designed”. Same model, different harness, different result. Next time you see a score, look for the harness name first.
Next
The worked example that introduces harnesses is entry 1 (how to read benchmark numbers); cache pricing is entry 2 (prompt caching).
🧪 All glossary entries →Glossary ・ How to read benchmark numbers ・ Prompt caching ・ Pricing ・ Release timeline
⚠️ Disclaimer
- Quotes and figures were checked against primary sources (DeepSeek's official page, the ARC Prize blog) on 2026-09-19. ARC-AGI-3 scores and costs are ARC Prize's published figures, not our measurements.
- Vendor statements are quoted verbatim; the surrounding explanation is ours. Where the two differ, the vendor's wording governs.
- DeepSeek Harness is in developer preview; features and APIs will change. Check the official page and repository before adopting it.
- We are not affiliated with any provider, and this is not a recommendation of any product.