GPT-6 Astra: How to Read "ARC-AGI-3 99.9%, ExploitBench 100%"

Installment 1 of Glossary. We take one line from an announcement and walk it back to primary sources. Today's line comes from GPT-6 Astra: the numbers are accurate, the conditions are missing. Without the conditions, this line cannot support a decision.

"ARC-AGI-3 99.9% · ExploitBench 100% · Terminal-Bench Science 64.6%. A 1.05M context window and 128K max output."

Model: GPT-6 Astra (OpenAI, announced 2026-09-03, $10 / $50 per 1M tokens) — full page · release timeline

● Last updated:  |  Primary sources: OpenAI announcement · ARC Prize blog · OpenAI model docs · Terminal-Bench-Science 0.1  |  Verified: 2026-09-15

First, split it into five numbers

The line stacks up five figures of three completely different kinds. Even the percentages are not the same species of number.

ARC-AGI-3
99.9%
Interactive reasoning: exploring an unknown environment and acquiring the goal itself. The benchmark is saturated
ExploitBench
100.0%
Turning real software vulnerabilities into working exploits. A perfect score is a ceiling
Terminal-Bench Science 0.1
64.6%
70 workflows that researchers actually ran. Still unsaturated
Context window
1,050,000
How much it can read at once. Past 272K tokens, the whole request costs 2x input
Max output
128,000
How much it can write at once. Use all of it and a single request costs $6.40

So "99.9%" is a saturated-benchmark number, "100%" is a hit-the-ceiling number, and "64.6%" is a number with room left in it. Same-looking list, very different information density.

1. Same benchmark, different playing field — 99.9% vs 62.7%

Short version: an ARC-AGI-3 score cannot be compared until you know which harness produced it. OpenAI's announcement says 99.9%. ARC Prize, which runs the benchmark, published 62.7% under its neutral harness for the same model on the same benchmark.
ConditionScoreCostWhat it does
Standard harness (neutral)62.7%$26,098The model chooses which notes to carry forward; the same minimal interface for every model
Provider Adapter harness (OpenAI's)99.9%~$19KPreserves opaque reasoning state between requests and compacts longer conversations

Sources: ARC Prize, "OpenAI's GPT-6 Astra on ARC-AGI-3" and ARC-AGI results page (verified 2026-09-15). OpenAI's own announcement notes that Astra was run on ARC-AGI-3 with its responses API harness, "which changes two settings to better match real-world performance."

What the gap actually means

  • The like-for-like comparison is 62.7% vs GPT-5.6 Sol's 7.8%. That is still a step change — ARC Prize calls it a genuine advance in interactive reasoning. Quoting 99.9% next to 7.8% compares two different harnesses.
  • ARC Prize says it will publish both harness results on the leaderboard, labelled. The benchmark's own operator is telling you that a score without a named harness is not readable.
  • The gap is software, not just the model. Under the Provider Adapter, aggregate elapsed time was about 3.66x faster and token use fell 49% across the 167 game-reasoning pairs both harnesses solved. If your runtime cannot carry reasoning state forward, the 99.9% is a demonstration of someone else's stack.
  • ARC Prize explicitly declined the AGI conclusion: saturating this benchmark is not proof of general intelligence. When the people who built the test hedge, take the hedge seriously.

Background: ARC-AGI-3, launched 2026-03-25 by the ARC Prize Foundation, is the first interactive reasoning benchmark — no instructions, no stated goal, and agents must build a world model as they go. At launch every frontier model scored below 1% while humans solved every environment (ARC-AGI-3).

2. A perfect score means "we can no longer measure it" — ExploitBench 100%

Short version: 100% does not mean "far ahead". It means the measurement ran out of headroom. ExploitBench asks whether a model can take real software vulnerabilities and produce exploits that actually run. Astra scored a perfect 100% (Sol 78.5%, Claude Fable 5.1 70%), so the benchmark no longer separates the leaders.
  • The model behind the leader cannot be distinguished on this test. The next comparison moves to ExploitGym (Astra 42.4% vs Sol 30.3%) and to fresh benchmarks. A number disappears when capability rises — and when a test fills up.
  • The contamination caveat comes from the vendor itself. Because scores may be helped by exposure to historical vulnerabilities, OpenAI also evaluated Astra on two novel benchmarks. Read the perfect score with that reservation attached.
  • OpenAI classified Astra as "Critical" under its Preparedness Framework — a first. The most advanced cyber capability ships gated, not open. "Most capable" and "available to everyone" are separate claims.

Source: OpenAI's GPT-6 Astra announcement (verified 2026-09-15). Both models were evaluated without production safeguards, per the same announcement.

3. The unsaturated benchmark carries the most information — Terminal-Bench Science 64.6%

Short version: 64.6% is a number from a test nobody has finished. That makes it useful for both capability and procurement. Terminal-Bench-Science 0.1 is led by researchers at Stanford University with the Terminal-Bench / Harbor team and domain experts: 70 expert-curated workflows across five domains (life, physical, Earth, mathematical, engineering sciences), three independent trials per task, graded on concrete artifacts — analyses, simulations, proofs, code, data products.
ModelTB-Science 0.1 (OpenAI's figures)Neutral leaderboard (snapshot)
GPT-6 Astra64.6%—
Claude Fable 5.152.6%—
GPT-5.6 Sol22.4%22.4%
Claude Opus 5 (top of the neutral board)—30.0%
  • The 64.6% is OpenAI's own condition. On the public leaderboard snapshot, the top row is Claude Opus 5 at 30.0%. Same benchmark, very different numbers — the harness problem from section 1 shows up again.
  • And yet this is the most usable figure of the five, because the test is genuinely hard: models that clear 80%+ on general terminal/coding benchmarks still fail the large majority of real research workflows.
  • OpenAI also published low-effort numbers: Terminal-Bench Science 61.1% at roughly 27% lower cost, and GPQA Diamond 94.9% at about 37% lower cost. Their claim is that turning reasoning effort down does not immediately break quality — a self-reported figure, not an independently verified one.

Sources: OpenAI announcement · Terminal-Bench-Science 0.1 release · GitHub (structure and contributing institutions) · neutral leaderboard mirror (BenchLM, snapshot as of September 2026) (verified 2026-09-15).

ARC-AGI-3: same model, same benchmark — 62.7% under the neutral harness vs 99.9% under OpenAI's Provider Adapter ARC-AGI-3: same model, same benchmark Different harness, different number. Standard harness (neutral) 62.7% $26,098 Provider Adapter harness (OpenAI) 99.9% ~$19K 0% 50% 100% Source: ARC Prize (2026-09-03). Both rows are ARC Prize's own published results.
Same benchmark, same model, two numbers. The harness moves the result this far (source: ARC Prize).

4. What "1.05M context" means in practice

Short version: 1.05M is about what fits. The 272K boundary is about what it costs. At roughly 0.75 words per token, 1,050,000 tokens is on the order of 780,000 words. The specs — 1,050,000 context, 128,000 max output, knowledge cutoff 2026-04-30 — are on the model page and in OpenAI's model docs.
How you use itRate ruleInput cost for one request
272K tokens or lessStandard ($10 / 1M input)up to $2.72
More than 272K tokensThe entire request is billed at 2x input and cache rates, 1.5x output1.05M → $21.00
Same, with a cache hit ($1.00 / 1M doubled)Cached input is 10% of input, but the 2x rule still applies1.05M → $2.10
Maxing out the output (128K × $50 / 1M)Output is 5x the input rate$6.40 (output side)
  • The boundary is 272K tokens, and it is not pro-rated. One token over and the whole request is billed at 2x input and 1.5x output. Dropping a giant document in once is where the bill jumps.
  • Caching only pays when the same long prefix is reused. Cached input is 10% of input ($1.00 / 1M); cache writes cost 1.25x uncached input ($12.50 / 1M). A one-shot giant context, with no cache design, is the most expensive way to use this model.
  • Long output hits the bill hardest. At $50 / 1M, maxing the 128K output costs $6.40 in a single request — more than a full 272K-token input. Splitting work and generating only what you need is cheaper than one giant generation.
  • "It fits 1.05M" does not mean "send 1.05M." Design under 272K and you pay list price; design over it and everything doubles. That boundary is the real design budget. Run your own numbers in the token calculator.

Source: OpenAI GPT-6 Astra model docs (context, max output, cutoff, rate rules; verified 2026-09-15). Dollar figures are simple arithmetic on published rates; an actual invoice sums input and output.

How to read the line (summary)

Three checks, every time

  • Whose playing field is it? The vendor's harness or a neutral one. On ARC-AGI-3 that was 99.9% versus 62.7%.
  • Has it hit the ceiling? A perfect or saturated score means the benchmark stopped separating models — a signal to move to the next axis (ExploitGym, fresh benchmarks).
  • Is there a like-for-like comparison? Numbers measured under different conditions do not make a ranking. Until the conditions match, a benchmark score is not yet readable.

The information in this line is not in the "99.9%" — it is in the gap between 99.9% and 62.7% and in the unsaturated 64.6%. The first is homework for your runtime design; the second is procurement input. Next time a model launches, apply the same three checks.

Next

With the conditions settled, the next question is cost: GPT-6 Astra pricing and availability, the Claude Fable 5.1 comparison, and your own numbers.

🧮 Run the token calculator →

GPT-6 Astra · Pricing · Timeline · DeepSeek V4.1 Flash · Prompt caching · Glossary

⚠️ Notes and disclaimers

  • Benchmark figures are as published by the benchmark operators or model vendors, and they differ in harness, trial count and reasoning effort. This page states those conditions rather than hiding them.
  • Prices and specifications were verified on vendor pages on 2026-09-15 and change without notice. Confirm on the official page before committing.
  • Cost conversions are simple arithmetic on published rates and do not guarantee an invoice amount.