GPT-6 Astra: How to Read "ARC-AGI-3 99.9%, ExploitBench 100%"
Installment 1 of Glossary. We take one line from an announcement and walk it back to primary sources. Today's line comes from GPT-6 Astra: the numbers are accurate, the conditions are missing. Without the conditions, this line cannot support a decision.
"ARC-AGI-3 99.9% · ExploitBench 100% · Terminal-Bench Science 64.6%. A 1.05M context window and 128K max output."
First, split it into five numbers
The line stacks up five figures of three completely different kinds. Even the percentages are not the same species of number.
So "99.9%" is a saturated-benchmark number, "100%" is a hit-the-ceiling number, and "64.6%" is a number with room left in it. Same-looking list, very different information density.
1. Same benchmark, different playing field — 99.9% vs 62.7%
| Condition | Score | Cost | What it does |
|---|---|---|---|
| Standard harness (neutral) | 62.7% | $26,098 | The model chooses which notes to carry forward; the same minimal interface for every model |
| Provider Adapter harness (OpenAI's) | 99.9% | ~$19K | Preserves opaque reasoning state between requests and compacts longer conversations |
Sources: ARC Prize, "OpenAI's GPT-6 Astra on ARC-AGI-3" and ARC-AGI results page (verified 2026-09-15). OpenAI's own announcement notes that Astra was run on ARC-AGI-3 with its responses API harness, "which changes two settings to better match real-world performance."
What the gap actually means
- The like-for-like comparison is 62.7% vs GPT-5.6 Sol's 7.8%. That is still a step change — ARC Prize calls it a genuine advance in interactive reasoning. Quoting 99.9% next to 7.8% compares two different harnesses.
- ARC Prize says it will publish both harness results on the leaderboard, labelled. The benchmark's own operator is telling you that a score without a named harness is not readable.
- The gap is software, not just the model. Under the Provider Adapter, aggregate elapsed time was about 3.66x faster and token use fell 49% across the 167 game-reasoning pairs both harnesses solved. If your runtime cannot carry reasoning state forward, the 99.9% is a demonstration of someone else's stack.
- ARC Prize explicitly declined the AGI conclusion: saturating this benchmark is not proof of general intelligence. When the people who built the test hedge, take the hedge seriously.
Background: ARC-AGI-3, launched 2026-03-25 by the ARC Prize Foundation, is the first interactive reasoning benchmark — no instructions, no stated goal, and agents must build a world model as they go. At launch every frontier model scored below 1% while humans solved every environment (ARC-AGI-3).
2. A perfect score means "we can no longer measure it" — ExploitBench 100%
- The model behind the leader cannot be distinguished on this test. The next comparison moves to ExploitGym (Astra 42.4% vs Sol 30.3%) and to fresh benchmarks. A number disappears when capability rises — and when a test fills up.
- The contamination caveat comes from the vendor itself. Because scores may be helped by exposure to historical vulnerabilities, OpenAI also evaluated Astra on two novel benchmarks. Read the perfect score with that reservation attached.
- OpenAI classified Astra as "Critical" under its Preparedness Framework — a first. The most advanced cyber capability ships gated, not open. "Most capable" and "available to everyone" are separate claims.
Source: OpenAI's GPT-6 Astra announcement (verified 2026-09-15). Both models were evaluated without production safeguards, per the same announcement.
3. The unsaturated benchmark carries the most information — Terminal-Bench Science 64.6%
| Model | TB-Science 0.1 (OpenAI's figures) | Neutral leaderboard (snapshot) |
|---|---|---|
| GPT-6 Astra | 64.6% | — |
| Claude Fable 5.1 | 52.6% | — |
| GPT-5.6 Sol | 22.4% | 22.4% |
| Claude Opus 5 (top of the neutral board) | — | 30.0% |
- The 64.6% is OpenAI's own condition. On the public leaderboard snapshot, the top row is Claude Opus 5 at 30.0%. Same benchmark, very different numbers — the harness problem from section 1 shows up again.
- And yet this is the most usable figure of the five, because the test is genuinely hard: models that clear 80%+ on general terminal/coding benchmarks still fail the large majority of real research workflows.
- OpenAI also published low-effort numbers: Terminal-Bench Science 61.1% at roughly 27% lower cost, and GPQA Diamond 94.9% at about 37% lower cost. Their claim is that turning reasoning effort down does not immediately break quality — a self-reported figure, not an independently verified one.
Sources: OpenAI announcement · Terminal-Bench-Science 0.1 release · GitHub (structure and contributing institutions) · neutral leaderboard mirror (BenchLM, snapshot as of September 2026) (verified 2026-09-15).
4. What "1.05M context" means in practice
| How you use it | Rate rule | Input cost for one request |
|---|---|---|
| 272K tokens or less | Standard ($10 / 1M input) | up to $2.72 |
| More than 272K tokens | The entire request is billed at 2x input and cache rates, 1.5x output | 1.05M → $21.00 |
| Same, with a cache hit ($1.00 / 1M doubled) | Cached input is 10% of input, but the 2x rule still applies | 1.05M → $2.10 |
| Maxing out the output (128K × $50 / 1M) | Output is 5x the input rate | $6.40 (output side) |
- The boundary is 272K tokens, and it is not pro-rated. One token over and the whole request is billed at 2x input and 1.5x output. Dropping a giant document in once is where the bill jumps.
- Caching only pays when the same long prefix is reused. Cached input is 10% of input ($1.00 / 1M); cache writes cost 1.25x uncached input ($12.50 / 1M). A one-shot giant context, with no cache design, is the most expensive way to use this model.
- Long output hits the bill hardest. At $50 / 1M, maxing the 128K output costs $6.40 in a single request — more than a full 272K-token input. Splitting work and generating only what you need is cheaper than one giant generation.
- "It fits 1.05M" does not mean "send 1.05M." Design under 272K and you pay list price; design over it and everything doubles. That boundary is the real design budget. Run your own numbers in the token calculator.
Source: OpenAI GPT-6 Astra model docs (context, max output, cutoff, rate rules; verified 2026-09-15). Dollar figures are simple arithmetic on published rates; an actual invoice sums input and output.
How to read the line (summary)
Three checks, every time
- Whose playing field is it? The vendor's harness or a neutral one. On ARC-AGI-3 that was 99.9% versus 62.7%.
- Has it hit the ceiling? A perfect or saturated score means the benchmark stopped separating models — a signal to move to the next axis (ExploitGym, fresh benchmarks).
- Is there a like-for-like comparison? Numbers measured under different conditions do not make a ranking. Until the conditions match, a benchmark score is not yet readable.
The information in this line is not in the "99.9%" — it is in the gap between 99.9% and 62.7% and in the unsaturated 64.6%. The first is homework for your runtime design; the second is procurement input. Next time a model launches, apply the same three checks.
Next
With the conditions settled, the next question is cost: GPT-6 Astra pricing and availability, the Claude Fable 5.1 comparison, and your own numbers.
🧮 Run the token calculator →GPT-6 Astra · Pricing · Timeline · DeepSeek V4.1 Flash · Prompt caching · Glossary
⚠️ Notes and disclaimers
- Benchmark figures are as published by the benchmark operators or model vendors, and they differ in harness, trial count and reasoning effort. This page states those conditions rather than hiding them.
- Prices and specifications were verified on vendor pages on 2026-09-15 and change without notice. Confirm on the official page before committing.
- Cost conversions are simple arithmetic on published rates and do not guarantee an invoice amount.