AI Governance & Data Protection  ·  4 September 2026

The Harness Does the Talking

OpenAI's own adapter scored 36 points higher than the benchmark maintainer's standard harness, on the identical model. The difference between the two is now a product, a price list, and a processing activity nobody can fully account for.
By Alan Wright  ·  The Haunted Lighthouse Limited  ·  Peel, Isle of Man

OpenAI's GPT-6 Astra scored 98.6% on ARC-AGI-3, a benchmark built specifically to resist the kind of brute-force scaling that inflates most AI leaderboards. When ARC Prize, the benchmark's own maintainer, ran the identical model through its standard test harness instead of OpenAI's, the score dropped to 62.7%. Same weights. Same reasoning setting. A 36-point swing produced entirely by the software wrapped around the model, not by the model itself.

That gap is this week's clearest demonstration of a pattern that has been building all year: the model is becoming the commodity, and the harness around it is becoming the product. It is also, on closer reading, a data protection problem that none of the parties selling harnesses seem in a hurry to discuss.


What a harness actually is

Strip the branding and a "harness" is ordinary infrastructure: what state the system carries between requests, what it is permitted to touch, what it forgets, and when it has to stop and check with a human. ARC Prize's standard harness lets a model carry forward only the notes it explicitly chooses to write out. OpenAI's Provider Adapter instead preserves the model's internal reasoning state between calls and compresses long conversations, so the system resumes its own thinking rather than reconstructing it from scratch each turn.

The published numbers, broken down by reasoning effort, show how far that lever moves the outcome:

ARC-AGI-3, GPT-6 Astra, by reasoning effort (score / cost per full evaluation)
Reasoning effortARC Prize standard harnessOpenAI Provider Adapter
Max62.7% for $26,09898.6% for $17,332
High54.8% for $40,70599.9% for $18,817
Medium38.6% for $48,09098.4% for $19,285
Low17.5% for $38,16698.0% for $21,298
None35.2% for $49,79196.7% for $23,457

The 98.6% headline figure is itself worth a second look before anyone runs with it, because it is not even the best number in the table. That score belongs to Astra at Max reasoning effort, the setting OpenAI's own marketing quoted. One row up, at High reasoning effort, the same adapter scored 99.9%, higher, for less money. Inside the vendor's own harness, turning reasoning down a notch from Max to High produced a better result. Run the full spread and the adapter scores 96.7% with no reasoning at all, climbs to 99.9% at High, then drops back to 98.6% at Max, which is not a capability curve, it is the harness quietly flattening whatever the reasoning dial was supposed to be measuring. Paying for Max reasoning effort inside this system buys a worse score than High, at a higher price. That is not the model thinking harder. That is the adapter erasing the difference between settings the vendor sells as meaningfully distinct.

Astra with no reasoning effort at all, inside OpenAI's harness, still beat the same model at maximum reasoning effort inside ARC Prize's harness, and cost less doing it. Across the 167 game-reasoning pairs both harnesses solved, ARC Prize measured the Provider Adapter runs using 49% fewer tokens and running roughly 3.66 times faster. The reasoning dial, the thing every vendor markets as the knob that matters, was outperformed outright by the plumbing around it, and, on this evidence, quietly scrambled by it too.

None of this is secret in the sense of hidden code; OpenAI's adapter runs on documented Responses API capabilities anyone can call. What is not for sale is the assembled system that produced the 98.6% figure. You can buy the parts. You cannot buy the thing that was measured.


The commercial layer

Anthropic, Google, and Microsoft have each concluded the same thing and are pricing accordingly. Anthropic meters its Managed Agents at $0.08 per session hour on top of token costs. Google and Microsoft itemise memory, code execution, and observability as separate line items. OpenAI is currently giving its Agents SDK away with no runtime fee, which reads less like generosity and more like a land grab ahead of the same metering everyone else has already introduced.

The clearest signal of where the money is actually going sits outside any single lab. Stripe paid a reported $8 billion in August for OpenRouter, a gateway that routes roughly 10 trillion tokens a day across more than 400 models for 10 million developers. That is not a model acquisition. It is the purchase of the routing and observability layer that sits in front of the models, the same category of thing every major lab is now building and metering internally.

Worth noting for the record: The New Stack's owner, Insight Partners, is a disclosed investor in both OpenAI and Anthropic. The publication ran the comparison anyway and let ARC Prize's own numbers do the talking, which is the right way to handle it, but it is exactly the kind of institutional-claim-versus-operational-reality gap that belongs in a Theatre Pulldown footnote of its own.

That gap showed up again days later, when OpenAI's president told a press briefing it was not unreasonable to feel the industry had entered the "AGI era." The claim describes a benchmark result produced by one particular, non-purchasable system. It does not establish anything about the underlying model. Maybe the industry is somewhere close to that threshold. This benchmark does not prove it, and the 36-point gap between two configurations of the same model is a strange foundation to build the claim on.


The part the benchmark story missed: your data protection exposure

Buried inside "the adapter preserves the model's opaque reasoning state between requests" is a sentence that should stop any DPO cold, because it describes a live processing activity that neither the vendor nor the customer can fully see into.

Run it through UK GDPR mechanically:


Where the liability actually sits

This is the asymmetry that makes the commercial pitch worth reading twice. Vendor SLA remedies for agentic platforms are typically service credits, with total liability capped at fees paid over the preceding twelve months. Your actual exposure if that opaque state mishandles personal data, in breach notification costs, regulatory fines, and reputational damage, will not be capped at anything close to that figure. You carry effectively all of the downside and rent a small fraction of the risk mitigation, in exchange for a benchmark number you cannot even reproduce yourself.

The industry has quietly converged on a version of "copilot, not root" for the harness itself: keep the state, the permissions, and the audit trail somewhere you can actually see them, or accept that the number on the leaderboard was never describing the system you were about to buy.


Sources


Editor's note: Companion piece to "If You Can't Trust the People Who Built It" and "The Sign-Out That Doesn't Sign You Out", same underlying pattern: a vendor's own architecture outrunning anyone's ability to audit it. Framing term drawn from "The Theatre Pulldown".

Questions about this analysis, or interested in working with The Haunted Lighthouse?
consultancy@haunted.lighthouse.co.im

The Sovereign Auditor covers digital sovereignty, cybersecurity governance, and data protection policy, with particular focus on Isle of Man jurisdiction and Crown Dependency issues.

Support independent analysis. Subscribe directly, or scan on your phone.

Payments via PayPal. Credentials delivered by email. No Substack. No Stripe. No middlemen.