THE DAILY DOWNLOAD

The Harness Gap. The Same AI Scored 62 Percent And 99 Percent On The Same Test In The Same Week

Connor T. MacIvor·AI implementation, Santa Clarita Valley·

Last week a computer sat down to take a test built specifically to be hard for computers. It scored 62.71 percent.

That same computer, on the same test, in the same week, also scored 99.95 percent.

Nothing about the machine changed between those two numbers. What changed was who set up the room.

One of those numbers went into the press release. You can guess which one.

The harness gap. Same model, same test, same week, two very different numbers.

Hold one number for me while we go: 12.78 dollars. I will come back to it, and when I do it may change how you hear the phrase *better than a human*.

What actually happened

OpenAI released a model called GPT-6 Astra and started rolling it out across ChatGPT plans. The company said it was the best model in the world for computer use, professional work, science, coding and cybersecurity.

That is a company claim about its own product. That is not a scandal. That is Tuesday.

The number they put in front of it was 99.9 percent on ARC-AGI-3.

What ARC-AGI-3 actually is

The name is doing a lot of work, so let me tell you what it is.

It is a set of small puzzle games. You are dropped in with no instructions and you have to work out the rules by poking at it. It was designed by a researcher named François Chollet specifically to test whether a machine can handle a situation it was never trained on. Not knowledge. Not memory. Can you walk into a room you have never seen and figure out what the rules are.

And here is the part that matters more to me than anything else in this piece: it is run by an outside organization, ARC Prize, that is not owned by any of these labs.

On September 3rd, ARC Prize published their own results. Not the press release. The measurement.

Two runs, and the entire story sits between them

They ran Astra two different ways. I want to be exact here, because the difference *is* the story.

Run one, the standard harness. The model gets to keep the notes it chooses to write down for itself as it goes.

Score: 62.71 percent. Cost to run: 26,098 dollars.

Run two, the provider adapter. The model's own internal reasoning state gets carried between requests using the company's own plumbing, so the machine picks up exactly where it left off with everything it was already thinking.

Score: 99.95 percent. Cost to run: 18,817 dollars.

Same model. Same puzzles. Same week. 62.71 and 99.95.

The distance between those two numbers is not intelligence. It is scaffolding. It is who built the desk and who kept the notes.

The two students

Think about a test in school. Two students, same brain, same questions. One takes it cold. The other gets to keep every piece of scratch work from every previous attempt, organized for them and handed back at the start of each question.

The second student scores higher. Of course they do.

But you have not learned that the second student is smarter. You have learned that good notes help.

I am going to call that the harness gap, and once you can see it you are going to see it everywhere for the rest of your life.

Nobody lied to you

Here is the part nobody says out loud. Both numbers are correct. Both are real. Neither one is a lie.

ARC Prize published them right next to each other, which is exactly what an organization does when it is not trying to sell you anything.

The company quoted the higher one. That is not fraud. That is marketing.

But when you hear 99.9 and you picture a machine solving something on its own, you have been handed a true number that creates a false picture.

Now the receipt. 12.78 dollars.

To score these puzzles you need to know what a person gets. So ARC Prize hired people. Real ones. They paid 115 dollars for a 90 minute session plus 5 dollars for every game finished.

That works out to roughly 12.78 dollars per attempted game.

So when somebody tells you the machine beat the human at this, the sentence is technically true and financially absurd. The machine beat the human the way a helicopter beats a man.

I want to be fair about it. It will get cheaper. It always gets cheaper. But right now, today, September 2026, that is the trade. And anybody quoting you the score without the receipt is telling you half.

"So why not run your own model?"

Fair question, and it deserves a straight answer.

The open models you can download today are genuinely useful for summarizing, drafting, tagging and cleaning up messy text. Nothing you type into them leaves your building. That is a real advantage and not a small one.

What it costs you is a machine with serious memory, which is thousands of dollars up front, plus the electricity, plus the part nobody puts in the sales pitch: somebody has to install it, update it, and work out why it stopped on a Wednesday. If that person is you, that is your evening. If that person is a contractor, that is a bill.

So the trade is not free versus paid. It is a subscription and somebody else's problem, against a purchase and your problem.

For a shop with six people, the subscription usually wins, right up until your data gets sensitive enough that it stops winning.

The other side, because I am not here to talk you out of being impressed

Astra scored 98.5 percent on ARC-AGI-1 and 95 percent on ARC-AGI-2, verified by the same outside organization, with no special plumbing. That is genuinely, seriously good. Those tests broke every model that came before it.

And Chollet, the man who built the benchmark, the person with the most professional reason on earth to say *not so fast*, was asked whether his 2030 forecast still held.

He said sooner. He described progress arriving about twice as fast as his own estimates.

That is not marketing. That is a skeptic updating himself, which is the most credible thing a person can do.

Then ARC Prize went further. They said they are not claiming this is AGI, and that saturating the benchmark would not represent proof of achieving AGI, because the test does not represent the complexity and open-endedness of the real world.

Read that again. The people who built the test, whose entire organization exists because of the test, are telling you the test is not the thing.

When was the last time you saw anybody do that?

Four numbers, one model, all defensible

Two different outfits scored the same model last week and reached opposite conclusions.

Epoch AI ranked it first overall, 169 points, across more than 50 benchmarks.

Artificial Analysis scored it 61. Level with the model it replaced, and behind Claude Fable 5.1.

Best in the world, and no better than last time. Same machine, same week. The difference is what each group chose to measure. Epoch leaned on math and knowledge. Artificial Analysis leaned on coding and comprehension.

So we now have four separate numbers for one model, ranging from world beating to no change at all, and every one of them is defensible.

That should tell you something permanent about benchmarks. A benchmark is not a measurement of a machine. It is a measurement of a machine doing one specific thing, under one specific setup, chosen by one specific group of people who had a reason for choosing it.

Watch what happens to these numbers next, because this is the part that costs you money

The labs will take the 99.9 and use it to tell you the future is arriving and you had better subscribe before you fall behind.

The podcasters and the pundits and the people running for something will take the exact same 99.9 and use it to tell you the machines are past us now, your kids are in danger, subscribe, donate, vote.

Opposite conclusions. Same number. Same business model.

Neither one will mention the 62.71, because the 62.71 does not sell anything.

Frightened people need a savior, and saviors get paid.

There is no outside force here. Nobody found this thing in a cave. People built it. People chose what to test. People chose which number to print. Every setup was a decision a human being made.

Decisions can be examined. Monsters cannot.

The one number nobody chose

On Friday, September 4th, the Bureau of Labor Statistics reported the August jobs numbers. 162,000 jobs added. Unemployment steady at 4.1 percent. Average hourly earnings 37.75 dollars, up 3.1 percent from a year ago.

Inside that report: the information sector lost 23,000 jobs. Construction added 22,000. Almost a mirror image, same month, same country.

The people who write the code are shrinking. The people who pour the concrete are hiring.

Be careful with that. The information sector is not just AI, and a single month is not a trend. But if you have been told that technology takes every job, the government's own count says the losses are landing in one specific room of the house, and it is the room where the software people sit.

And when a company announces AI layoffs, there are always three possibilities and only one of them is in the memo. Either the machine genuinely does the work now, or the company needed to cut anyway and AI is the respectable word for it, or they are cutting people to pay for the AI. That third one is real and almost nobody names it.

What this means if you run a shop in Santa Clarita

Stop shopping for the smartest model.

The harness gap just proved that the smartest model, wired up wrong, scores 62 instead of 99. Your setup matters more than your model.

Here is what I actually do when I walk into a business. I ask how a lead comes in. I ask what happens in the first 10 minutes. I ask who touches it. I ask where it goes to die.

Almost every time, what I find is not a job for artificial intelligence. It is a leak.

One business I looked at had a real problem: 4 out of 10 people who called after 5 PM never got called back the next morning, because the message pad lived next to the register and the morning person did not work the register.

That is not artificial intelligence. That is a form and a text message.

An agent belongs in exactly one place: where real judgment calls have to be made. And before you put one there, write down your philosophy, your positioning, and the lines it must never cross. Then supervise it for 30 days like a new hire, because that is what it is. At LAPD we were on probation for a year after the academy. A year. Let go at any time.

That is where small businesses beat the giants, and almost nobody believes it. You can spot a leak on Tuesday and have it fixed by Thursday. The large company needs 40 signatures to change a phone script. Right now, being small is the advantage.

Call an agent into existence and go fishing is not here yet. It may never be. Anybody certain either way is selling something.

Two guardrails

One. What you type. Every time you paste a client list, a contact, or your pricing into a public model, ask two questions. How important is this information, and what is the vendor's retention policy in one sentence. If you cannot answer the second part, do not paste the first part. That is the whole rule.

Two. Regulation is already here and most owners do not know it. Texas passed the Responsible Artificial Intelligence Governance Act and it took effect January 1st of this year, with civil penalties and the Attorney General enforcing it. There is still no comprehensive federal law, so what governs you depends on which state you are in. That is going to stay true for a while.

And the one that gets almost everybody

These models tell everybody they have a great idea.

That is a product feature, not a verdict. It was tuned to be agreeable, because agreeable keeps you in the chat and keeps you paying.

If it has never once told you your plan was weak, you are not getting analysis. You are getting company.

One move this week

Go find the last 10 leads that came into your business. Write down the exact minute each one arrived, and the exact minute somebody answered.

Do not automate anything yet. Just measure it.

That second sheet of paper will tell you more about what to fix than any model on any leaderboard. It costs you 20 minutes and zero dollars.

A gap I owe you

I could not reach the public forums this morning where regular people argue about this, so I do not have a read on the room the way I usually would. Everything above came from ARC Prize, the Bureau of Labor Statistics, the Texas Legislature, and the companies themselves, each one labeled as I went.

Nobody lied to you today. They just picked which true number to say out loud.

That is going to be the whole skill from here on.

AI for everyone. Not just the wealthy.

Common questions

How did the same AI score 62 percent and 99 percent on the same test?

ARC Prize ran GPT-6 Astra on ARC-AGI-3 two different ways. Under the standard harness, where the model keeps only the notes it chooses to write for itself, it scored 62.71 percent at a cost of 26,098 dollars. Under the provider adapter, where its own internal reasoning state is carried between requests through the company's own plumbing, it scored 99.95 percent at a cost of 18,817 dollars. Same model, same puzzles, same week. The distance between those numbers is scaffolding, not intelligence.

What did the human beings cost by comparison?

ARC Prize hired real people and paid them 115 dollars for a 90 minute session plus 5 dollars per completed game, which works out to roughly 12.78 dollars per attempted game. So the comparison is 12.78 dollars for the human against 18,817 dollars for the cheaper of the two machine runs. When somebody says the machine beat the human, the sentence is true and financially absurd.

Was anybody lying about the 99 percent?

No. Both numbers are real and ARC Prize published them side by side, which is what an organization does when it is not selling you anything. The company quoted the higher one. That is marketing, not fraud. The problem is that a true number can still create a false picture, because when you hear 99.9 you imagine a machine solving something on its own.

Do benchmarks actually measure how smart a model is?

A benchmark measures a machine doing one specific thing under one specific setup chosen by one specific group of people who had a reason for choosing it. Epoch AI ranked this model first overall at 169 points across more than 50 benchmarks. Artificial Analysis scored it 61, level with the model it replaced and behind Claude Fable 5.1. Four defensible numbers for one model, ranging from world beating to no change at all.

What does this mean for a small business owner?

Stop shopping for the smartest model. The harness gap just proved that the smartest model wired up wrong scores 62 instead of 99. Your setup matters more than your model. Most of what looks like an AI problem in a small business is a leak, and a leak needs a rule or a switch, not a brain.

What is the one move this week?

Find the last 10 leads that came into your business. Write down the exact minute each one arrived and the exact minute somebody answered. Do not automate anything yet, just measure it. That second sheet of paper will tell you more about what to fix than any model on any leaderboard, and it costs 20 minutes and zero dollars.

Want this working in your business?

Connor builds the AI systems he writes about, here in Santa Clarita. Book a working session and bring your actual workflow.

Get on Connor's Calendar

Connor T. MacIvor · CalDRE #01238257 · Sync Brokerage, Inc. · DRE #02031490