Skip to main content
🎓 Claude Code Masterclass Learn AI-assisted development on Udemy — plus the companion book on Leanpub & Amazon. Start Learning
Ballard's Test, and why today's AI still fails it
AI

Ballard's Test, and Why Today's AI Still Fails It

Ballard's Test: reason without language proves understanding. AlphaGeometry and Coconut rediscover a 144-year-old Smithsonian idea — LLMs aren't reason engines.

LB
Luca Berton
· 5 min read

It is worth reintroducing Ballard’s Test right now, because the conversation around artificial intelligence keeps confusing the two things Ballard’s Test was built to separate.

An entity does not possess capacity for understanding until reason is demonstrated in the absence of language.

Read that twice. The test is not “can it talk.” The test is “can it reason when you take the talking away.” If all you have ever observed is the talking, you have not yet observed understanding. You have observed fluency.

What Ballard’s Test actually separates

Most of what passes for “AI is smart now” in 2026 is a fluency observation. A large language model produces coherent, confident, grammatically flawless text. That is genuinely impressive as a statistical achievement. But fluency is not the evidence Ballard’s Test asks for. The evidence is reason that survives the removal of language — a system that arrives at a correct inference it was never trained to narrate.

We have not collected much of that evidence. What we have collected is a great deal of fluent narration about inferences the system did not actually perform.

The 144-year head start

Here is the part that ought humble everyone in the room. The principle that language is the output of understanding, not the engine of reason, was presented at the Smithsonian 144 years ago. We are not inventing this distinction. We are rediscovering it, slowly, with GPUs.

Every few months a lab announces a breakthrough that “shows models can reason,” and the announcement quietly confirms Ballard’s Test rather than defeating it: the moment they want real reasoning, they stop relying on the language model alone and bolt on something that reasons outside the language channel. The frontier is converging on the Smithsonian, not leaving it.

The frontier is finally going reason-first

Two recent efforts make the rediscovery explicit.

Google DeepMind’s AlphaGeometry is a neuro-symbolic system. A neural language model proposes useful geometric constructs — where to add a point, a line, a circle — and a symbolic deduction engine proves the theorem step by step. The reasoning that matters does not happen in the language model. It happens in the formal engine the language model is merely steering. The language model is the intuition; the symbolic engine is the proof. That is Ballard’s Test being respected by architecture: the reason is demonstrated where language is not doing the work.

Meta’s Coconut — Chain of Continuous Thought — goes further and says so out loud. Instead of generating reasoning as text tokens, it keeps the reasoning in the model’s continuous latent space, using the last hidden state as the “thought.” The authors describe this as freeing reasoning from the language space. A research team at one of the largest AI labs looked at chain-of-thought prompting and decided the next step was to remove language from the reasoning loop. That is not a coincidence. That is the same insight, arrived at from the engineering side: language is the output, not the engine.

So when people say the frontier is “moving toward reasoning,” the precise and more interesting claim is that the frontier is moving away from language as the substrate of reasoning. That is exactly what Ballard’s Test predicts a serious attempt at understanding would have to do.

So — are we at General Artificial Intelligence?

Maybe. But we are not there yet. Not close. And there is no guarantee we ever get there.

I mean that as a statement of probability, not pessimism. The reason-first architectures are early, narrow, and brittle outside their training envelope. A system that can prove Olympiad geometry under tight constraints is not a system that understands the world. It is a system that is very good at one sliver of inference and silent everywhere else.

And if we do get close — if some future system genuinely passes Ballard’s Test across domains — I would not be shocked if a steely being from the future showed up to halt further progress. I am only half joking. The people building this are not stupid, and they are not unaware of what they are building. The pause button, if it is ever pressed, will not be pressed by accident.

Why an LLM could not write this essay

Most people who read my essays — Palantir bulls excepted — come away saying, “AI couldn’t have written this.” They are right, and not because I write better sentences. They are right because the content is not retrievable.

I apply my own creativity to capital markets. The connections in these pieces are original to me. An LLM cannot search for an idea that does not exist in its training data, and the idea does not exist until I read a 144-year-old paper on a Saturday morning and decide it explains something nobody had connected it to before. The retrieval engine has nothing to retrieve, because the synthesis was never published. Novel combination is not a search problem. It is an authorship problem.

That Saturday-morning move is the whole point. Understanding shows up as a link between two things that were not previously linked. You cannot prompt a system to make a link it has no reason to suspect exists, because the existence of the link is the reason. Until I made it, there was no token for it anywhere.

A note to the thinking adults

I am not writing this to sell you on AI, and I am not writing it to scare you off AI either. I take the subject seriously enough to tell you exactly where I think the line is.

Large language models are not reason engines. They are the most fluent autocomplete the world has ever built, and fluency is a real and useful thing. But fluency narrating a reasoning process is not the same as a reasoning process. The labs that want real reasoning know this — that is why AlphaGeometry and Coconut exist. The rest of us should hold the line too: praise the fluency, demand the proof, and do not hand the word “understanding” to a system that has never passed Ballard’s Test.

Before we call anything intelligent, let it reason without the words. Most of what we have built still cannot.

#ai #ai-infrastructure #Production AI #LLM #reasoning
Share:
AI Integration & GPU Platforms

Need help with AI Integration & GPU Platforms?

Need help deploying AI/ML platforms? Get expert consulting on OpenShift AI, GPU orchestration, and MLOps.

Learn more about AI Integration & GPU Platforms

Want to operate this yourself, in production?

Take the free AI Platform Engineer Readiness Scorecard to see which skills transfer — then build a production-shaped AI platform in the 4-week Bootcamp.

Take the Scorecard →
Luca Berton — The Production AI Expert, Docker Captain

Luca Berton

The Production AI Expert · Docker Captain · KubeCon Speaker

15+ years in enterprise infrastructure. Author of 8 technical books, creator of Ansible Pilot (1M+ YouTube views, 648K site users). Former Red Hat engineer. Speaker at KubeCon EU 2026 and Red Hat Summit 2026.

Free 30-min Production AI consultation

Book Now