Large language models can draft legal contracts, explain quantum mechanics to children, and produce serviceable sonnets in the style of Shakespeare. Ask one to count the windows in a photograph of a building, and it will confidently report eight when there are twelve. This is not a temporary limitation awaiting the next parameter increase. It is a window into what these systems fundamentally are — and are not.

The disconnect feels almost comical. A model that can discuss the philosophical implications of Gödel's incompleteness theorems struggles to determine whether a picture contains more apples than oranges. Users encounter this daily: upload an image of a chessboard mid-game and ask how many pieces remain, and the response will be articulate, confident, and wrong. The failure mode is consistent enough to constitute a signature.

The architecture of eloquence

Language models learn by predicting the next token in a sequence, trained on vast corpora of human text. They become extraordinarily skilled at pattern completion — recognizing that certain words follow others, that arguments have structures, that style has fingerprints. This is genuine capability, not mere mimicry. The models develop internal representations that capture something real about language and reasoning.

But counting objects in an image requires something different. It demands maintaining a running tally while systematically attending to each item exactly once, avoiding double-counting, handling occlusion, and distinguishing foreground from background. This is not pattern completion. It is procedural execution with working memory constraints that current architectures handle poorly.

Why more parameters will not solve this

The intuition that scaling will eventually fix everything has proven remarkably durable despite mounting counterevidence. Larger models do not count better; they simply express their incorrect counts with greater confidence. The issue is architectural, not computational. Transformers process information in parallel, which makes them fast but poorly suited to inherently sequential tasks like enumeration.

Human children learn to count through explicit instruction and practice — pointing at objects one by one, reciting numbers in order, developing the concept of cardinality. This is not absorbed passively from exposure to text describing counting. It is a skill built through embodied interaction with the world, refined through correction and repetition.

What this reveals about intelligence

The counting problem illuminates a broader truth: intelligence is not a single dimension that improves uniformly. The systems we have built are savants — spectacular in their domain, brittle outside it. They excel at tasks that can be framed as sophisticated pattern matching over learned distributions. They struggle with tasks requiring systematic procedures, spatial reasoning, or maintaining precise state over multiple steps.

This is neither a criticism nor a dismissal. Pattern matching over learned distributions turns out to be extraordinarily useful — useful enough to transform industries and augment human capability in countless domains. But it is a specific kind of intelligence, not intelligence entire.

Our take

The counting limitation is clarifying rather than damning. It reminds us that we have built something genuinely new — not a digital human, not a general reasoner, but a pattern engine of unprecedented sophistication. Understanding what these systems actually are, rather than what we imagine them to be, is the prerequisite for using them well. The poetry is real. The counting errors are also real. Both facts matter.