The most sophisticated artificial intelligence systems ever built, capable of passing bar exams and composing symphonies, reliably fail at tasks a five-year-old masters without effort. Ask a leading language model how many letters appear in the word "strawberry" and it will confidently answer incorrectly. Request that it multiply two four-digit numbers and watch it stumble. This is not a temporary limitation awaiting the next model release — it is a fundamental consequence of how these systems process information at the most basic level.

The paradox illuminates something essential about intelligence itself. We have built machines that mimic understanding so convincingly that millions use them daily for complex professional tasks, yet these same machines lack the computational substrate for operations humans consider trivially simple. The gap between appearance and mechanism has never been wider.

The tokenization problem

Language models do not see text the way humans do. Before any processing occurs, input text is broken into tokens — chunks that might be whole words, word fragments, or individual characters depending on the tokenizer's training. The word "strawberry" might become "straw" and "berry," or "str," "aw," "ber," and "ry." The model never perceives the word as a sequence of individual letters available for counting.

This design choice was deliberate and sensible. Processing text character-by-character would be computationally prohibitive and would lose the semantic relationships between words that make language models useful. The architecture optimizes for meaning at the expense of granular symbolic manipulation. When you ask the model to count letters, you are asking it to perform surgery with mittens on — the tool was not designed for that level of precision.

Pattern matching versus computation

Arithmetic presents a related but distinct challenge. Language models can correctly answer "What is 7 times 8?" not because they compute the multiplication but because they have encountered "7 times 8 equals 56" countless times in training data. They are retrieving a memorized pattern, not performing calculation. When numbers grow large enough that the specific multiplication has never appeared in training — say, 3,847 times 2,916 — the model must either guess based on similar patterns or attempt something resembling actual computation, at which it performs poorly.

The transformer architecture underlying modern language models excels at identifying statistical relationships across vast contexts. It is optimized for the question "Given everything I have seen, what token is most likely to come next?" This is fundamentally different from the question "What is the deterministic output of this mathematical operation?" The former requires pattern recognition; the latter requires symbolic manipulation. These are different cognitive tasks, and the architecture serves one brilliantly while largely failing at the other.

What this reveals about intelligence

The limitations are instructive rather than merely embarrassing. Human intelligence is not a single general-purpose engine but a collection of specialized systems — we have distinct neural machinery for language, for spatial reasoning, for numerical cognition. Language models have developed something analogous to our linguistic faculties while lacking anything resembling our innate number sense.

This suggests that artificial general intelligence, if it arrives, will not emerge from scaling language models alone. The path forward likely requires hybrid architectures that combine the semantic fluency of transformers with symbolic reasoning systems capable of reliable computation. Some researchers are already pursuing this integration, though the engineering challenges are substantial.

Our take

The counting problem is not an indictment of language models but a useful corrective to inflated expectations. These systems are genuinely remarkable at what they were designed to do — they have transformed how humans interact with information and will continue reshaping industries. But they are tools with specific affordances and specific blind spots. Treating them as oracles capable of any cognitive task courts disappointment and, more dangerously, misplaced trust. The wise user understands that the machine writing elegant prose about quantum mechanics cannot reliably count the r's in "strawberry," and adjusts accordingly.