Ask a large language model to write a sonnet about quantum mechanics and it will produce something serviceable in seconds. Ask it how many times the letter 'e' appears in the word 'strawberry' and it will confidently answer three. The correct answer is one. This is not a glitch awaiting a patch. It is a fundamental feature of how these systems process language, and understanding it clarifies both the remarkable capabilities and the hard limits of the technology reshaping white-collar work.
The paradox cuts to the heart of what large language models actually are: prediction engines trained on text, not reasoning machines that happen to use words. When a model encounters a counting question, it does not see individual letters. It sees tokens—chunks of text that might be whole words, fragments, or punctuation marks. The word 'strawberry' arrives as something like 'straw' + 'berry' or even stranger subdivisions depending on the tokenizer. The model has no internal representation of the letter 'r' as a discrete object to be tallied. It has statistical associations learned from billions of documents where counting was discussed, not performed.
The illusion of understanding
This architectural reality produces what researchers call the 'capability overhang'—the gap between what a model appears to understand and what it actually computes. When GPT-4 or Claude writes a legal brief, it is not applying legal reasoning in any meaningful sense. It is predicting what tokens would plausibly follow the prompt based on patterns absorbed from legal texts. The output looks like reasoning because legal writing has a distinctive statistical signature that the model has learned to reproduce with high fidelity.
The same mechanism explains why models hallucinate with such conviction. A confident wrong answer about letter counts and a confident wrong answer about historical dates emerge from the same process: pattern completion without verification. The model has no internal fact-checker because it has no internal facts—only weights representing the probability that certain tokens follow others.
What this means for real-world use
The counting failure is a diagnostic tool. Any task that requires precise manipulation of discrete elements—exact arithmetic, character-level string operations, step-by-step logical deduction with many dependencies—will push against the same architectural constraint. Models can approximate these tasks when the answers appear frequently in training data, but they cannot reliably perform them from first principles.
This explains the pattern of AI adoption in professional settings. Models excel at drafting, summarizing, translating, and brainstorming—tasks where approximate pattern-matching produces useful output. They struggle with auditing, debugging, and verification—tasks where a single discrete error invalidates the whole. The most effective deployments use models for generation and humans for validation, a division of labor that respects the genuine asymmetry in capabilities.
Our take
The counting problem is not an embarrassing limitation to be minimized in marketing materials. It is an invitation to precision about what we have built. Large language models are the most sophisticated pattern-completion systems ever created, and pattern completion turns out to be far more powerful than anyone anticipated. But they are not general reasoners, and pretending otherwise leads to deployment disasters. The firms and individuals who thrive with this technology will be those who understand its actual nature: genuinely useful, genuinely limited, and genuinely unlike anything that came before.




