Large language models produce Shakespeare-quality sonnets on demand, debug complex code, and synthesize research papers across disciplines. Ask one how many times the letter 'r' appears in 'strawberry' and watch it confidently announce two. The machine that seems to know everything cannot perform a task mastered by seven-year-olds.
This is not a software glitch awaiting a patch. It is a direct consequence of how these systems process language at the most fundamental level — and understanding why illuminates both the genuine brilliance and the hard limits of the technology reshaping industries worldwide.
The tokenization problem
Before a language model reads your prompt, it breaks text into tokens — chunks that might be whole words, word fragments, or individual characters depending on the vocabulary it learned during training. The word 'strawberry' might become 'straw' and 'berry', or 'str', 'aw', and 'berry', depending on the tokenizer. The model never sees the individual letters you see. It processes abstract numerical representations of these chunks.
When you ask about letter frequency, you are asking the model to reason about a level of granularity that exists below its perceptual floor. Imagine being asked to count the atoms in a photograph you can only view at billboard resolution. You might make educated guesses based on what photographs typically contain, but you cannot actually see what you are being asked to count.
This tokenization scheme exists for good reasons. Processing text character-by-character would make training vastly more expensive and context windows impossibly short. The chunking approach lets models handle longer documents and learn higher-level patterns. But it creates a permanent blind spot.
Fluency without foundation
The counting failure points to something deeper about how language models generate text. They predict the next plausible token based on statistical patterns learned from training data. When asked to count letters, they do not execute a counting algorithm — they generate text that resembles what a correct answer would look like based on similar questions they encountered during training.
This is why the errors are so confident. The model has seen countless instances of questions about letter frequency followed by numerical answers. It knows the format of a correct response perfectly. It simply has no mechanism to produce the actual count. The fluency is real; the reasoning is simulated.
The same dynamic explains why language models struggle with precise arithmetic, calendar calculations, and spatial reasoning. These tasks require operations on structured data that the model only encounters as serialized text. Multiplying large numbers demands carrying digits in sequence — but the model generates all digits in parallel, predicting each based on what multiplication answers typically look like rather than computing the result.
The memorization workaround
Models often get simple calculations right because they have memorized common results. Ask for 7 × 8 and the answer appears instantly — not computed, but recalled from the vast corpus of text where this multiplication appeared. Ask for 7,847 × 9,263 and performance collapses because this specific product rarely appears in training data.
This creates an uncanny valley of competence. The model handles some quantitative questions flawlessly and others disastrously, with no obvious pattern to a user who does not understand the underlying mechanism. A student might trust its arithmetic on homework problems that happen to match memorized examples while receiving confident nonsense on novel calculations.
Our take
The letter-counting failure is not an embarrassment to be fixed but a window into what these systems actually are: extraordinarily sophisticated pattern-completion engines that produce human-like text without human-like cognition. The gap between generating plausible prose and performing reliable reasoning remains vast. Anyone deploying these tools in consequential contexts should understand that fluency and accuracy are entirely separate properties — and that the most dangerous errors come wrapped in the most confident language.



