Large language models are remarkable at generating plausible prose, summarizing documents, and translating between languages. They are terrible at counting the letter 'r' in 'strawberry.' This is not a flaw to be patched in the next release. It is a fundamental consequence of how these systems process language, and grasping this distinction separates informed users from disappointed ones.
The confusion stems from a category error. When humans count letters, we examine each character sequentially, maintaining a running tally. When a language model encounters the word 'strawberry,' it does not see individual letters at all. It sees tokens — chunks of text that the model has learned to recognize as meaningful units. The word might be split into 'straw' and 'berry,' or processed as a single token entirely. The model never examines the raw characters because it was never designed to.
The tokenization wall
Tokenization is the preprocessing step that converts text into numerical representations the model can process. Different models use different tokenization schemes, but all share a common trait: they optimize for predicting the next likely token, not for character-level analysis. A model trained on billions of words develops sophisticated intuitions about which words follow which, but it has no mechanism for inspecting the internal structure of those words. Asking it to count letters is like asking a chess grandmaster to count the threads in the carpet — technically possible, but requiring a completely different cognitive apparatus.
This explains why language models struggle with precise arithmetic, reliable citation, and tasks requiring exact recall. They are statistical pattern-matchers of extraordinary sophistication, not databases or calculators. When a model produces a correct numerical answer, it is typically because similar problems appeared frequently enough in training data that the correct response became statistically likely. It is not performing computation in any meaningful sense.
What the architecture actually does well
The same architecture that fails at counting excels at tasks humans find laborious. Summarizing a dense legal document, identifying the emotional register of a customer complaint, generating plausible dialogue in a specified style — these leverage the model's core competency of predicting statistically appropriate continuations. The model has absorbed enough examples of good summaries, appropriate emotional responses, and stylistic variations that it can generate convincing new instances.
This is genuinely useful. A tool that can draft a first version of nearly any text, identify likely errors in code, or explain complex concepts in accessible language has obvious applications. The mistake is expecting it to also be a reliable calculator, a fact-checker, or a system that 'knows' things in the way humans know things. It knows nothing. It predicts.
Our take
The AI industry has a marketing problem masquerading as a technical one. Calling these systems 'intelligent' invites comparisons to human cognition that the architecture cannot support. A more honest framing would acknowledge that language models are extraordinarily powerful autocomplete engines — useful precisely because they have absorbed the statistical regularities of human text at a scale no person could match. The disappointment users feel when their AI cannot count letters or remember what it said three prompts ago is not the model's failure. It is the failure of an industry that preferred mystification to explanation.



