It's just probability.
Yes. And so is the brain you are comparing it to. Most of the ways a language model is "just guessing," you are guessing too.
You have heard the objection: we cannot trust a language model, because all it does is pick the most likely next word. It does not know anything; it guesses. A person, the argument assumes, is a different kind of thing.
Half of that is right. A language model really is, underneath, a machine for guessing what comes next. But the other half, that this is what makes it different from you, is where it falls apart. Because the brain is a guessing machine too. Nearly every way the model is "just probability" turns out to be a way that you are.
This is not the claim that models are as reliable as people. It is a smaller and sharper one: "it's just probability" is not the difference between them. It is what they have in common.
Your brain guesses the next word before it arrives
The leading picture of the brain in neuroscience is not a camera that records the world. It is a machine that constantly predicts what is about to happen and corrects itself when it is wrong. When scientists record a person's brain while they listen to speech and compare it to a language model, they find the brain doing the same thing the model does: guessing the next word before it comes, and reacting more strongly the more surprising it turns out to be1. It is not settled, and it is not identical. But the direction is clear: your language system is running on prediction, not on looking things up in a dictionary.
Your certainty is a feeling, not a fact
It feels as if a belief is a solid thing: you either know it or you don't. But that is not how a decision is made inside you. When you judge something true or false, evidence piles up, noisily, until it crosses a line, and then you feel sure2. Nudge the noise or move the line and the same evidence gives a different answer. What you feel as knowing is really a race between options that happened to finish where it did. A model's list of odds for its next word is not a poor imitation of your certainty. It is the same thing, just out in the open where you can read it.
You make things up too
The failure people fear most in a model is the confident lie: the fluent, detailed answer that was simply never true. It gets called the model's signature flaw. It is also exactly how your memory works. You do not play back a recording. You rebuild a likely story, and the story bends to the question you were asked. Ask people how fast cars were going when they "smashed" instead of "hit," and they remember more speed, and broken glass that was never there3. Every time you remember something, you rewrite it a little4. A memory is a likely reconstruction told with confidence. When a model does the same, we call it a hallucination.
So what actually separates them?
None of this makes a model trustworthy. It just means "it's only probability" is the wrong reason to distrust it. If you want to say a particular model is less reliable than a particular person on a particular job, the real questions are ones you can actually check:
- Does its confidence match its accuracy? When it sounds sure, is it right?
- What do its mistakes look like? Random and easy to catch, or confident and all leaning the same way?
- Can you check the answer against something true?
- Is there someone there over time who carries the consequences, or a fresh stranger every time you ask?
These are the differences that matter. And on the first three, the machine often comes off better, because you can read its odds directly and score how well-calibrated it is, while a person's confidence is a closed box you can only guess at from the outside.
Many fear we overrate language models. The bigger, better-supported worry, the one nobody says out loud, is that we overrate humans.
The one real difference on that list is the last: a model, as we run it today, forgets everything between conversations, so there is no continuous "someone" building up over time. A brain has that. But that is about what we have built so far, not a wall the machine can never cross. Give it a lasting memory and the gap starts to close, and whether that produces a real self or just a good memory is an open question, and an honest one. It is a prediction, not a finished result. Naming it is the point: an argument that hides its weak spot is easier to knock down than one that shows you where it is.
A few of the sources
The full argument lists all twenty-one. These are the load-bearing few.
- Goldstein, A., et al. (2022). Shared computational principles for language processing in humans and deep language models. Nature Neuroscience, 25(3), 369–380. 10.1038/s41593-022-01026-4
- Ratcliff, R. (1978). A theory of memory retrieval. Psychological Review, 85(2), 59–108. 10.1037/0033-295X.85.2.59
- Loftus, E. F., & Palmer, J. C. (1974). Reconstruction of automobile destruction. Journal of Verbal Learning and Verbal Behavior, 13(5), 585–589. 10.1016/S0022-5371(74)80011-3
- Nader, K., Schafe, G. E., & LeDoux, J. E. (2000). Fear memories require protein synthesis in the amygdala for reconsolidation after retrieval. Nature, 406, 722–726. 10.1038/35021052