The Dunning-Kruger curve, seen from inside a language model

Paper 21 · Pødenphant Lund (2026) · Read on Zenodo

I was working on something else entirely, looked at a curve, and blurted out: "that's Dunning-Kruger!" I thought I had found it in a language model. It turned out others had seen it in language models before me (Ghosh and Panday described it in early 2026).The Dunning-Kruger effect is the observation that people are often most confident when they know the least, and it has been argued about for decades because nobody can see the thing that actually drives it. In a person, confidence is something you have to guess at from what they say and how well they do. In a language model you can look straight inside and watch how strongly the answers are competing before it commits. What is new here is that you can follow the curve itself from the inside, and see where it comes from.

The thing you can never see in a person

The Dunning-Kruger curve has four landmarks. A beginner starts out appropriately unsure. Then comes a steep climb to a peak of confident cluelessness, often nicknamed "Mount Stupid." Then a dip, "the valley of despair," as the learner starts to see how much they were missing. Then a slow climb back as real skill catches up.

The whole debate is about one hidden quantity: how hard a person's competing answers are fighting it out inside their head before they pick one. You can never watch that directly. You reconstruct it from confidence ratings and test scores, and critics have shown that reconstruction can manufacture the curve on its own, through statistical quirks, before any real overconfidence enters the picture.

So we changed where we looked

A language model picks each word by running a kind of race between candidates, and you can read the scoreboard. For every answer it gives, you can see how many options were live and how far ahead the winner was. That is the quantity nobody can read in a human brain, sitting right there in the numbers.

Before using it, we checked it measures what we think it measures. The gap between the top answer and the runner-up (we call it the balance of evidence) predicts whether the model is right, beats the simpler measure the field has been using, and even tells correct from incorrect answers apart when that simpler measure says they look the same. It works on knowledge questions and on a visual "is this mostly red or green?" task too. It is reading a real decision variable, not a fluke.

Then we built a Mount Stupid on purpose

To get the curve you need a place where a confident-but-wrong belief forms and then gets corrected. Picture a teacher with a hidden grading rule: a pupil's grade equals their number on the class list, except every pupil whose number is a multiple of five gets 100 added, so pupil 10 scores 110, not 10. Show the model only the easy cases (1→1, 2→2, up to 9→9) and never a multiple of five. It does exactly what a person does: it spots "number = grade" and becomes sure of it. Ask what pupil 10 scores and back comes "10," confidently, sure and wrong. That is the peak: one answer won easily because nothing was competing with it yet, and an easy win feels like confidence.

Then the truth arrives a little at a time ("actually, pupil 5 scored 105"). A second answer gets laid down, and being contradicted makes it land harder, so now two answers compete and the easy win shrinks. That is the valley. Keep going and the correct answer eventually wins cleanly. That is the climb back up.

What we found

Three parts of the curve showed up with no biology needed at all. They come straight from the way learning competes:

Two other parts behaved differently, and the difference is the interesting bit. They were not really missing. They were hidden by how the model is normally made to answer:

An application: believing you can do it

The measuring instrument reaches further than the curve. In psychology, the belief that you can carry out a particular task is called self-efficacy. Albert Bandura described it in 1977, and it is the judgement you make before you start: can I do this? Read from inside the model, it is not a self-belief filed away somewhere apart from performance. It is a readout of the same race between answers that the rest of the paper measures.

That gives two things which stay tangled in a person and lie separately here. The first: the race pulls in two directions at once, and it is the same quantity doing both. Competition between the candidate answers degrades the answer, because the harder they fight, the more often the wrong one wins. And that very same gap between winner and runner-up is what the model reads when it judges whether it can answer at all, and says "I'm not sure" instead. So the friction that degrades an answer and the signal that reports confidence in it are not two things that happen to move together. They are one quantity, read twice.

The second: the two sides can be pulled apart, and that is the part no human study can run. The degrading side is cheap. Even a small model gets worse when the candidates compete hard. The reporting side is expensive, because reading your own race well enough to decline takes capacity. The smallest model we ran declines at every stage, including where it could answer perfectly well, while models from around 70 billion parameters up get it right. A small model keeps the degrader and loses the reporter: the signal is still there in the numbers, but the system cannot act on it. You cannot take away a person's ability to read their own doubt and leave the doubt intact.

One detail is worth having, because it points at why we read the numbers rather than asking the model. When another team gave ten language models the questionnaire people fill in about their own capability, the models answered consistently from one administration to the next, but the answers did not match what they could actually do. That channel, the spoken self-estimate, is not the one used here. The quantity we read, the margin just before the model commits, does predict whether the answer comes out right, and does drive the decision to decline.

This is an application of the measuring instrument rather than a new experiment. It is the same results as above, set out under a construct psychology already uses. What is inherited from Bandura is that a capability judgement is the readout of a decision variable, and that calibration improves with capacity. What the substrate adds is that the degrader and the reporter are the same quantity, and that they come apart when capacity is small.

Why this matters

The direction is reversed here, compared with the rest of the work. Paper 0 uses the race between answers to explain how a person makes a decision. Here it runs the other way: the model’s readable numbers become the measuring instrument for the kind of bounded decision people make too. Confidence tracks how the competing answers resolve, not how good you actually are. Becoming wiser is nothing more mysterious than going from one answer that wins too easily to several answers that compete until the right one wins. The confident climb turned up in every model we tried, from the smallest to the largest, so Dunning-Kruger is not a strange human flaw. It is the shape of learning that starts from a rule that is too simple.

It also sorts the curve into two piles. The confident climb, the recognition, and the recovery are properties of the decision machinery itself. No neurons required. The felt humility and the felt despair are the part that the biological brain seems to add on top. That is a more useful answer than "models do or don't show Dunning-Kruger": it says exactly which pieces are mechanical and which are human.

The cite

Pødenphant Lund, T. (2026). Mount Stupid in the machine: how evidence competition explains the Dunning-Kruger curve in a language model. Zenodo. https://doi.org/10.5281/zenodo.20562415

Read on Zenodo → · Technical version · Dansk version

Related on this site: