The Physics of Learning

Paper 16 · Pødenphant Lund (2026) · Read on Zenodo

I study language models to understand people.Sweller, Bjork and Edmondson built three of the most important theories of learning, but the three literatures almost never talk to each other. Each is about its own thing: how much working memory can hold, why effortful retrieval makes things stick, and why you learn less when you do not feel safe. This paper argues that the three describe the same underlying mechanism from three different angles. If you teach, design courses, or have to explain something complicated to another person, it is the same physics you are working with every time.

Three traditions, one mechanism

If you read learning research, you run into three big literatures that have mostly grown up in isolation:

Each of the three is empirically excellent. And each is mechanically thin: it tells you what happens very well, but not why it happens at the level where the learning actually takes place. Cognitive load theory says working memory has limits, but not why those limits exist or why they have exactly the shape they do. Desirable difficulties says effortful retrieval helps, but not what actually happens during retrieval that makes it stick. Psychological safety says safety beats content, but not what mechanism makes safety the thing that comes first.

The paper argues that all three are local consequences of the same constraint. It calls it a limited-capacity race architecture, and that is worth unpacking, because it is the idea the whole of the rest hangs on. Picture the system that does the learning (a brain, a neural network) as always having several possible answers in play at once. They compete, and the system has to pick one. But it can only hold a limited number in play at a time, because it costs resources. That is what is meant by a "race" under limited capacity. The things the three traditions describe fall out as consequences of exactly that constraint. The claim is a conditional one, and the paper is precise about it: it holds for systems that actually run such a race, and it is not a law about everything that learns. The paper reads the race architecture as a resource-rational task-analysis in Lieder and Griffiths' sense, a proposal about what a system with limited resources should do and what it pays along the way.

Four ideas, and what they explain

To get from "there is a constraint" to concrete predictions, the paper uses four ideas:

Out of those four ideas, five classic findings from learning research fall out almost on their own:

Why language models

Here is why I work with language models. You cannot look inside a brain while it learns, not without anaesthetic, and even then you do not see the single choice being made. A language model you can look straight into. It is a mechanical mirror, where the constraints I have described lie open. You can watch the competition between possible answers word by word. You can see that when you push the capacity hard enough, the system's ability to learn collapses suddenly rather than gradually. And you can see the difference between dense material (many answers in play, high load) and thinned-out material (few answers in play, low load).

One measurement can stand for the rest, because it is the paper's Figure 1. Three models of different sizes were given the same chemistry task, with 0, 1, 3 and more worked examples in front of them. The smallest ran flat: it sits below the capacity at which a race between strategies can run at all, so examples made no difference either way. The middle one behaved the way the textbook says, with examples helping and the gain tapering off. The largest, a 70-billion-parameter model, dipped: 73 per cent correct with no examples, 52 per cent with one, 61 per cent with three. The model that best knew how to solve the task was the one a single example disrupted most.

And you can see why, because the competition between possible next words was measured directly in the same runs, 100 answers per condition. It peaked exactly where accuracy dipped: 1.114 competing routes on average at one example, against 1.052 with none and 1.073 at three. A single example reopens a strategy race the model had already closed, and the extra competition is paid as friction. An independent rerun on a second provider hit the same peak to within 0.007.

The language models are a mirror, not the load-bearing evidence. The load-bearing evidence is the human literature itself, and it is strong: half a century of findings on cognitive load, retrieval practice, and psychological safety. You can take the account here on that ground alone. The language models show the mechanism at a resolution you cannot otherwise get, and what they transfer is where in a stretch the friction shows up, and what opens and closes a race. The size of an effect does not transfer, and the numbers above belong to the model they were measured on.

Three ways teaching fails

The same single constraint gives three different ways teaching can go wrong. You will recognise all of them:

Each has its own fix. Too much calls for cutting down. Too thin calls for concentrating the material so the competition gets going. Never-a-decision calls for making some choices for the learner, so the field of possibilities closes.

The principle of matched resistance from Paper 6 shows up here in a variant: do not explain too thoroughly. When you explain everything, you remove the work the learner was meant to do, and it is the work that does the learning.

Practical implications

What would knock it down

An idea is only worth something if it can be wrong. So the paper says plainly what would knock it down, and it is honest about which of the four conditions are the sharp ones:

  1. A system that learns turns out not to run parallel routes under bounded resources with a choice that cannot be taken back. That would not knock the account down but narrow where it applies, and the paper counts this one as the soft one of the four.
  2. You scan a person mid-learning and find no difference between the moments where the choice lands and the moments where it does not. The paper puts the number on it: pupillometry, EEG or fMRI, at least 60 participants, a difference of at least 0.3 in effect size, replicated in an independent sample. If it is not there, the prediction about friction at the commit moment is withdrawn.
  3. Plain repetition of the same form sticks as well as varied repetition. Four groups, at least 80 per cell, measured on recall and transfer at 7 and 28 days. If the two land within 0.1 of each other at 28 days, the prediction falls.
  4. The window where the resistance is just right turns out to be symmetric around its peak. If overshooting costs the same as undershooting, the claim that a trace laid on the wrong route makes the upper side dearer falls.

The last two are the sharp ones. If both fail in pre-registered studies with enough participants, it is not one point that needs correcting but the account as a whole that is in difficulty.

Why it matters

The inverted-U is not just one finding. The paper shows that four classics share the shape for the same reason: the Yerkes-Dodson arousal curve of 1908, Vygotsky's zone of proximal development, Tulving and Thomson's encoding specificity, and expertise reversal itself. A trace forms only where the race both can close and has something new to close on, and the two walls around that region are what the U looks like from four sides. On Vygotsky the paper deliberately holds back: it accounts for why a difficulty band exists, but not for the zone's own contribution, which is what a more knowledgeable other does inside the band.

What I do not know

Here is where the line runs. The account itself stands on the human literature, and that is a solid floor: half a century of findings on cognitive load, retrieval practice, and psychological safety. What I can measure directly myself is the mechanism on a language-model substrate, where the race lies open. What has not yet been measured is the fingerprint in a person, whether the friction can be watched rising and falling in the very moment the choice is made. That measurement lies ahead, and it needs laboratories I do not sit in.

Nor do I know exactly where the window lies where the resistance is just right, because it moves with the material, with the learner, and with how far that person has already come. The paper gives a language for thinking about it, not a formula that tells you where to draw the line in your concrete situation. That is future work, and some of it can only be done together with people who measure in biological systems.

The cite

Pødenphant Lund, T. (2026). The Physics of Learning: How Race-Architecture Constraints Explain What We Know About Teaching, Communication, and Understanding. Zenodo. https://doi.org/10.5281/zenodo.20416959

Read on Zenodo → · Technical version · Dansk version

Related on this site: