You remember the work, not the material
Paper 4B · Pødenphant Lund (2026) · Read on Zenodo
Why does help that serves a novice harm an expert?The expertise reversal effect is one of the best-established findings in instructional psychology, but no one has been able to say what happens inside. The paper proposes a mechanism and tests it in language models, where it can be read word by word. The title’s claim, that what is encoded is experience and not information, is the frame for all of it: what sticks is the work that had to be done with the material, not the material itself.
The job that started it
I was asked to build training material on malnutrition in elder care. The brief was to teach care staff the nine clinical signs of malnutrition and every illness that can lead to it. In short: a great deal of information.
It was hard. However we turned the material, there was a lot of content, and none of it told anyone what to do. Care staff already get information from many directions, and nine more things to remember would not help.
The breakthrough came when we stopped asking “what should they know?” and started asking “what should they do?” The guidelines already said staff should offer monthly weighing. It simply was not done consistently, because responsibility was diffuse and the action was not built into the routine.
The new material was one sentence: “Remember to weigh, it is good care. If you see weight loss of more than a kilo, act.” A campaign was built around that sentence, and the nine signs and the illnesses went on the back, there when needed. Training went from 15 to 20 minutes of information to about three minutes of action.
Completeness is a property of the sender. Learnability is a property of the receiver. They are not on the same axis.
The idea is not new, and that comes first
That people remember what they do with material rather than what they are shown is established. Craik and Lockhart showed in 1972 that what is retained depends on the processing performed, not on the information presented. Morris, Bransford and Franks added in 1977 that recall succeeds to the extent that the processing at test resembles the processing at learning. The paper contests neither.
What it adds is a proposal for what exactly is encoded. Not how deeply something is analysed, but the friction that arises when several possible routes compete and have to be resolved. And the two can be told apart: information that opens no race leaves no trace, however thoroughly it is worked through. The word “experience” in the title is shorthand for exactly that. It carries no claim about feeling or awareness.
Under that mechanism, Nick Shackleton-Jones’s theory that we remember what mattered to us emotionally takes a particular place: emotional intensity becomes one visible face of high-friction processing. What feels significant is what cost something to settle. That is a proposed connection, and the paper does not test it in people.
The mechanism
A system that can hold several strategies for a problem opens a competition between them whenever new information does not fit the strategy it has already committed to. The competition costs for as long as it stays unresolved, and the system closest to the edge of its competence pays most.
That explains why the same help works in opposite directions. A novice has no strategy, so an example shows the way. An expert already has one, so an example of something the expert can already do opens a rival to the strategy that was working.
What the experiments showed
- A correct example can harm the capable model. On a task where two given pieces of information have to be multiplied together, one example of a procedure the model already masters left the smallest model unaffected, helped the mid-sized one and harmed the largest. Llama-3.3-70B fell from 73 to 50 per cent with one example and recovered to 61 with three, and friction peaked exactly at the switch. That larger models are more sensitive to what they are given in context is known from Wei and colleagues and Shi and colleagues. What is new is that the harm comes from a correct example that contradicts nothing. That is the mark of expertise reversal, not of a clash with something the model already believed.
- The pattern is not a law. A second, structurally different task showed correct examples harming again, but on a different model and with a different trace in the words. Which size of model is hit depends on the task.
- Clarity and length are not the same. An elaborated example that showed the route to the answer gave 16 percentage points more, at lower friction, than a short one showing only the answer. The naive expectation ran the other way, that more content means more load. Friction tracks whether the race gets closed, not how many words there are.
- Quantity is not what costs. 1,600 tokens of meaningless filler cost less than nine tokens of plausible elaboration. The filler opens no race and can simply be ignored. That shows the quantity of material presented is not what encodes. It is not yet a test of friction against depth of analysis.
- Conflicting instructions cost. When the system message asked for one format and the example showed another, accuracy fell to 48 per cent against 70 with a matching example and 78 with none. The result is carried by accuracy. The friction signal pointed the same way but was inconclusive at 50 items.
Clarity and help are two different things
This is the result that belongs with the weighing story. If redundant help to someone who has already decided opens a race and does harm, what happens when you instead remove an ambiguity, so the receiver can decide at all?
The paper built pairs of texts describing the same action either clearly or buried, matched to be equally easy to read. Clarity sharply lowered uncertainty about what to do, while readability carried nothing on its own. On 17 real before-and-after rewrites from a plain-language archive, the after version made the action clearer in 15 cases, and in some of them the text actually became harder to read while the action became clearer.
And the most important part: the benefit of clarity did not vanish with capability. Even where the largest model hit the ceiling for correct answers, uncertainty about the action still dropped with a clear text. An unclear action leaves the race open, whatever the receiver knows.
So redundant help and unclarity are not two points on one scale. They are two signs of the same mechanism, and the sign is set by whether the race was open or closed before the help arrived. The common assumption that a capable audience does not need clarity holds for redundant help and fails for unclarity. “Remember to weigh” closed a race that had been left open.
What it does not claim
This is measured in language models, and the models are used as a model system for a mechanism whose confirmation in people is still to come. Whether the same split between clarity and help holds for people is a prediction the paper passes on to Paper 16, not a result.
The clarity measurement has its own limits. The set of hard items is small, only eight, and carries the direction rather than the size. It is the 17 real rewrites that carry the ecological weight. And part of the drop in uncertainty could come from a clear text binding the action verb more tightly, a possibility not fully ruled out.
Finally, a language model with no memory between conversations lacks the layer that lets the trace build up over time. The mechanism within a single piece of processing can be compared, but how long the trace lasts differs from one system to another.
Related papers
- Paper 16 (The physics of learning) — where expertise reversal, cognitive load and desirable difficulties are brought together, and where the prediction about people belongs.
- Paper 20 (Compliance is behaviour) — the same shift from “what should they know” to “what should they do”, applied to rules.
- Paper 0 (BFT) — the behavioural framework behind the notion of friction.