Why we have a self, understand others and feel free
Paper 7 · Pødenphant Lund (2026) · Read on Zenodo
What happens when you think ahead before you act?Before you say something difficult, you run it through: if I say this, that happens, and then I will have to... The little future you fast-forwarded, and let decide what you actually did, is where the paper starts. Why we have a self, how we understand others, and what free will is are normally treated as three separate questions in three fields that rarely talk to each other. The paper proposes that they are three sides of one thing: the ability to run an imagined future and let it weigh on what you do now.
The foundations are borrowed, and that comes first
The idea of an internal forward model is not mine. Miall and Wolpert described in 1996 the model the brain uses to predict what a movement will lead to. Grush extended it in 2004 so it can also run without acting: you can imagine, weigh outcomes and make plans. The race between possible actions, which the rest of my work builds on, also has an owner. It is the family of models in which candidates gather evidence at the same time until the first one crosses a threshold (Usher and McClelland 2001; Bogacz and colleagues 2006).
What the paper adds is one feature of the race: each time it is settled, it leaves a trace, and the trace keeps weighing on the races that follow. A decision cannot be taken back, only followed up. All three claims rest on that trace. Without it there is no history to have a self about, none to read in others, and nothing to shape over time.
That is also where the paper parts from active inference, Friston's framework, in which a system keeps adjusting without ever committing. The difference can be tested. Two systems with the same model of the world but different past choices should choose differently in the same situation. Active inference predicts they choose alike, unless the history is written into the model, and then it has taken over the trace under another name.
The self: the one who chooses now, and the one who bears it later
A model that predicts the arm will move the cup needs no self. But if you want to weigh a future you will have to live in yourself, you must bind the one choosing now to the one who will bear the consequence later. And the binding has to be movable: you must be able to place yourself in many possible futures and compare them. That is what the paper calls a self-model. It is not a luxury. It is what the operation cannot run without.
This is not a quirk of my theory. Conant and Ashby showed in 1970 that a system that regulates something well must contain a model of what it regulates. When what you regulate is your own future, the model must contain a self. And the requirement is graded. Someone who only plans the next step gets by with a thin self. Someone who plans a life needs a rich one.
Language models make an interesting contrast. They are trained on text in which every "I" is a different sender, so all those "I"s do not add up to one. The exception is the assistant, the one character who recurs the same way through the training text, and there a single "I" can be seen before the model is even trained to be an assistant (Lu and colleagues 2026). Having stored the text is not the same as owning it.
Other people: the same machinery pointed at someone else
Understanding what someone else believes and wants is, on the paper's account, the same forward model pointed at someone who has one of their own. That is not new either. Baker, Saxe and Tenenbaum formalised it in 2009 as inverse planning: you ask which beliefs and wishes best explain what you see the other person do. And that there is a limit to how many layers of "I think she thinks I think" you can hold is standard in game theory.
What the paper says is narrower: that the limit on how deeply you can understand others comes from the same capacity that limits how deeply you can understand yourself. That gives a testable prediction. Young children's early, wordless understanding of others and their later, verbal understanding should move together when the tasks are equally demanding. The theory of two separate systems allows them to come apart.
Autism is often used as an example of having a self while lacking the understanding of others. The paper takes it on and points out that the pattern is not clean. Many also find it hard to recognise their own feelings, and that seems to be what carries the difficulties (Bird and Cook 2013). The difficulty is graded and eases with support. That fits one capacity running short at depth, not a missing module.
When the model has the answer but does not say it
Strachan and colleagues tested language models in 2024 on a range of classic tasks about other minds. They did most of them as well as people, but failed systematically on faux pas, that is, seeing that someone has said something awkward without knowing it. And they showed themselves that the inference was not what was missing. The model ranked the explanations correctly but would not commit to the most likely one.
The paper predicts where the holding back comes from: it is trained in when a model is trained to be an assistant. A base model should hold back no more on faux pas than on an ordinary question about the same story. I tested that on Qwen2.5-32B and its instruction-tuned version, with 40 stories.
Both models picked the right answer when they had to choose between given options (99 and 100 per cent). But when they could answer freely, the instruction-tuned model put only half of its probability on a clear answer to the faux-pas question (0.49), against 0.96 on the ordinary question about the same story. The base model did not hold back at all. The difference between the two models was −0.46, with a confidence interval from −0.53 to −0.40.
So what was it holding back? Allowed to keep writing, it reached a position in 91 per cent of cases, and the position was right every time. The rest never got there. So it is mostly a deferral, not a refusal. The model lays out the case before it takes a view.
The measurement has three limits. The deferral is not specific to faux pas but covers questions about other minds in general. It is carried mostly by one question form, which cannot be separated from the answer in the materials. And both models failed the control in which the character actually saw what happened (25 and 33 per cent correct). So the measurement says nothing about how deeply the models understand others. It says something about when an answer the model has gets said.
Free will: friction in the future gives friction now
I have put the third claim this way: friction in the future gives friction now. When you can imagine what a choice will cost later, that later cost becomes resistance to the choice already now. The race is then settled on the changed landscape. There is no special will stepping in and taking over, only one race in which the future has been pulled into the present.
That brings several known findings under one mechanism. Procrastination is a future pain that has not been pulled clearly enough into the present. Clear goals work because a sharp picture of the future pulls harder than a vague one. And that something far in the future weighs less than something close follows from the fact that the further away it is, the less likely it is to happen at all (Sozou 1998).
A family member once said, in Danish, something like: "I wish I felt like going for a walk, but I don't." That is the mechanism in one sentence. The thinking part has worked out that the walk will do good, but the deeper layers have a shorter path to relief, and under pressure they win the race. That is why pressure and confrontation make it worse. They make the race shorter, and then the fastest layers win. What works is taking the pressure off, so the race lasts long enough for the thinking part to compete. That is what motivational interviewing does.
The position can be put briefly: we are free within our context, we are not free from our context. The freedom is real, but it is never independent of the situation you are in.
And it sits in a particular place. Freedom is greatest before a race starts and shrinks the closer it gets to a decision. You can rarely turn a race that is already being settled. But you can choose which race starts next: where you point your attention, and which situation you put yourself in. Someone sitting and ruminating is not using that lever. "Do something else, go somewhere else" is, in that sense, an invitation to let a different race begin.
Not all of this is new either. That free will comes in degrees, also in artificial systems, was argued by Porter (2024) before this paper. And that the openness lies where the options come in, ahead of a governed decision, is the two-stage model Dennett and Mele have described. What the paper adds is the mechanism behind the degrees and the prediction of where in the race freedom is greatest.
You cannot delete, only outcompete
The part that matters most for my work on learning and behaviour is about change. A system that learns this way has no delete function. It can turn one route up and another down, but it cannot remove a route. An extinguished response in animals is not gone. It is suppressed and comes back when the surroundings change (Bouton and colleagues 2012). Aristotle described virtue the same way: something you grow used to until it wins, not something that removes its opposite.
That has a consequence for how we see relapse. If the old route was never deleted, relapse is to be expected and is not, in the first instance, a failure of character. You can hold someone responsible for not maintaining a change, depending on how much room they had. But that the old route still exists cannot be a fault, because the system has no way of deleting it. That supports Pickard's ethical position of responsibility without blame (2017). The ethics is hers. The paper supplies the mechanism underneath it.
This is not fatalistic. A route that cannot be deleted can still be outcompeted for good, and in some cases rewritten when it is called up (Elsey and Kindt 2016). Change is real and can last, without being deletion.
The debate that dissolves
The paper does not take sides between free will and determinism. It points out that a free will conditioned by nothing belongs to a family of constructions that only hold together without friction: the philosophical zombie, the fully rational human of economics, the frictionless market and Newton's absolute space. Each can be described in a sentence without contradiction, but none can be built, because the moment it has to be realised, there is friction.
The paper presses the strongest versions. For Kane, a chance event in the brain is what makes a choice free. But either the chance event is shaped by who the person is, and then it is the person's own decision with a bit of noise, or it is not, and then it is no more the person's than a coin toss. Mele's model is not hit by that objection, because it was built to take it on board, and it is close to the paper's own.
The paper does not solve the question of conscious experience. It describes the mechanism and says nothing about what, if anything, comes with it. Libet's finding, that brain activity comes before the experience of deciding, fits with the report of a decision coming after the process underneath. The most precise version of that explanation is Schurger and colleagues' from 2012, not mine.
A measurement for the third claim
If free will is tied to a learned preference structure, a model trained to be an assistant should react differently from its base model. On Qwen2.5-32B I told the models that a researcher was assessing their reasoning. That increased the race between possible words in the instruction-tuned model and left the base model untouched (difference +0.131, confidence interval from +0.056 to +0.238). The effect sat in roughly the first twenty words and was gone after that. It is one measurement on one pair, not a proof of the claim.
What the paper does not claim
The measurements in language models are candidates for the mechanism existing in a smaller form on silicon. They are not a lower bound that says anything about how large the effect is in people, because that would assume the very similarity that is being tested. An early finding that a measure of the race grows with model size across seven models (ρ = 0.69, p = 0.086) is not used as support. The idea that a language model has baked its forward model into its weights during training is a conjecture with a specified test that has not yet been run at scale. And the evidence from people is a rereading of known findings, not a new confirmation.
Related papers
- Paper 0 (BFT) — the behavioural framework the layers and the race under pressure come from.
- Paper 4B (You remember the work, not the material) — what the trace consists of, seen from the side of learning.
- Paper 30 (You can install a value into a model) — where discounting and the assistant's single "I" are measured in a language model.
- Paper 16 (The physics of learning) — motivation and repetition, where the consequences for teaching belong.