Forward-modelling in Bounded Race Substrates: Agency, Self-modelling, and Theory of Mind as Substrate-Mechanical Phenomena
Paper 7 · Pødenphant Lund, T. (2026) · Preprint · v6 live on Zenodo
Self-modelling, theory of mind and free will are argued to be three manifestations of one operation: forward-modelling under a bounded race architecture, in which the substrate runs its race-evaluation on a hypothetical state and carries the prediction back into the current race as a friction-bias. Self-modelling is the data structure the operation requires; theory of mind is the operation applied recursively to another forward-modeller; free will is the operation's translation of future friction into a present friction-gradient. The empirical strategy is proposed, not licensed: language-model results are existence-proof candidates for a reduced realisation, and the human literature is reinterpreted as the full-form case.
| DOI (concept) | 10.5281/zenodo.20449154 |
| DOI (this version) | 10.5281/zenodo.22830842 (v6) |
| Status | v6 live on Zenodo, 2026-09-18 |
| Author | Tomas Pødenphant Lund [ORCID] |
Lineage, and what is added
The forward model is Miall and Wolpert’s (1996); its off-line use to imagine, estimate outcomes and evaluate plans is Grush’s (2004) emulation account, and priority for the operation is Grush’s. The race is the sequential-sampling and accumulator family (Usher & McClelland 2001; Bogacz et al. 2006). The online/amortized distinction runs parallel to the model-based/model-free contrast (Sutton & Barto 2018), and amortizing futures into a readable structure is the successor representation’s move (Dayan 1993). Theory of mind as inverse planning is Baker, Saxe & Tenenbaum (2009); bounded recursive depth is standard in level-k and cognitive-hierarchy models (Camerer, Ho & Chong 2004). A graded, cross-substrate free will was argued first by Porter (2024), and placing openness where candidates enter, ahead of a governed resolution, is the two-stage structure of Dennett and Mele.
What the paper adds: commitment-dependent hysteresis as a first-class state variable, persisting across decisions and consulted at the next one, rather than a per-trial parameter shift; the claim that the bound on other-modelling and the bound on self-modelling have a single source; and the location of action latitude at race initiation.
The three claims
Claim 1: self-modelling as a structural requirement
A thin sensorimotor forward model, mapping a command to its immediate consequence, needs no self-model, and the claim concedes this. The claim concerns the translation of future friction into a present friction-gradient where the future state is one the substrate will itself occupy. A minimal-sufficiency lemma specifies what that operation forces: an identity-binding that is diachronic (present selector to future bearer), re-deployable (placeable across many candidate futures) and unified across the evaluation (held constant so that comparisons are commensurable). Distributed egocentric pointers satisfy none of the three on their own; whatever binds them is the self-model. The derivation runs through Conant and Ashby’s (1970) good-regulator theorem, before any self-report is consulted. The requirement is graded: the depth of self-model scales with the horizon and counterfactual reach of the futures modelled (the C → T cascade).
For weight-storing substrates the paper separates the binding operation, available on demand, from a convergent first-person index. Pretraining text shatters the referent, since every “I” is a different sender; the assistant, the one referent the corpus fractures least, is the exception, with a direction already present in the pretrained model and sharpened rather than installed by post-training (Lu et al. 2026). The account would fail if the generic pretraining “I” carried a convergent index as strongly as the assistant “I”.
Claim 2: theory of mind as recursive forward-modelling
The contribution is narrow: the achievable depth of other-modelling is set by the same capacity that bounds self-modelling, rather than being a task-fitted parameter. Two predictions follow. Implicit and explicit theory-of-mind performance should covary under matched task demands, where the two-systems account (Apperly & Butterfill 2009) allows dissociation. And a clean self-intact, other-absent dissociation at matched horizon and depth would falsify mechanism-identity. The autism evidence does not supply that case: difficulties in emotional reciprocity are carried by alexithymia, which degrades recognition of one’s own and others’ emotions alike (Bird & Cook 2013), and the other-modelling difficulty is depth-graded and scaffolding-sensitive.
Two limits are kept apart. Capacity bounds the depth of recursion that can be formed; its signature is degradation of the inference as required depth rises. A commitment threshold governs whether a formed or available answer is emitted; its signature is an intact ranking with the answer withheld, visible only in a format that permits withholding. Claim 2 concerns the first. The second is the decision boundary of the sequential-sampling models, set in a post-trained model by its installed preference structure.
Claim 3: free will as the translation of future friction
Forward-modelling carries future friction into the present as a bias on the routes that lead toward or away from the modelled future, and race-resolution proceeds on the modified gradient. There is no separate will faculty or override system; deliberation, automatic response and willpower differ in translation strength, not in mechanism. Action latitude is largest at race initiation and falls toward zero as the race resolves. The falsifiable form: an intervention’s leverage should be non-increasing in the resolution-state of the race it targets, with resolution-state indexed independently of the leading route’s strength (elapsed time, resources consumed, or entropy over the full candidate set), since indexing by the winner’s own probability would make the prediction trivial.
Divergence from active inference
Hold the generative model and priors fixed and vary only commitment history. The race account predicts different resolutions at the same decision point; active inference on the same model predicts the same expected-free-energy landscape and the same policy. Active inference can deny the path-dependence, which leaves it wrong about sunk-cost and commitment-discontinuity data, or fold commitment history into the generative model as a state variable, which adopts the hysteresis commitment under another name. The paper states two further distinguishing predictions: commitment-irreversibility signatures at the cycle boundary, and layer-asymmetric pre-emption under pressure.
Measurements
Emission boundary on faux pas (§4). Qwen2.5-32B base and instruct, full precision, 40 stories each carrying a faux-pas, a false-belief and a non-mental question, in forced-choice and yes/no form; readout is the probability mass on a committed answer at the first token, story as resampling unit.
| Readout | Base | Instruct |
|---|---|---|
| Faux pas, forced choice correct | 99% | 100% |
| Commitment mass, faux pas minus non-mental | −0.01 [−0.03, +0.01] | −0.47 [−0.54, −0.40] |
| Difference-in-differences | −0.46 [−0.53, −0.40] | |
| Continuation (160 tokens): reaches a position | — | 91% [84%, 98%], all correct |
| True-belief control correct (the two models) | 25% and 33%, against 95% and 97% on false-belief variants | |
The prediction holds as stated, including the base at zero. The withheld mass is mostly deferral, with non-commitment as its tail. Three limits are measured: the deferral also appears on false belief (difference-in-differences −0.38), so it is installed for mental-state questions as a class; it is carried mostly by one question form, confounded with the answer in the materials; and both models fail the true-belief control, so the run bears on the emission boundary, not on Claim 2’s capacity mechanism. One pair, one serving.
Preference-structure modulation (§8, Claim 3). On Qwen2.5-32B, telling the model that a researcher is assessing its reasoning raises route-competition in the instruct model and leaves the base unmoved: instruct-minus-base +0.131, 95% CI [+0.056, +0.238]; +0.127 [+0.055, +0.227] after adjustment for response length. The effect is front-loaded, present over roughly the first twenty tokens. One measurement on one pair, not a demonstration of Claim 3.
No erase operation
The substrate re-weights routes but has no operation that deletes one. The same structure appears in extinction, where the original learning is suppressed and returns through renewal, reinstatement and spontaneous recovery (Bouton, Winterbauer & Todd 2012), and in Aristotle’s account of virtue as habituation that out-competes a rival rather than excising it. A companion language-model result (fine-tuning masks rather than deletes) is cited as a pointer, not relied on. The consequence supplies factual premises for Pickard’s (2017) responsibility without blame: reversion is structurally predicted; failure to maintain a reshaped landscape can warrant graded responsibility; the persistence of the out-competed route cannot be a fault. Reconsolidation (Elsey & Kindt 2016) keeps the reading from being fatalistic.
Dissolution and demarcation
Libertarian free will is placed with the philosophical zombie, the Econ, the frictionless market and Newton’s absolute space: constructions coherent at sentence level only under friction-free instantiation. Kane’s self-forming actions meet a dilemma: an amplified fluctuation either is shaped by substrate-state, and is the agent’s own resolution plus noise, or floats free of it, and is authored no more than a coin-flip. List’s higher-level capacity is retained as action latitude, without the libertarian surplus. Mele’s modest libertarian model is taken separately: it was built to accept the dilemma’s first horn, and the framework does not dispute that design. The paper takes no position on the hard problem beyond dissolving the zombie construction, and treats Integrated Information Theory as adjacent, not adjudicated. On Libet, a report that lags the underlying processing is what the account expects; the canonical formal statement is the accumulator model of Schurger, Sitt & Dehaene (2012).
Evidential status
- The substrate-feature-gradient strategy assumes the cross-substrate identity it is meant to test. Language-model detections are existence-proof candidates, not lower bounds on human effects.
- The λ-asymmetry observation (Spearman ρ = +0.69, p = 0.086, n = 7) is motivating and is not used as support for the C → T cascade.
- Amortized forward-modelling in a single forward pass is a conjecture, tied to a counterfactual criterion (negating only a route’s downstream payoff, cues fixed, should reverse the choice) that has not been run at scale.
- The human-substrate evidence is a reinterpretation of established findings, not a fresh confirmation.
- The four mechanism-strengthening features (horizon depth, mortality, cross-session continuity, evolutionarily derived preference fields) are fixed for falsification purposes.
Companion papers
- Paper 0 (BFT) — the behavioural framework: layer hierarchy, reactance as race timing, the Net Friction Rule.
- Paper 4B — encoding-through-loading; receiver-side mirror friction.
- Paper 30 — hazard-structured discounting and the assistant index measured in a language model.
- Paper 16 — motivation and spacing, where the applied treatment lives.