Routes, not modalities: coincidence of traces, not sensory channels, drives multimodal learning
Paper 26 · Pødenphant Lund, T. (2026) · Empirical note · Live on Zenodo
A robust rule of thumb, add an image and learning improves, is applied outside the case it was tested on: the image that is incongruent with the text. This note argues that the active variable in multimodal benefit is not the sensory modality but the coincidence of traces, the number of inputs whose routes converge on the same answer, and reads it as friction in a language-model substrate rather than as accuracy alone. Congruent traces summate; an incongruent input does not vanish, it imposes a rejection-race cost. The phenomenon is not the claim: the negative effects of irrelevant material, and the race-vs-summation formalism, are inherited. The contribution is the mechanism, what governs it, and two things the neighbours do not carry.
| DOI (concept) | 10.5281/zenodo.20679145 |
| Type | Empirical note (LLM/VLM substrate; friction as the readout) |
| Order variable | Coincidence of traces (congruent routes converging on one answer) |
| Readout | Competing-routes friction (retrieval difficulty), not accuracy alone |
| Author | Tomas Pødenphant Lund [ORCID] |
TL;DR
Standard accounts credit multimodal benefit to cross-channel integration, the modality pairing itself. This note reframes the active variable as the coincidence of traces: each input (a word, an image) opens a route toward an answer, congruent routes stack, and an incongruent input still forces a race to establish that it is noise. Length itself starts races; irrelevant length is pure cost. Modality is one reliable trace-supplier, not the mechanism.
Read as friction rather than accuracy, four results follow, and the cost turns out to be capacity-relative: a large model runs the rejection race for free while a small one pays measurably. The formalism (Miller's race-model inequality, redundancy gain) and the negative phenomena (distraction by irrelevant text/image) are inherited and cited; the contribution is the mechanism, the size-dependence, and a retention effect.
Findings
- No redundancy gain from a redundant image. When the answer name is already in the text, a congruent image adds essentially nothing across three setups and two model families; the point estimate is marginally negative. In race-model terms, the congruent-channel gain humans reliably show is absent here, and that contrast is what becomes informative once the question is posed in the field's own vocabulary.
- An incongruent image imposes a measurable cost. Attaching a foreign image to a text-answerable question raises retrieval friction (statistically reliable). The mechanical reading: the race that must run to reject the noise. This is the substrate-level account of the "seductive details" intuition.
- Recognition dissociates from use. A symbol-to-referent mapping can be learned to near-ceiling while the model remains unable to use the symbol to retrieve the referent's property, distinguishing acquiring an association from being able to compose over it.
- Multi-trace encoding aids retention. Facts encoded through multiple congruent traces show a consistent (small) resistance to overwriting under subsequent unrelated fine-tuning. Direction replicates across five setups; effect size is small and reported as indicative, not settled.
What governs the cost: capacity
Decomposing the incongruent-image cost gives a crowding term (monotone in the token-footprint of the irrelevant image, matched by an equally long irrelevant text) plus a small image-specific term. The decisive result is that the magnitude is capacity-gated: the large model attends to the irrelevant image as much as the small one yet does not pay, and within one model family the cost collapses roughly tenfold from the small to the large size (order 0.12 to 0.013 on the paper's scale). The image-tokenisation account of the between-family difference is tested and rejected. This instantiates the series' capacity-relativity: the rejection race runs on every model, but only the resource-pressured one pays a measurable friction cost.
Structure, not image
On a multi-hop retrieval task, performance is near-floor when the premises are given as running text in random order but recovers when the same premises are co-located in a structure, either a bulleted list or a diagram. The active variable is co-location, not modality, and the human ordering inverts: structured text is as good as or better than the diagram for the model, and the diagram is the most expensive (an image-reading cost that grows with graph size). Single-family, single-task, and reported as an early indicative result, it points the same way as the main line.
Prior art and what is inherited
The race-versus-summation question, and its standard test, are not new: the race-model inequality is due to Miller (1982), and the gain from an extra congruent channel is redundancy gain. The negative phenomena, distraction of a language model by irrelevant text and degradation of a vision-language model by an irrelevant image, are already reported by others. On the human side, the two lead findings correspond to the redundancy effect, the split-attention effect (Sweller, Chandler and colleagues), and the seductive-details effect (Harp & Mayer 1998), organised under Mayer's multimedia-learning principles. Recent work has also transferred learning-psychology constructs to language models. The paper claims none of these. It claims the substrate-level mechanism (the race), the friction readout, the capacity-gating of the cost, and the retention effect.
Scope and honesty
The core mechanism (capacity-gated rejection-race cost, with the tokenisation alternative tested and rejected) is settled within the tested range; the retention result and the structure-not-image result are small-scale and indicative (few model sizes, few facts, single task families) and are labelled as such. The paper places itself in the cross-substrate family (Papers 23, 24): a sorting exercise over which human learning-patterns reappear in a non-human system as measurable friction, and which (e.g. expertise reversal) the small scale cannot yet resolve.
Connections to other papers in the series
- Paper 16 (The Physics of Learning) — the human-facing companion: desirable difficulties, cognitive load and the dump/dilute failure modes as race-architecture consequences.
- Paper 1 (Friction Theory) — competing routes under bounded resources; the friction readout used here.
- Paper 24 and Paper 23 — the same cross-substrate family (visual illusions; social reactions).
Read the paper
The full note is on Zenodo (concept DOI 10.5281/zenodo.20679145):