Routes, not modalities: coincidence of traces, not sensory channels, drives multimodal learning

Paper 26 · Pødenphant Lund, T. (2026) · Empirical note · Live on Zenodo

A robust rule of thumb, add an image and learning improves, is applied outside the case it was tested on: the image that is incongruent with the text. This note argues that the active variable in multimodal benefit is not the sensory modality but the coincidence of traces, the number of inputs whose routes converge on the same answer, and reads it as friction in a language-model substrate rather than as accuracy alone. Congruent traces summate; an incongruent input does not vanish, it imposes a rejection-race cost. The phenomenon is not the claim: the negative effects of irrelevant material, and the race-vs-summation formalism, are inherited. The contribution is the mechanism, what governs it, and two things the neighbours do not carry.

DOI (concept)10.5281/zenodo.20679145
TypeEmpirical note (LLM/VLM substrate; friction as the readout)
Order variableCoincidence of traces (congruent routes converging on one answer)
ReadoutCompeting-routes friction (retrieval difficulty), not accuracy alone
AuthorTomas Pødenphant Lund [ORCID]

TL;DR

Standard accounts credit multimodal benefit to cross-channel integration, the modality pairing itself. This note reframes the active variable as the coincidence of traces: each input (a word, an image) opens a route toward an answer, congruent routes stack, and an incongruent input still forces a race to establish that it is noise. Length itself starts races; irrelevant length is pure cost. Modality is one reliable trace-supplier, not the mechanism.

Read as friction rather than accuracy, four results follow, and the cost turns out to be capacity-relative: a large model runs the rejection race for free while a small one pays measurably. The formalism (Miller's race-model inequality, redundancy gain) and the negative phenomena (distraction by irrelevant text/image) are inherited and cited; the contribution is the mechanism, the size-dependence, and a retention effect.

Findings

What governs the cost: capacity

Decomposing the incongruent-image cost gives a crowding term (monotone in the token-footprint of the irrelevant image, matched by an equally long irrelevant text) plus a small image-specific term. The decisive result is that the magnitude is capacity-gated: the large model attends to the irrelevant image as much as the small one yet does not pay, and within one model family the cost collapses roughly tenfold from the small to the large size (order 0.12 to 0.013 on the paper's scale). The image-tokenisation account of the between-family difference is tested and rejected. This instantiates the series' capacity-relativity: the rejection race runs on every model, but only the resource-pressured one pays a measurable friction cost.

Structure, not image

On a multi-hop retrieval task, performance is near-floor when the premises are given as running text in random order but recovers when the same premises are co-located in a structure, either a bulleted list or a diagram. The active variable is co-location, not modality, and the human ordering inverts: structured text is as good as or better than the diagram for the model, and the diagram is the most expensive (an image-reading cost that grows with graph size). Single-family, single-task, and reported as an early indicative result, it points the same way as the main line.

Prior art and what is inherited

The race-versus-summation question, and its standard test, are not new: the race-model inequality is due to Miller (1982), and the gain from an extra congruent channel is redundancy gain. The negative phenomena, distraction of a language model by irrelevant text and degradation of a vision-language model by an irrelevant image, are already reported by others. On the human side, the two lead findings correspond to the redundancy effect, the split-attention effect (Sweller, Chandler and colleagues), and the seductive-details effect (Harp & Mayer 1998), organised under Mayer's multimedia-learning principles. Recent work has also transferred learning-psychology constructs to language models. The paper claims none of these. It claims the substrate-level mechanism (the race), the friction readout, the capacity-gating of the cost, and the retention effect.

Scope and honesty

The core mechanism (capacity-gated rejection-race cost, with the tokenisation alternative tested and rejected) is settled within the tested range; the retention result and the structure-not-image result are small-scale and indicative (few model sizes, few facts, single task families) and are labelled as such. The paper places itself in the cross-substrate family (Papers 23, 24): a sorting exercise over which human learning-patterns reappear in a non-human system as measurable friction, and which (e.g. expertise reversal) the small scale cannot yet resolve.

Connections to other papers in the series

Read the paper

The full note is on Zenodo (concept DOI 10.5281/zenodo.20679145):

Pødenphant Lund, T. (2026). Routes, not modalities: coincidence of traces, not sensory channels, drives multimodal learning. Zenodo. https://doi.org/10.5281/zenodo.20679145

Read on Zenodo → · Plain English version · Dansk version