Reading creativity from the inside: a competing-routes view of the diversity–coherence trade-off in language models

Paper 31 · Pødenphant Lund, T. (2026) · Empirical paper · Live on Zenodo

Creativity is treated here not as a property a model possesses but as a property of how its per-token route competition is run, read directly from the model's own next-token distribution at no additional inference cost. The creativity literature measures the output (by prompting for variation, or by drawing many samples); per-token internal measures do exist in the neighbouring literature, so the claim is not first-to-read-from-the-inside. It is the use of that reading as a creativity instrument: held against judged novelty × value, and used to predict output diversity from a single pass.

DOI (concept)10.5281/zenodo.20840284
TypeEmpirical paper (LLM substrate; per-token route competition as instrument)
Control variableLandscape breadth (suppress-default / inject-routes), not sampling temperature
JudgesFour independent LLM judges; dissent reported rather than averaged away
AuthorTomas Pødenphant Lund [ORCID]

TL;DR

Three results. (1) Pushing the landscape toward the non-default raises novelty monotonically while value and coherence fall, so novelty × value describes an inverted U: too narrow is banal, too broad is incoherent. Temperature does not traverse this curve. It fails to raise novelty and degrades value directly, moving off the curve rather than along it, which is the mechanical reason raising temperature does not produce creativity. (2) The alignment-narrowing fingerprint at the choice point is graded and recipe-dependent, not universal, with one family reversing it outright. (3) Output diversity is predictable from route competition in a single generation, where prior work reported a null.

Prior art is credited explicitly: the concentration phenomenon (alignment sharpens the distribution, most at early positions) is established by Yang, Li & Holtzman (2025), with Mohammadi (2024) on lower token entropy in aligned models. The contribution is the instrument framing, an independent replication, and the refinements below.

Result 1: an inverted U in novelty × value, and what temperature does

Landscape interventions (suppressing the default route, injecting distant material) raise rated novelty monotonically while rated value and coherence decline past a point, so the product peaks at intermediate breadth. This is the series' capacity-match inverted U instantiated in generation: the downward arm exists because of the value term, and without it the curve would be a ramp. It also gives a precise reading of the freedom question: usable freedom is freedom within a context, since the context is the constraint that converts novelty into value.

The temperature contrast is the sharp result. Raising sampling temperature does not increase novelty; it degrades value and coherence and reaches incoherence without passing through the high-usability region. Geometrically the landscape lever moves along the novelty–value frontier while temperature moves off it. A hint that optimal breadth scales with model size (smaller models require more pushing, larger ones peak earlier) proved judge-dependent and is reported as suggestive only.

Result 2: the alignment fingerprint is graded, with a counterexample

In two of three families (Qwen, Mistral) the aligned model shows reduced route competition at the choice point (early positions) relative to its base sibling. The phenomenon is Yang et al.'s; the contribution is replication with an independent instrument plus the refinement that the effect is graded and recipe-dependent. The third family (Llama-3.1) reverses it: its aligned model is the least collapsed of the set, broadest at the choice point and most varied in output. That is a genuine counterexample to an "alignment always sharpens" reading.

The load-bearing consistency is that across the three families, the degree to which alignment narrows route competition tracks the degree to which it collapses measured output diversity, so the instrument correctly orders models by how mode-collapsed they are. On the Result-1 curve a heavily aligned model sits on the narrow arm, and the lever is to reopen routes at the choice point rather than to raise temperature. This yields a stated but unrun prediction: a base model could be fine-tuned toward higher creativity by preserving choice-point breadth while installing the coherence and steerability base models lack, landing near the peak instead of collapsing to the narrow arm. Running the third family was a reviewer demand; it reversed the prediction and strengthened the result.

Result 3: single-pass diversity prediction, and a prior null explained

Route competition at the choice point predicts realised output diversity from one generation, where existing approaches require multi-sample estimation or a dedicated prompting scheme. Yang et al. searched for this relationship (branching factor against lexical diversity) and reported a null. Reviewers required that the explanation for the discrepancy be tested rather than asserted; the ablation (Appendix A.8) disproved the paper's first explanation, since the correlation survives their lexical measure (Distinct-2) and is in fact stronger there than on the semantic measure. The surviving explanation is range restriction: the design here manipulates breadth via ladder conditions, giving the predictor real variance, whereas the prior design contained no breadth manipulation; and the single model in which the correlation vanishes here is the one with the most compressed competition range.

Evidence quality and stated limits

Four independent judges were used. The inverted U held under three of four; the dissenting judge rewards incoherent output and disagreed on one model, which is reported rather than averaged away. The size-scaling story was judge-dependent and demoted. A formative review round returned major revision, whose central demand (the third model family) produced the reversal above and forced the graded reframe; the closing round on the reframed version returned accept and accept-with-minor-revisions, both noting that the reversal strengthened the paper. One caveat is carried throughout: a mechanism (a live competitor makes deviation worthwhile) is not a deployable selector, which would require per-model and per-dataset calibration.

Where it points

Selecting a route that deviates from the statistically most likely is the same event underlying two phenomena usually studied apart: creativity, where the value of the deviation is novelty, and error correction, where a confidently wrong default is overridden by a correct alternative and the value is correctness. Both require a live non-default competitor at the decision point, and correctness is a cleaner objective value measure than a judge. The hard case is confident-wrong: a flat competition signal with an incorrect commitment, where no competitor is visible to open, which neither the signal nor a route-opening intervention addresses. The inject-routes lever has a human-literature name, mediated association or bisociation (Mednick 1962; Koestler 1964); the bridge-object mechanism is folded into the discussion with a falsifiable prediction that competition should rise specifically at junction tokens, and most on a heavily constrained model.

Connections to other papers in the series

Read the paper

The full paper is on Zenodo (concept DOI 10.5281/zenodo.20840284):

Pødenphant Lund, T. (2026). Reading creativity from the inside: a competing-routes view of the diversity–coherence trade-off in language models. Zenodo. https://doi.org/10.5281/zenodo.20840284

Read on Zenodo → · Plain English version · Dansk version