The beige in the machine
Morten Münster's mediocrity, and why the same middle turns up somewhere without an ego
I read Morten Münster's Afdelingen for magisk tænkning og det utrolige potentiale i middelmådighed with great pleasure. The title translates roughly as The Department of Magical Thinking, and the remarkable potential of mediocrity, and it is in Danish only. I have read his other books too, and I have been on his masterclass. There he talked about the AFL method, which in Danish stands for behaviour, friction, solution, and that is where the reason I am sitting with this at all comes in: friction was the hard part. Describing the behaviour precisely is difficult enough. But getting hold of where the friction actually sits is often harder still. And what is friction, really? I have struggled a fair bit to explain it any better than "something that is a hassle". My search for an answer rather ran away with me, and it ended in a whole scientific theory, behavioural friction theory.
Back on track. Morten Münster's points are sharp and entertaining. What follows here is the dry version alongside: a mechanistic account of the phenomena he describes, as they follow from behavioural friction theory and the race mechanics underneath it. And one more thing worth bringing along: the same holds for language models.
It is not a design flaw
Here is the actual reason I wrote this down. When Morten Münster writes that the brain falls into black-and-white thinking, or that it goes for whatever is easy, it is tempting to hear that as a design flaw. As though something were wrong with the machinery, and we would be cleverer if only it had been assembled better.
That is not the case. Any system that has to arrive at one answer with limited time and limited resources will end up behaving in certain ways. That holds for a brain and it holds for a machine. It is the underlying physics, and it would be strange if our brains had not evolved to reflect it. They came about inside those limits, not in spite of them.
That is worth holding on to all the way through, because there is no ill will anywhere in this. The media are not evil or stupid. People are not evil or stupid. Nobody decided any of it. It is races, and it is physics.
That is why the mechanism is worth having beside the points. It says why they work, and it says one thing more, which a description cannot: when it stops working.
The underlying picture is the same one used everywhere else on this site. Inside you, small races run constantly between possible routes, and the route that crosses the threshold first wins. For the pictures behind that, see the water page.
What follows is a walk through the mechanisms underneath the points, in the same division as the book.
One: good enough is best
In short: stop trying to be extraordinary. Returns diminish, and if you spend too much time where they diminish, you cannot spend it on everything else. Then you become lopsided. Winning an Olympic medal in one thing costs you on every other measure.
Mechanically there are two things in that, and they are often run together.
One is that the return on further training falls away. The groove keeps deepening; that is not what stops. But once the route wins reliably, the extra training is no longer buying the win. It is buying refinement of how well it goes, and each further percent there costs disproportionately more than the one before. That is why a world-class performance takes enormous quantities of extra work for a few per cent. The mechanism also says when this sets in: not after a set number of hours, but once the route starts winning reliably.
And then the main point: those enormous quantities of extra work are paid for somewhere else entirely. What you put into the one route is taken from every other race you are running. So the curve does not turn because the skill gets worse. It turns because the whole becomes lopsided. That distinction matters: the axis is not effort, it is allocation.
And there is an uncomfortable footnote. You cannot subtract a groove, only lay a new one beside it. Over-investment is therefore not merely wasted time. It sticks.
This has also been measured, somewhere you can watch it happen. We took a large language model and trained it thoroughly in one particular way of answering. It got very good at that. It simultaneously got so much worse at everything else that its general ability fell from 86 per cent to 3. Not because it forgot anything, but because it started pasting the one learned manner onto tasks that had not asked for it. That is "we won the Olympics in one thing and it cost us everything else", measured on a machine.
Münster also brings in Barry Schwartz here, and his distinction between people who maximise and people who settle for good enough, the satisficers. The latter are consistently happier. The former spend far more time turning decisions over, both the ones taken and the ones not, and that turning-over has a price in regret and in mood. Mechanically it is the same bill as above, just paid in attention rather than in hours: holding a race open costs something, and holding it open after it is effectively settled costs without returning anything.
The curve has a floor as well as a ceiling. Below a certain level nothing happens at all: a model that is too small cannot apply a rule even with just the one rule to keep track of, and simplifying the task does not rescue it. That is the left-hand side of the figure.
The same pattern is familiar in teaching. Support and explanations that help a beginner will harm someone experienced, because the experienced person already has the route and now has to hold the explanation as well. The effect is well known in the literature as expertise reversal, described by Slava Kalyuga and colleagues in 2003. Good enough is not only better for the whole. Past a point, more is actively worse for the thing itself.
Two: brilliant basics
In short: solve the basic thing. Do not fight your way out onto the decimal places while the obvious sits unfixed. This is bike-shedding, the phenomenon where a committee approves a nuclear plant in ten minutes and then spends two hours on the colour of the bike shed.
Bike-shedding usually gets an explanation about vanity: people want to contribute, so they contribute where they can. Mechanically there is a simpler reason.
A race is settled by which route accumulates fastest, not by which one matters most. The bike shed is a cheap route: everyone can judge it, everyone has a view, the evidence pours in, and it crosses the threshold first. The nuclear plant is an expensive route: almost nobody can judge it, so it never accumulates enough to win. The outcome looks like a failure of priorities. It is an ordering.
And this is not an analogy for us, it is something we have measured. Hand a system an entire rulebook at once so as to be covered, and the very rule you wanted followed loses, because it has to compete with all the others. A short behavioural recipe wins under pressure where the exhaustive manual does not. That is brilliant basics seen from the other end, and it is the whole point of the page on compliance as behaviour.
The urge to live out on the decimals has the same shape. The elaborate solution is many routes, each of them weak. The basic one is a single deep route. Under pressure, deep beats broad.
Three: the mediocre middle way
In short: the case against black-and-white thinking. Either I am on a diet or I am not. Either I am someone who trains or I am not.
Here the mechanism comes closest to what I mean by calling it physics. Black-and-white thinking is not a fault in the thinking. It is what a race under high pressure looks like from outside. Pressure shortens the race. The first route across the threshold wins, and the alternatives never accumulate anything. From the inside it does not feel like cutting corners. It feels like being right. It has been measured in language models: give them little time and the competition between possible answers collapses onto whatever already weighed most, and the wide search simply stops. Pepsi or Coke with the waiter standing there is the picture of the same thing in us, not the measurement. That is on the page about what pressure does to a choice.
Any system that has to deliver one output from several candidates in limited time will look categorical from outside. No intent is required. And from that follows a piece of advice a description cannot give: you do not fix black-and-white thinking by arguing for nuance. That is information fed into a race that is already settled. You fix it by lowering the pressure so the race has time to finish.
The book has an example that sits right inside that mechanism and is easy to recognise. The body can take neither everything nor nothing; it likes the middle. Yet those two extremes are exactly what we do: either we train until we are deep in the red, or we do not train at all. Münster calls it the black-and-white brain overcompensating in both directions, and he names why the middle feels wrong: "a bit" looks pathetic. We are raised to believe it has to hurt before it counts as training. So the easy version, which is what works for most people, is not merely a weaker choice. It is a choice that does not feel like one.
And a distinction that decides what to do
This is the part I want to add myself, because it has practical consequences. "Black-and-white" covers two different mechanisms, and they call for opposite responses.
Case one: the bar of chocolate. You promised yourself you would not. You end up eating one square. Then you eat the whole bar. Here the reward has gone. The promise was the only thing paying you anything for abstaining, and the moment it breaks, the day is lost anyway. The race still runs, but the abstaining route has no winnings left to accumulate. So the other route wins unopposed.
Case two: "I am someone who runs three times a week." Here the competition has gone. You do not weigh up whether to run today. The identity settled it in advance, so the competing alternative is never set up in the first place.
The two look alike from outside, and they need opposite things.
If the reward has gone, the goal has to be made graded rather than binary, so one square does not reset the day. Seven days out of ten is a target you can fail at without losing it. That lever sits in the toolkit under behaviour design.
If the competition has gone, you actually have a strength, and it can be built deliberately. But it carries a built-in fragility worth remembering: when the identity breaks, and it will, because you miss a week, you fall into the first case on top of it. "I am not that person after all." An all-or-nothing identity is fragile precisely because it removed the competition instead of beating it.
That is why the distinction matters. Give the grading lever to someone with a competition problem and nothing happens. Give the identity lever to someone with a reward problem and you have built them a trap.
Why we always add instead of subtracting
One of the strongest parts of the book draws on Leidy Klotz. Ask people to improve a text and they write more. Ask them to improve a recipe and they add ingredients. Ask them to improve an itinerary that is already crammed and they cram more in. And it turns out the opposite was usually better: the text improved by getting shorter, the food by fewer ingredients, the trip by two churches a day instead of four.
Klotz points among other things to an evolutionary reason, and Münster reports it. In race terms you can put it more precisely, and it is, I think, the sharpest single application on this page: an addition is a route, a removal is not.
When you look for an improvement, the candidates get set up and then they race. "Add X" is a candidate you can think. "Remove Y" requires you first to form a picture of something that is not there, and an absence does not enter by itself. So the removal does not lose the race. It is not in it. That is exactly the same mechanism as WYSIATI below, just on the actions rather than on the information.
And then there is a twist that makes it worse, and which I think is the finest detail in the whole mechanism: you can only remove something by adding something. To get the removal into play you have to add the candidate "remove Y". So the field of options can only grow. You cannot subtract from the very set that races.
And that addition has a price. The moment you say "remove Y", you have named Y. Now Y is in the race, and a route that has been named is easier to activate than one that has not. So proposing to remove something also strengthens the thing you wanted gone. That is exactly the mechanism we have measured in language models elsewhere: an instruction never to do X makes X more likely rather than less, because the prohibition has to name the forbidden thing. It sits under compliance as behaviour.
That gives a neat account of why a rule beats a discussion. "Shouldn't we drop the Y project?" adds a candidate and switches Y back on at the same time. "One project in, one project out" makes removal a standing slot in the process, so nobody has to argue for removing anything in particular. The rule avoids the naming, and so it avoids the price of it.
And then it becomes self-reinforcing, which Klotz also found: thinking in less takes more energy than thinking in more. Under pressure the race is shorter, so the cheap route wins. The cheap route is to add. What you added raises the load. So there is more pressure next time, and the cheap route wins more clearly still. You cannot think your way out of it by pulling yourself together, because being under pressure is what produces the addition.
It also explains why the levers the book suggests work. Teddy in, teddy out, which Münster has from the consultant René Bomholt, or one project in and one project out. They do not make people cleverer. They enter the removal as a candidate, every single time something is added. It is a rule that changes the field rather than the thinking, and that is why it survives pressure where good intentions do not.
The thread under all three
Münster draws on Daniel Kahneman's WYSIATI, what you see is all there is: we build a coherent story out of what is in front of us and rarely weigh in what is missing. That is not a fourth point beside the three. It is what ties them together.
Mechanically it is close to trivial, which is why it is strong: a race can only run between the routes that have been set up. An absence does not carry low weight. It carries none. There is no route for "the thing I did not think of" that could lose. So confidence is computed over whatever happens to be present.
Run that down through the three. The bike shed is not only cheap to judge, it is also present; the expensive route often never gets set up, so it does not lose, it does not take part. Black-and-white feels well founded because the two present options are everything there is; the nuance is not a weaker candidate, it is one nobody raised.
And sharpest on the first section: the price of over-investing in one thing is paid by all the other races, and those are not running at the moment you decide. They are idle. The bill is therefore structurally invisible exactly where it is signed. That is why the trap feels sensible from the inside, and it is, I think, the best answer to the question you are left with once you have nodded along to good enough being best: why is it so hard to see?
Why the spectacular always wins
At one point Münster brings in Hans Rosling's point from Factfulness: normality is the media's kryptonite. What fills the news is the unusual, and so we end up believing the world is far more dangerous than it is. Nobody clicks on things being slightly better than yesterday.
It is the same law as the bike shed, just on a different axis. A race is won by the route that accumulates fastest, and for attention the rate is set by surprise. The expected carries almost no information. That route was the favourite already, so a confirmation moves nothing. The unusual moves a great deal. So the spectacular crosses the threshold, and the liver-paste-coloured normality never gets going.
And then WYSIATI closes it. What reaches you is exactly what won those races. Afterwards you work out how dangerous the world is over precisely that field. The ordinary case is not a weaker candidate in that calculation. It was never entered. So your answer is not the product of poor thinking. It is the right answer over the wrong field.
That also says what works. You do not fix it by asking people to think more critically, because the thinking is not what fails. You fix it by making the absent present, by putting the denominator into the field. That is exactly what Rosling does when he sets a curve beside the story, and it is the same lever as under behaviour design.
And then a point that turns back on us, because it is easy to miss. The sentence "the media love sensation" is itself a black-and-white statement. It sounds as though someone came up with it because they are bad at their job. But the same physics holds for an editorial desk as for everyone else. Nobody decided that normality should go. It simply loses a race, every single day, to something that accumulates faster. The difference matters, because an accusation is something you can only agree or disagree with. A mechanism is something you can do something about.
It happens in the machines too
And here is why the page is called what it is. The same patterns turn up in language models. A language model has no character to have a flaw in, no vanity to tend and no clicks to chase, and yet its best place is the same undramatic middle. What Münster calls the liver-paste-coloured normality, the beige, is in the machine too.
The clearest case is all-or-nothing. We have tried training a model to say "I do not know" when it is uncertain. That is a sensible goal. But trained hard enough, it collapses into saying it almost every time, including about things it knows perfectly well. It stops being of any use. That is "I broke the rule, so the rule no longer applies" in machine form, and it is not an impression, it is a number.
It suggests that what Münster describes is the shape of bounded computation, and that we are one instance of it among several.
A detour: what we like
Here I step away from the book, because this is not in it. But it is in keeping with Münster's larger point as I read it: if we know the mechanisms, we have a chance of working against them. And this is the consequence I find the most striking of them all.
In friction theory, negative valence is not a thing in its own right. It is reactance, the resistance that arises when something presses on a route you already have. And the positive is then, in the ordinary case, the absence of that resistance. The absence of friction, in other words. Good is not a property of the thing. It is what it feels like when a route runs without hitting anything. The friction itself can sit in four different fields, which are set out under behaviour design.
So you can ask where the resistance goes, and the answer is familiar from a quite different literature: exposure. Robert Zajonc showed as early as 1968 that being exposed to something repeatedly makes people like it better on its own, with nothing else having happened. Mechanically that is straightforward. Every time the route runs it becomes easier to run. At some point it runs without resistance. And then it is good.
That is worth knowing if you want to raise tolerant children, or build a society with room in it. Because then tolerance is not an attitude you argue people into. It is what happens once the route has been run enough times to stop producing resistance. And the converse comes with it: intolerance need not be malice. It can be a route that was never run.
There is a condition, and the neat part is that it falls out of the mechanism itself. Exposure only works if it does not itself arrive as pressure. Pressure on a defended route strengthens it rather than weakening it, and that is precisely reactance. So forced contact under threat should push the other way. That fits what is known: Thomas Pettigrew and Linda Tropp's large meta-analysis across 713 samples finds that contact between groups usually reduces prejudice, that it works better still under good conditions, and that the open question is what prevents it from working. The mechanism here offers a candidate answer: contact that arrives as pressure.
How sure is this
The inverted U itself is not my discovery, and it is more than a century old. Robert Yerkes and John Dodson described it in 1908, and in neuroscience it has since been tied to how signalling chemicals such as dopamine and noradrenaline act on the prefrontal cortex, where both too little and too much degrade performance. Amy Arnsten has pulled that thread together. Nor is it new that the best point for a bounded system sits in the middle rather than at the end: Herbert Simon called it satisficing in 1955, the idea that a system with limited view does not optimise, it finds something good enough.
What I add is the mechanism underneath, and that it can be read directly in a system we can measure. That is a claim which can turn out to be wrong, and the places where it already did not hold are in the section above.
Münster ends his own book arguing for making things more complex rather than simpler. That is the right place to end, and the mechanism says why it is hard: holding a nuance open means keeping a race running, and that costs. Under pressure it closes on its own. So nuance is not an attitude you can decide to hold. It is a state you have to make room for.