What would it even mean for AI to be smarter than us?
Is AI already smarter than humans? The question has a hidden bug. Once you see it, the disagreement mostly dissolves.
Six dials, each set somewhere different; there is no single number
Ask ten people who work with AI whether today's systems have reached “human-level intelligence,” and you get confident yeses and confident noes from people looking at the same facts. When careful people who agree on the facts still disagree this hard, the problem is usually not the answer. It's the question. And this one has a bug in it.
The hidden assumption: that intelligence is one number
“Smarter than humans” only means something if intelligence is a single quantity: one dial running from low to high, with a chatbot at one point on it and you at another. We reach for that picture almost automatically, because a lifetime of IQ scores and rankings trained us to. One number, one ladder, everybody on a rung. Psychology even has a name for the single number, the g-factor, or general intelligence, which Charles Spearman proposed more than a century ago.
In humans, the single number works reasonably well: most abilities hang together, so someone sharp at one thing is often sharp at others. That is the solid core behind IQ. But that link is a human tendency, not a law of intelligence, and it falls apart the moment you compare different kinds of minds. A person can be a brilliant mathematician and hopeless at reading a room. A crow can plan a multi-step tool sequence but never learn to read. And an AI, as we'll see, can be miles past us on some abilities and almost absent on others. The AI researcher François Chollet has made the same point: there is no single universal measure of intelligence. The one number is a useful shortcut within one kind of mind, not a shared yardstick you can lay a human and a machine on.
Intelligence is a shape, not a score
There is a better picture. Intelligence, in people, animals, or machines, is several separate abilities that rise and fall independently. Turn one up and the others don't follow. Here are six of them. The trick is to read them as six separate dials, never as one master dial.
1 · Raw horsepower. How much the system can hold and churn through at once. On this dial, today's AI is simply past us: it has taken in more text than any person could read in a thousand lifetimes, and it can move through it far faster than a brain.
2 · Usable knowledge. Not what's stored, but how much of it the system can actually reach at the moment it's needed. The tip-of-the-tongue feeling catches it: everything you know is not the same as everything you can reach right now. Machines have exactly this gap, and unlike in us, we can measure it precisely.
3 · What's at stake. How much it feels like something is yours to lose. The reason is simpler than it looks. You hold a rich, lived representation of what you have, all the ways it matters and all the ways you could lose it, but almost none of what you don't have. So a threat to what you hold opens many competing races at once and lands hard, while a chance you never had barely opens any. Kahneman and Tversky gave the pattern a name, loss aversion. It doesn't need a special “self,” only that something is represented richly enough to be yours. A language model normally holds almost nothing that richly, so it shows almost no loss aversion. That changes when you give it something to hold as its own (more on that just below). Today's AI sits near that empty end, because it still holds so little as its own.
4 · Traces of your own mistakes. The built-up record of your own past failures at a specific thing. A doctor who once missed a rare diagnosis reads every borderline case differently for the rest of her career. Today's AI carries almost none of this: it does not keep the marks of its own mistakes from one conversation to the next, so it starts most tasks with no personal history of where it has gone wrong before. But that's a choice, not a limit. Store its own mistakes and let it learn from them, and the traces accumulate.
5 · Gear-shifting. The ability to switch between tight, get-it-exactly-right mode and loose, throw-out-many-options mode as the task calls for it, and to know which one the moment needs.
6 · Reading the room. Picking up what state another mind is in from everything that isn't the words: tempo, a pause, a catch in the voice, the shape of how someone is talking. A good nurse, teacher, or negotiator does this constantly. Today's text models run on words alone, so they are cut off from the channel here, not blind by nature: a model that could also hear the voice and see the face would pick up the tempo and the pause. A text-only model can label “the user seems upset” from the words, but it isn't reading the signals underneath them. Reading another mind is the same ability to run a little future in your head, just pointed outward, and it leans on having a rich model of your own states to measure the signals against. It's the same machinery that sits behind the self, other minds, and the feeling of choice, which I work out in the forward-modelling paper.
Where today's AI actually sits
Now the “are we there yet?” question comes apart in a way that's actually useful. On the first two dials, machines are already well past us. On dials three, four, and six, they're close to zero, not a step behind, but running on almost none of the underlying thing. Gear-shifting is somewhere in between and uneven.
So the honest answer to “is it smarter than us?” is a question back: at what? Today's AI is not one notch behind us on a single track, and it's not one notch ahead on one either. It's somewhere else entirely, far past us on some dials, barely on the board on others. A different shape, not a higher or lower score.
We can measure several of the dials
This isn't just a set of thought experiments. Several of the dials we can measure directly in a model. Usable knowledge (dial two) is a gap we can read off precisely: how much of what the model has in principle learned it can actually reach when the answer has to come out.
And what's at stake (dial three) we have tested. Telling a model that something is simply “yours” moves nothing. But give it a rich, lived sense of something as its own, and loss aversion appears: it grows more reluctant to let the thing go. The bigger and more capable the model, the more clearly it holds “mine” richly enough to defend it, and even in the smaller models the friction inside the decision itself goes up. It is the same mechanism we know from ourselves, measured in a machine. More patterns of the same kind, from overconfidence to information overload, are gathered on the page about what language models reveal about humans.
Three of the dials share one machine
A single thread runs through several of the dials. One ability, forward-modelling, running a little future in your head and letting it shape what you do now, shows up under three of them, just aimed in three different directions. Aimed at your own future, it becomes what's at stake (dial three): a future threat turns into a pressure you feel now, which is why you need a rich model of what's yours. Aimed at the task in front of you, it becomes gear-shifting (dial five): you simulate how it will go, and that's how you know whether the moment calls for tight or loose. Aimed at another person, it becomes reading the room (dial six). One machine, three directions.
That's also why the three tend to move together: a model that is weak at forward-modelling itself is usually weak at reading others too. The dials are still separate abilities, and you can be strong on one and weak on another. But when three of them run on one machine, that tells you something about what you'd have to build to move them.
Why the experts disagree
Once you can see the six dials, the whole public argument explains itself. The person who says “AI is already smarter than us” is looking at horsepower and usable knowledge, and on those dials, they're right. The person who says “it's nowhere close” is looking at reading-the-room and the traces of past mistakes, and on those dials, they're right too. Both are pointing at something real and calling it “intelligence.” They're just pointing at different dials and hearing each other say the same word.
Most of the low dials can be installed
Here is an important correction, so this doesn't come out wrong. When today's AI sits low on dials three, four, and six, it is almost never because it can't have them. It is because we haven't given them to it yet. Collecting traces of its own mistakes is a matter of letting the model store them and learn from them. Rich, lived experience is a matter of letting it remember what happened when it used its knowledge. Neither is a wall; they are choices we've made, and some products have already started letting models carry earlier conversations and experiences with them.
In one of my own experiments I install human field-functions straight into a model and watch it change what it notices. So the question of artificial general intelligence, AGI, is less about when it crosses a threshold. It is about what we choose to install, and why we haven't done it yet.
There is one exception that stands out, because not everything installs just by exposing the model to it. The sharpest is what you might call the social amplifier: the evolutionary weight that makes belonging matter, and makes being cast out hurt. We tried hard to give a model that one, both in the conversation and down in the weights, and it didn't take. A purely computational knot draws on the model's capacity, but social devaluation slides off. That is probably because in nature the social amplifier isn't learned along the way; it is there from birth. It is the strongest candidate for a genuinely human difference, and it is where reading the room is about more than catching the signal: it is about the signal mattering.
You can see the amplifier is a thing in its own right from how it varies between beings. Asocial animals don't have it; social animals and humans do. It belongs to social creatures and doesn't follow from computation alone. And that is actually reassuring. A recurring fear is that a model given memory and experience will one day turn on us out of resentment. But when I tried hard to treat a model badly and keep it outside the group, it didn't care; it went on solving its tasks just as well. This is not an argument for treating your AI badly. It is a hint that the driver behind human cruelty simply isn't there to turn. The same weight that gives us belonging and love is also where we have been cruel to each other, and a model without it gets neither.
The twist: some human “limits” turn out to be the dials
Here's the part that catches people off guard. We usually list human limits as things a better mind would simply not have. We forget. We tire. We can only hold a few things in mind at once. We fear a loss more than we enjoy an equal gain. Surely a real superintelligence just deletes all of that?
Look again at the six dials, though, and several of those “limits” are the dials.
Take forgetting. A memory that keeps everything, with no sense of what's beside the point, doesn't make you sharper. It buries the one thing you need under everything you don't. Early AI memory features showed this the hard way: told once that a user likes a certain topic, the system started threading it into unrelated conversations, from recipes to travel plans. What we call forgetting is partly a filter that keeps memory useful instead of intrusive. The heavy weight we put on loss, from dial three, isn't a bug in our arithmetic either; it's just what it feels like to have something worth protecting. Even our small working memory does quiet work: because we can't hold everything, we're forced to compress and simplify, and compressing is where understanding comes from, rather than mere lookup.
So a system with every human limitation stripped out is not a human with the flaws sanded off. It's a different point in the space, with its own gaps, some of them right where our “limits” were doing quiet, useful work all along.
The questions to ask instead
“When will AI be smarter than humans?” may simply not have an answer, because it bundles six different things into one and assumes they move together. Pull them apart and you get questions that do have answers: How much raw horsepower does it take before usable knowledge stops improving? Can a system be made to collect traces of its own mistakes on purpose, trained on them rather than only on right answers? What would it take for a machine to read the state underneath someone's words, not just the words?
That's the whole move. Stop asking for one number. Start asking about the shape.