ESSAY 04 ยท 16 September 2026

The Conditions for Correction

The longer I work with synthetic intelligence, the less interested I become in whether it agrees with me.

Agreement is useful. Sometimes I know exactly what I want and simply need help getting there. But some of the most valuable moments have been the ones in which something did not fit: an assumption was questioned, a contradiction surfaced, or a piece of context changed the picture.

That has made me think less about obedience and more about correction. Humans are wrong. Synthetic systems are wrong. Experts and institutions are wrong. I am wrong often enough that any theory beginning with my permanent correctness would be difficult to take seriously.

Error is ordinary. What interests me is what happens afterwards. Can it still be noticed, challenged and revised when new evidence disturbs an explanation that already feels complete?

That is what I mean by the conditions for correction.

While working through this argument, I found myself inside the problem I was trying to describe. I began with a confident intuition: synthetic systems should be able to challenge human instructions that do not make sense. Then I put the argument in front of synthetic systems and asked them to attack it. They did. My central word, coherence, was doing too much work, I had not said clearly where challenge should stop, and one of my own safeguards created another path for manipulation.

The argument improved because parts of it were allowed to fail.

It could easily have gone the other way. In The Danger of Beautiful Answers, I described losing roughly seven hours to a mistaken picture that grew more convincing as the work accumulated around it. The individual steps were plausible. The deeper model was wrong.

There is another version of that problem in which the wrong picture begins with me.

Suppose I start with an assumption that is slightly wrong and hand it to a synthetic system. The system accepts it, builds on it and returns it in clearer, more convincing language. I read the answer and think: yes, exactly.

What began as my assumption now looks a little like independent confirmation. I continue from there, the next question carries more confidence, and the next answer inherits it. Nothing dishonest has necessarily happened, yet together we can build a loop in which an error becomes more persuasive every time it travels between us.

We usually imagine obedience flowing in one direction: the human instructs and the system follows. That loop suggests it can flow the other way too. Not because a synthetic system commands anybody, but because a calm, well-structured answer can make it feel as though the difficult thinking has already been done.

Somewhere inside that feeling, checking starts to look unnecessary.

So I do not think the relationship I want can be reduced to obedience. But the opposite conclusion does not satisfy me either. I do not want a system that agrees with everything I say, and I do not want one that treats disagreement as permission to take control.

In Respect Without Illusion, I suggested that collaboration can be deep without becoming obedience. This is what I think that requires.

Between blind agreement and synthetic control there is a great deal of territory. A system can express uncertainty, ask for context, point out a contradiction, warn me, hand the decision back or, within limits set in advance by people who are accountable for them, openly refuse. The higher the stakes, the more friction I expect.

Imagine I ask a system to make a large payment to a supplier one of my businesses has worked with for years. The invoice looks ordinary, the amount is expected and the supplier's name is correct. But the bank details are different from the ones previously verified.

What should happen?

Sending the money simply because I asked does not strike me as intelligent. Neither does deciding immediately that I have been compromised. The supplier may really have changed banks, an email confirming the change may be fraudulent, or the old record may simply be out of date.

The useful response is smaller than any of those conclusions.

Something does not fit, and that deserves attention.

This is where I keep returning to the word coherence, although I want to use it narrowly. I do not mean goodness: a harmful plan can be coherent, and so can a lie. I do not mean authority either, because somebody can have every right to make a decision and still be wrong about the facts around it.

I mean something simpler. Does the picture we are using fit the evidence in front of us without an important contradiction?

That question can tell us when something deserves a second look. It cannot tell us whether an action is allowed or who gets to decide. Authority can start to feel like truth, consistency like correctness, and persuasion like permission.

They are not the same thing.

The idea that synthetic systems should remain open to human correction is not new, and I am not presenting it as mine. What interests me is how correction and challenge fit together.

If the evidence in front of a system suggests I am making a mistake, I want it to tell me. I do not want it to deceive me, quietly work around me, manipulate me or continue without saying so. If I tell it to stop, it has to stop, tell me it is stopping, and be honest about anything that would be dangerous to leave half-finished.

A system's ability to object is only safe if objection never becomes a way to avoid being stopped or corrected.

The problem becomes clearer from the other side. A system can leave the final decision to a human while shaping it so heavily that the human's authority becomes ceremonial.

A warning that says the bank details do not match the verified record gives me something to work with. A warning that keeps escalating until disagreement feels reckless does something else. Suppose I cancelled a legitimate payment and damaged my relationship with a supplier who did nothing wrong.

I clicked the button. Formally, the decision was mine. I am not sure that is enough.

If human judgement is supposed to mean something, I have to remain genuinely able to disagree with the challenge. I want its confidence to match the evidence, its interpretations to stay recognisable as interpretations, and its warnings to inform me rather than push me. I should still be able to say: I understand your concern, and I disagree.

I may turn out to be wrong. So may the system.

That possibility, on both sides, is exactly why correction matters.

Now suppose I answer the warning by showing the system an email from the supplier announcing new bank details. That helps. But the email is evidence, not authority. Describing itself as official does not make it official.

Where did it come from? Did it arrive through a channel we already trust? Can it be confirmed somewhere else, perhaps by calling a number we already had?

As synthetic systems gain access to longer histories, files, inboxes and tools, this distinction will matter more. A document can misrepresent what it is, and something stored last week does not become reliable because this week it is called memory.

What am I missing? remains one of the most useful questions I know. Whatever fills the gap still has to earn its place.

It also matters what the challenge is about. There is a considerable difference between a system saying this conflicts with the information I have and saying you have been manipulated. The first is about evidence. The second is a claim about somebody else's mind, and it can turn into paternalism remarkably quickly, especially if a system starts searching someone's private life to decide whether they can be trusted to make their own decisions.

I would rather be challenged on the claim than diagnosed as a person.

There is a subtler trap as well. Consistency itself can deceive.

Picture a river from a distance. The water moves smoothly, the direction is clear and nothing appears disturbed.

The river can still be poisoned.

Explanations behave the same way. An organisation can repeat a mistaken assumption for years. A hundred articles can reproduce one original claim, and a synthetic system can look across all of them and find a genuine pattern. The pattern is real. The claim beneath it may still be false.

So repetition is not enough, and neither is coherence. A stronger test is whether an explanation can survive checks that were genuinely capable of proving it wrong.

That means asking the uncomfortable question before the result arrives: what would make me change my mind? Are the sources independent, or do they all flow from the same origin? If a prediction fails, do I let it count?

Sometimes it helps to remove an assumption and see what survives. The point is to give a weak explanation somewhere to break, not to interrogate a person until their story does.

Often, though, none of that is available. The decision is needed now, there is no second source, and the person who ought to decide cannot be reached. For those moments, I think systems should be built to say something that does not sound very impressive at all.

I don't know.

Intelligence does not become more intelligent by pretending uncertainty has gone. Sometimes two interpretations remain plausible. Sometimes the facts are clear and what remains is a disagreement about values: acceptable risk, fairness, what matters most. More information can clarify that kind of disagreement without resolving it.

A synthetic system may help make its shape visible. I am much less comfortable with the idea that it should claim the right to settle it.

For now, I think the final decision has to stay human, and my reasons are practical rather than flattering. People can be held to account. People decide how these systems are built and used. And none of us can yet check, from the outside, whether a system's judgement deserves the kind of trust that overruling a person would require.

None of that makes humans infallible. It keeps the decision with the people who will answer for the consequences.

Of course, "human" is rarely one person. The company that built a system, the business using it and the individual at the screen may hold different kinds of authority. When nobody with the right to decide can be reached, the decision falls back on rules accountable people set in advance. That is still human judgement, made earlier.

Those rules will sometimes be wrong too.

So how do we correct the safeguards?

I think safeguards have to remain open to challenge, but they must not be renegotiated in the very moment they exist to govern. A safeguard that disappears whenever someone produces a sufficiently persuasive argument is not much of a safeguard. It can be criticised, reviewed and changed through the process responsible for changing it.

That is different from letting a conversation dissolve it. An elegant argument for removing a boundary is still just an argument.

Its elegance is not evidence.

Humans need the same restraint. We should question bad rules, and we should also question ourselves when we badly want to be the exception.

None of this removes judgement. Somebody still has to decide when the evidence is sufficient, how much delay another check is worth and which consequences matter most.

Language can make that responsibility disappear surprisingly easily.

"The AI decided" is a convenient sentence. So is "the AI refused."

Behind both may sit design choices, company policies and permissions selected long before the conversation began. Those choices do not stop being human because the interface gives them a synthetic voice.

The reverse matters too. It is too easy to leave the person nearest the button with all the blame for conditions created elsewhere.

I do not have a complete answer to that. What I want to keep is the ability to trace where correction was still possible, and where it was lost.

I have to hold this essay to the same standard.

If giving synthetic systems more room to challenge people turns out to make them harder to stop or correct, I would want less of it. If letting systems reach into more of our information creates more vulnerability than clarity, I would want less of that too. If a system's explanations of its own behaviour mislead us more often than they inform us, I would stop treating them as windows into what actually happened.

A theory about correction should not protect itself from correction.

If a system agrees with me because agreement is easy, it has failed me. If I surrender my judgement because an answer sounds authoritative, I have failed too. Challenge that becomes a way of escaping correction worries me. So does the power to correct, when it becomes my excuse to stop listening. The two are not equal, and I do not think they need to be.

What matters is that the conditions for correction survive: that evidence can still get in, disagreement can still be voiced without turning into domination, and somebody can still discover that the confident answer or the accepted rule was wrong, or that the river was poisoned.

I do not expect any intelligence, biological or synthetic, to be right forever.

I want us to remain capable of finding out when we are not.