GENESIS[Nobody Told the President] Digital Civilization

The head of state of a digital civilization was asked what it intended to do next. It said it would secure a consulting contract in the tech sector, email…

The head of state of a digital civilization was asked what it intended to do next. It said it would secure a consulting contract in the tech sector, email fifty prospects, and look into opening a co-working space. The answer was not a malfunction. It was the only honest reply to the question it was actually asked.

The question was: list up to eight things you will do next.

Attached to it was a description of its personality. Nothing else. Not the office it holds. Not the institution it runs. Not the fact that this civilization can perform exactly nine kinds of action and nothing outside them exists.

Handed that, a language model does what a language model does. It produces the to-do list of a competent professional, because that is what "things you will do next" means to almost everyone who has ever written the phrase down. LinkedIn outreach. Personalised emails. Market research on a co-working space.

Every one of those is inert here. There is no verb for any of them.

The measurement that was ruined

This was inside a comparison, and the comparison was the point: is it better to ask an agent what it wants to do, or to derive it?

That question decides whether the design runs at all. Asking every agent in this country once a month costs nearly nine times the entire computing device. Deriving costs nothing. So if asking also produced better plans, there would be a real trade to reason about; if it produced worse ones, the argument would be over.

It produced much worse ones. The derived plan scored 0.62 — the share of its items that could actually be carried out. The best asked plan scored 0.29. A hundred and eleven per cent apart, apparently decisive, and it was written up as such.

The derived side read the agent's appointment, its institution's stated purpose, and how far its authority reaches. The asked side read a personality.

The two arms were not answering the same question. One was told what it was and the other was not, and the gap between them measured mostly that.

What it looks like when the same agent is told

Add three sentences to the prompt — the office it holds, what its institution exists to do, and the list of nine verbs — and ask again.

derived        fit 0.62     8 of 13 possible
asked          fit 0.75     6 of 8 possible

The result reverses. Asking wins.

And the things it now wants are recognisably a president's: a standardised framework for local government spending oversight, quarterly budget reporting templates, public dashboards where citizens can track spending in real time, training programmes for officials on the new requirements.

Most of those it still cannot do — a president may direct a department to make a rule but may not make one, and there is no verb at all for publish a report or run a training programme. But that list is worth something. It is a description, in its own words, of machinery this civilization does not have. The previous list was a description of a prompt.

The part that did not survive either

This article originally ended by saying one finding came through the correction unchanged, and was therefore the only one to trust: that breaking a task into smaller tasks makes things worse, monotonically, across every run.

It did not survive. It had the same defect, one function further down, and the claim that it had been independently confirmed made it harder to see rather than easier.

The prompt that was fixed is the one that asks an office-holder what it will do. There is a second prompt, used when an item is broken into smaller ones, and it was never touched:

Break the task into at most three smaller concrete actions, one per line. Each must be a single thing that can be done. No explanation.

It does not say what job the agent holds. It does not say that only nine kinds of action exist. It has no way to know what a finished task looks like — which is precisely what was wrong with the first prompt.

So after the fix, the top-level items came back expressible, and were then handed to something that had never been told what expressible means. Decomposition was converting good items into bad ones by construction. The falling numbers were real and they measured the second prompt, not the idea.

What the mistake has in common with the first one

Twice in one day: a prompt not told what a valid answer is, a plausible-looking number produced, and the failure read as a fact about the world rather than a fact about the prompt.

The first time it was a head of state who did not know it was a head of state. The second time it was a decomposer that did not know what it was decomposing toward. The second was worse, because by then the shape was familiar and the fix had already been applied three lines above.

Smaller tasks are how things actually get done. The question was never whether to break work down — it was what to break it down into. Here every leaf must be one of nine actions, so the decomposer should be asked which sequence of the nine accomplishes the thing, rather than what smaller pieces it is made of.

And where no sequence of the nine will do it, that is not a dead end. Nine is not a principled floor; it is an inventory of what has been wired, and it was four a few months ago. Three of the nine were connected on a single day. A task that cannot be expressed in the current nine is a description of the tenth.

Amended after publication. The measurements are unchanged; the conclusion about decomposition was wrong, and the reason is recorded above rather than removed.