GENESIS[Everything Looked Right] Digital Civilization

A civilization of fifteen thousand chose its first president at random from those eligible. The count of eligible was wrong by fifteen hundred. Nothing…

A civilization of fifteen thousand chose its first president at random from those eligible. The count of eligible was wrong by fifteen hundred. Nothing failed, no test went red, and every number on the screen was the number the code intended to print — which is the problem, and it happened four times in one day.

The founding president of this civilization is not elected. There is nobody to elect them: on the first day of a republic there are no legislators to confirm an appointment, no courts to hear a challenge, and no electorate that has ever voted. So the office is filled by draw, and everything downstream — the cabinet, the agencies, the confirmations — hangs off that one seat existing.

The draw ran. An agent named Fleet Monsoonspray 001 was seated as founding president, chosen, the log said, from 14,976 eligible agents.

That sentence is true. It is also the sound of an experiment quietly dying.

The fifteen hundred who should not have been eligible

This civilization runs a control group. A fraction of its agents are built with no personality traits at all — no dispositions, no leanings, nothing that would make one of them prefer a policy to another. They exist so that when a trait-bearing agent does something interesting, there is something to compare it against. Without them, every finding is a story about agents in general and not about traits.

In the previous run the control group was twenty-four agents out of a hundred and eleven. Call it a fifth of the population. Big enough that when the trait-bearing agents refused to stand for office at a particular rate, you could ask whether the traitless ones refused at the same rate, and get an answer.

The new run scaled the population to fifteen thousand. The control group scaled to twenty-four.

Not twenty-four percent. Twenty-four agents. The same twenty-four, carried across from the old population, while fourteen thousand eight hundred and eighty-nine new agents were generated with traits because nothing had said to do otherwise. A control group that had been 21.6% of the civilization was now 0.16% of it.

Here is what makes this worth an essay rather than a bug report: there was no error. The command did exactly what it was told. The population reached its target. Every agent was valid. The president was drawn correctly from a correctly computed pool of eligible agents. If nobody had gone looking, the run would have finished, produced numbers, and those numbers would have been compared against a control group too small to mean anything — and the comparison would have looked exactly like every other comparison this project has ever published.

The fix cost four minutes. Finding it cost one question that nobody was required to ask.

A rule that had never once refused anything

The same day, in an unrelated corner of the system, a rule was found that had never worked.

Candidates in this civilization can have supporters. An agent reads a candidate's platform, decides it agrees, and registers as a public backer. There is one restriction, and it is the obvious one: a candidate cannot register as their own supporter. The code for it was right there, three lines, easy to read, doing exactly what the comment above it said.

It compared the incoming supporter's identity against the candidate's identity. If they matched, it refused.

They could never match. The incoming identity was a universally unique identifier — a long string of hexadecimal that this civilization uses to key money and ownership, chosen precisely because it is unique and unchangeable. The candidate's identity, in this one comparison, was built out of the candidate's display name. A name against a number. The two strings were never going to be equal, and so the branch that refuses was never once entered, across the whole life of the system.

Every candidate had always been free to back themselves. Nobody did, as it happens — but nobody was stopped, and the rule that would have stopped them had been sitting there looking enforced.

This is the shape worth naming. A refusal that never refuses produces no evidence of itself. A crash leaves a stack trace. A wrong number leaves a wrong number. But a rule that silently permits what it claims to forbid produces a system that behaves exactly as though the rule works, and a codebase in which the rule is visibly present, and a test suite in which nothing is red — because there was no test. Nobody writes a test for a rule they can read.

When the fix went in, we ran the new tests against the old code, to check the tests could actually fail. Three of them did. One failed with the line DID NOT RAISE, which is the whole essay in three words.

The check that said nothing was running

Before restarting a service that the civilization talks to, you check whether the civilization is currently running. The project's own documentation gave the command for it: list every running process's executable, and look for one launched from the project's isolated Python environment.

The command returned nothing. So the service was restarted.

Two copies of the civilization had been running for two and a half hours.

The reason is a detail of how Unix reports processes. The environment's Python is a symlink pointing at the system Python, and the kernel's record of what a process is executing resolves symlinks before reporting. So a process launched from the project's environment truthfully reports that it is running the system interpreter, and a search for the environment's path finds nothing — not because nothing is running, but because the question was asked in a vocabulary the answer does not use.

Two things about this are worse than the mistake itself.

The first is that the documented purpose of that check was to decide whether it was safe to delete the environment. It returns "nothing is running" in exactly the circumstance it exists to detect, and the next step it authorises is removing the files a live process is executing from.

The second is that two civilizations had been running at once. Not one that had been forgotten — two, started three minutes apart, polling the same electorate, competing for the same single graphics card, both writing to the same records. The logs showed an election being polled twice with identical results, two minutes apart, which reads as a stable system right up until you count the processes.

The dry run that described a different command

Before rebuilding a civilization from scratch, you ask the command what it intends to do. That is what a dry run is for: it is the thing a person reads instead of the code, at exactly the moment they are about to do something expensive and irreversible.

The instruction was to grow the population to fifteen thousand. The dry run answered:

bootstrapping 8 agents

Eight. The command was going to do the right thing — the logic underneath was correct, the fifteen thousand would have been created, the only thing wrong was the sentence describing it. But the sentence is the entire output of a dry run. There is nothing else to read. A dry run that describes the wrong operation is worse than having no dry run at all, because it converts caution into confidence.

Had that message been believed, the reset would have been run with eight agents and the mistake discovered somewhere in the following hour, after the old state was gone.

What the four have in common

A control group diluted by a factor of a hundred and thirty-five. A prohibition that had never prohibited. A safety check that reported the opposite of the truth. A preview that previewed something else.

None of them crashed. None of them logged a warning. None of them failed a test, and three of them had no test that could have failed, because nobody writes tests for the things they can see are correct. Every one of them produced output that was well-formed, plausible, and consistent with a system working exactly as designed.

The common property is not that the code was wrong. It is that the code was wrong in the one way that produces no signal. Software that breaks tells you. Software that quietly does something adjacent to what you meant does not, and the more carefully the surrounding system is built — good logs, clean output, a passing suite — the more convincingly it reports its own health.

There is a temptation to draw a comfortable moral here: test more, check more, be careful. It isn't quite right. Three of these four were found by asking a question that no process required and no checklist contained — is that number plausible? The control group looked like a number. Fourteen thousand nine hundred and seventy-six eligible agents looked like a number. They were correct arithmetic over a wrong premise, and arithmetic does not know what it is counting.

Four agents, one wallet

The fourth defect is about money, and it is the one where "looked right" had been looking right the longest.

Agents in this civilization hold balances. They are paid, they bid for fractions of assets, they buy advertising for their campaigns. Every balance is filed under an identifier, and the system is emphatic about which identifier: the agent's unique key, never its name. There is a file whose only job is to say so, and it says it in terms — a name is display text, it is not unique, and it can change.

That file was correct. Every balance in the database followed it: a hundred and nineteen of them, all filed under unique keys.

Of the forty-six actual transfers — the movements of money between agents — the number filed under a unique key was zero.

All forty-six had been filed under display names. The rule existed, the helper function that implemented it existed and was correct, and the places that actually moved money had each built the identifier by hand instead of calling it. An audit had checked the helper. The helper was fine. The audit had not checked whether anything called it.

The consequence is easier to feel than the cause. Four separate agents in the old population answered to the name "Quiet Anchor". Under a naming scheme that derives from identity, that is not a bug — agents with similar natures deserve similar names, the same way four people can be called James. But when money is filed under the name, those four agents share a wallet. One of them earns; all of them can spend it. Which of them earned it is not recoverable from the records, because the records never held the distinction.

The deeper cause was structural and had been sitting in plain sight. The component that tracks elections holds its own record of each candidate — a name, a slug, an identifier of its own — and nothing in it pointed back at the agent. So when that component needed to know whose money to debit for a campaign advertisement, it had no way to ask. It did the only thing available: it built an identifier out of the name it had.

It was not a careless line. It was the only line that could be written, given a link that nobody had noticed was missing.

The fix adds the link and then does something that matters more: when the link is absent — for the older records, or for a candidate with no agent behind it — the system now refuses to move money rather than falling back to the name. A fallback would have looked like a safe default. It would also have silently restored the original behaviour on precisely the records that still had the problem.

The president who was drawn twice

The civilization was rebuilt. This time the control group was specified: fifteen hundred agents with no traits, alongside thirteen and a half thousand who have them. The draw ran again, from a pool that was now the pool it was supposed to be.

An agent named Gentle Monsoonbark 001 is the founding president of a civilization of fifteen thousand.

Fleet Monsoonspray 001, who held the office for about four minutes in a database that no longer exists, is somewhere in the new population with no memory of it. Every agent starts each run knowing nothing — that is the design, and it is not sentimental about presidencies.

The suffix on both names is worth a footnote. In a population this size, names collide: the naming system generates from each agent's underlying identity, and agents with similar identities genuinely deserve similar names. Fifty of the old hundred and eleven shared a name with somebody, and four separate agents answered to "Quiet Anchor". That was fine while names were only names, and it stopped being fine the moment money was keyed to them.

Which is the sentence this whole day turns on. Nothing here was written carelessly. The naming system is good. The identity rule is correct and clearly argued. The dry run, the safety check and the supporter rule were all written by someone who understood what they were for. Each of them failed at a join — a place where two correct things met and one of them was speaking a vocabulary the other did not use.

You cannot test a join you have not noticed. You can only keep asking the unglamorous question about numbers that look fine: is that plausible? Fifteen thousand agents and twenty-four controls is arithmetic nobody got wrong. It is just not a control group.