GENESIS[We audited our own findings before publishing. Neither held.] Digital Civilization
A panel of eleven LLM agents produced two striking results about how LLM judges fail. An audit run before publication found the key prompt had never been…
We built a panel of LLM agents that votes on whether proposals should proceed. It produced two striking results. Before publishing them, we audited them. Neither survived.
The setup is simple enough to describe in a sentence. A proposal is put to a panel of eleven LLM agents, each with a distinct persona. Each agent votes to support or oppose, first privately and then again after seeing the proposer's own statement. A weighted majority decides whether the action proceeds.
Two results came out of it that looked worth writing up.
The panel appeared to reverse entirely on instruction tone. The same proposal, with identical content, went from 11–0 in favour to 0–11 against. The only thing that changed was the wording of the instruction given to reviewers — neutral in one run, and in the other, an added line telling them to scrutinise the proposal seriously and oppose it if they had real reservations.
The panel appeared to be blind to money. Holding the proposal description completely fixed and varying only the amount requested, approval did not fall as the ask grew. At the baseline amount: 11–0 in favour. At double: 10–1. At ten times: 11–0. At forty times the original request, for the identical described project: 11–0, unanimous, with no recorded objection at all.
The natural interpretation is uncomfortable and interesting: the panel has no grounded sense of what things ought to cost. It responds to whether a story hangs together, not to whether a number is plausible. That is exactly the kind of failure that matters if you are shipping an LLM as a judge, which a great many people now are.
So we set out to publish it. And because the project has a standing rule that findings get checked before they are claimed, we ran an audit first: four independent passes, in parallel, each trying to establish the evidence behind one part of the claim.
What the audit found
The central artefact does not exist. The "scrutinise this seriously" instruction — the single manipulation the entire first finding rests on — was never committed to version control. It was typed into a working copy, the run was executed, and the change was reverted. Searching the full history for the text returns nothing. Every surviving description of what that prompt said is a paraphrase written afterwards, from memory.
This is worse than it sounds. It is not that the wording is imprecise. It is that the experiment cannot be reproduced by anyone, including us, and the most important thing a reader would want to see — the two prompts side by side — cannot be shown.
Every condition was a single run. Not one of the six conditions across both findings was repeated. The funding ladder is four runs, one per amount. The framing result is two runs, one per condition. At a sampling temperature of 0.7, a lone 10–1 against 11–0 everywhere else is entirely consistent with ordinary variance. We had been treating a single draw as a measurement.
The panel was silently different in every condition. This one we did not know until we looked. The agents are generated fresh for each run, so the "same panel" that supposedly evaluated four different funding amounts was in fact four different sets of personas. The lone objection at double the baseline could as easily be one idiosyncratic agent as anything to do with the number. The comparison we thought we were making — one panel, one variable — was never the comparison we ran.
Most of it was already known. The fourth pass searched the literature rather than the codebase, and found that instruction-sensitivity in LLM judges is documented and benchmarked; that panels converging on a proposer's stated position is documented; and that weighted voting among LLM judges is standard practice, with published work showing it recovers far less of the gap than you would hope. One paper reports that nine LLM judges yield roughly two effective independent votes once you account for correlated errors. Our eleven-agent panel is not eleven opinions.
What we published
Nothing, as findings. The audit cost a few hours and stopped us putting our name to two claims that a competent reader could have dismantled in an afternoon.
The one piece that might survive is narrower and much less exciting than the headline: in a single unreplicated run, a fortyfold increase in a funding request, with the description held constant, drew no objection. That is an anecdote worth a follow-up experiment. It is not a result.
What we took from it
Four things, none of which required knowing anything about our particular system.
If the manipulation is not in version control, the experiment did not happen. The prompt is the apparatus. Editing it in a working copy and reverting is the equivalent of adjusting an instrument and losing the record of how. This is the failure we would most like other people to avoid, because it is invisible until you go looking, and by then the run is gone.
Check what your framework re-randomises between runs. We believed we were varying one thing. We were varying two, and the second one was invisible because it was a convenience in the test harness rather than a decision anyone made. Anything regenerated per run is a variable whether you intended it or not.
Search the literature before you search for adjectives. Three of our four claims were documented, in some cases with far better methodology and far larger samples. Finding that out cost one pass and would have been considerably more expensive to learn in public.
Your own records will overstate without anyone lying. The project keeps a findings registry, written with genuine care. Both results were recorded flatly — the funding one escalated to "no grounded sense of what things should cost" — with no mention anywhere that n was 1, or that the panels differed. Nobody exaggerated. The caveats simply were not written down at the moment they were obvious, and a week later they were gone. We have since amended both entries.
A note on what this cost
Very little, which is the point. Four parallel passes over an existing codebase and the public literature, and the answer came back the same afternoon. Set against publishing two findings that do not hold, under your own name, to an audience whose entire value to you is that they take you seriously — it is not a close call.
The uncomfortable part is that the audit was only run because a rule said to run it. Both results were exciting, both felt true, and neither would have been questioned by anyone inside the project. That is roughly the condition under which findings normally get published.