Mosaic

How we count / Write-up

We built an LLM judge for scientific dossiers. It fabricated 41% of its evidence.

August 2026 · the run records are d_paired.json, d_paired_debiased.json and d_controls*.json in our working repo; every number below comes from them.

We generate target dossiers — structured intelligence documents about drug targets — and we wanted to know whether they were better than what a frontier model produces on its own. So we did the obvious thing: we built an LLM judge.

We did not build it carelessly. The judge was blind (documents labelled A and B, no product names), paired (always a head-to-head, never a lone score), multi-run (three runs per pair), and quote-backed: every score the judge emitted had to cite a verbatim supporting quote from the document it was scoring. The judge was Claude Sonnet 4.5. The comparison was seven oncology targets, our dossier against a no-tools baseline from the same model — a bias that favours us, and we noted it in the artefact.

Before trusting it with the real comparison, we ran a control: a real dossier against a mechanically degraded copy of itself. The judge preferred the real document in both orientations, 15–0 and 12–0, with zero quotes discarded. It passed. Twice.

210

scores the judge emitted

both-orientations run · 7 pairs × 2 × 3 runs × 5 parameters

86

supporting quotes in neither document

exact substring check, whitespace/case normalised

41%

of its evidence, fabricated

86 ÷ 210 — the title of this piece

source: d_paired_debiased.json · totals counted from the file

The control was insufficient, and that is on us

A real document against a gutted copy of itself is an easy discrimination — the gap is so large that content swamps every other signal. Passing it says the judge can detect a wreck. It says nothing about whether the judge can separate two plausible documents, which was the entire job. A control that only works when the answer is obvious is not much of a control. We designed it, it passed, and it validated nothing we needed. We say this first because everything that follows was found despite the control, not because of it.

The judge picked whichever document it read first

The first paired run scored our dossiers 27 and the baseline 45. Before writing that down as a loss, we noticed something: whichever document sat in slot A won five of seven pairs, and the baseline won every pair where it happened to be A.

So we re-judged every pair in both orientations and counted a win only when the same document won both ways. Six of seven comparisons flipped when the documents swapped places. Across the fourteen orientations, slot A won eleven, tied two, and lost one. Exactly one target produced a consistent verdict in both directions.

At the difficulty that matters — two fluent, plausible documents — the judge was not reading content. It was reading position. Neither the 27 nor the 45 means anything, in either direction, and we report no quality verdict from this run.

pairMosaic read firstbaseline read firstverdict
BRD41st doc wontieflips on swap
EGFR1st doc wontieflips on swap
EZH21st doc won1st doc wonflips on swap
FLT31st doc won1st doc wonflips on swap
KRAS1st doc won1st doc wonflips on swap
PDCD11st doc won2nd doc wonconsistent
RET1st doc won1st doc wonflips on swap
whichever document was read first won 11 of 14 orientations (2 ties, 1 loss) · 6 of 7 pairs flip when the documents swap places

86 of 210 scores rested on quotes that did not exist

The quote requirement was the one piece of the harness that earned its keep. Every score had to carry a verbatim quote from the document being scored, and the harness checked each quote against the document — an exact substring match after normalising whitespace and case. Not a semantic similarity check. Presence.

That both-orientations run emitted 210 scores. 86 of them — 41% — were discarded because the supporting quote did not appear in either document. Not mangled paraphrases — text that was not there. The judge, asked to ground its verdicts in evidence, invented the evidence, fluently, in well-formed JSON, with confident rationales attached. (The abandoned single-orientation run was no better: 33 of its 105 scores failed the same check.)

This is the failure shape we had spent months engineering out of our data pipeline: a confident, well-formed answer that is wrong. We had written that sentence into our engineering doctrine long before this run. What this run taught us is that the instrument you build to measure quality has exactly the same failure mode as the thing it measures — and if the harness had not verified quotes mechanically, we would have had 210 confident scores and no way to know that 86 rested on nothing.

One suggestive aside: in a smaller follow-up on three targets where the judged documents were denser with verifiable identifiers, fabrication fell from 41% to 13%. That is n=3 — a signal, not a result — but it points somewhere useful: models fabricate less when the document offers real text worth quoting. It does not point anywhere near zero.

quote found in the document · 124quote in neither document · 86
one square per score, in emission order by pair · 86 / 210 = 41%

The fix was mechanical, not motivational

We did not prompt the judge to be more honest. Prompting a model to not fabricate is asking the failure mode to police itself. The fixes were structural:

Quote verification is not optional. Any LLM output that claims grounding in a document gets its quotes checked against the document, by string matching, before the output counts. A score whose evidence fails the check is discarded, loudly, and the discard rate is itself reported. The 41% is only knowable because discards were counted, not silently dropped.

The judge was demoted to rejection-only. The limit line we had stamped on every judge artefact — written before this run, as a design principle — turned out to be exactly right:

An LLM judge rewards fluency and cannot determine that a claim is scientifically wrong. Use to reject, never to approve.

A judge that flags a gutted document is useful. A judge whose approval ships a document is an approval you cannot trust, because you now know what its approvals are made of. In our pipeline today, nothing ships because a model liked it.

Delivered documents are verified by rules, not by a model. Every dossier we deliver passes a rule-based verifier: no claim may rest on an axis whose fetch failed; precision figures must disclose their sampling; pipeline plumbing may not render into prose; and every figure in the document must trace to fetched evidence — figures that cannot be traced are listed in the delivered dossier as unsupported, rather than silently removed. Rules are dumber than a judge. They are also incapable of inventing a quote.

What we would tell anyone building an eval

Three things, none of which we followed until the data made us.

First, your control has to be as hard as your task. A judge validated on an easy discrimination is unvalidated.

Second, run both orientations, always. Position bias did not show up as noise; it showed up as a coherent-looking 27–45 result that would have survived any single-orientation analysis and been quoted forever after.

Third, verify grounding mechanically or assume it is not there. The single cheapest component of our harness — a substring check — was the difference between measuring our judge and being fooled by it. If your eval trusts a model's citations without checking them against the source, your fabrication rate is not zero. It is unknown.

We built the judge to find out whether our documents were good. It could not tell us that. It told us something more useful: exactly how an evaluation fails while looking like it is working — the same confident, well-formed wrongness we built this product to keep out of scientific answers, arriving inside the tool built to measure them.

Use it to reject. Never to approve.


Mosaic delivers coverage-honest target dossiers over MCP — every count carries a state, and every state carries the query that produced it. How we count · github.com/sourabhnk/mosaic-mcp