We built an LLM judge for scientific dossiers. It fabricated 41% of its evidence.
August 2026 · the run records are d_paired.json, d_paired_debiased.json and d_controls*.json in our working repo; every number below comes from them.
We generate target dossiers — structured intelligence documents about drug targets — and we wanted to know whether they were better than what a frontier model produces on its own. So we did the obvious thing: we built an LLM judge.
We did not build it carelessly. The judge was blind (documents labelled A and B, no product names), paired (always a head-to-head, never a lone score), multi-run (three runs per pair), and quote-backed: every score the judge emitted had to cite a verbatim supporting quote from the document it was scoring. The judge was Claude Sonnet 4.5. The comparison was seven oncology targets, our dossier against a no-tools baseline from the same model — a bias that favours us, and we noted it in the artefact.
Before trusting it with the real comparison, we ran a control: a real dossier against a mechanically degraded copy of itself. The judge preferred the real document in both orientations, 15–0 and 12–0, with zero quotes discarded. It passed. Twice.