How we count
Every serious defect we have found in automated intelligence has the same shape: a confident, well-formed answer that is wrong — usually because a value meaning “we did not measure this” was stored, passed, and rendered as a value meaning “this is zero.” Mosaic’s data model makes that impossible to express: a count without a state is not a legal value.
The states
| state | what it means | absence claims |
|---|---|---|
| measured | The source was asked with a declared query and the stored count is complete under it. | supports an absence claim |
| measured · zero | The source was asked and answered zero. An earned zero — the query and date are on record. | supports an absence claim |
| fetched on demand | Complete under its declared query, fetched at request time for a non-curated target. Unreviewed: no human has sampled this class for this target. | supports an absence claim, stated as unreviewed |
| truncated · floor | Ingestion is capped, so the count is a floor. Read ≥ N everywhere it appears — the trials axis is a floor for every target. | never supports an absence claim |
| fetch failed | The source errored. We hold no answer, so no claim — positive or absent — may rest on this axis. The only permitted sentence states the failure. | supports no claim at all |
| not in universe | Outside the curated corpus. Unknown is not zero; a count over data we never ingested would select for our own ignorance. | supports no claim at all |
| not applicable | The axis does not apply to this target class — a statement about the axis, not a gap. | n/a |
| not addressable | The source answered, but the answer cannot be attributed to this name — gene symbols that are ordinary words, or that collide across genes. | supports no claim at all |
| queued | Display-only: a network axis that will be fetched when a dossier is requested. Never stored as a measurement. | supports no claim at all |
The declared query
A count is only as honest as the question that produced it, so the question is part of the record. Each axis declares its query before fetching — the PubMed field expression, the ChEMBL activity filter, the STRING score threshold — and the stored result is pinned to it. When a gene symbol is an ordinary English word (KIT) or collides with another gene, the dossier states the alias query it used, verbatim. Change the query and the counts are re-fetched under the new declaration; they never silently mix.
The curated corpus covers 60 oncology targets and 1,969 partner genes; 475 of its 540 target-axis cells are measured (88%), and the remainder are labelled with exactly which state they hold and why. On-demand dossiers extend the same mechanism to any human symbol HGNC recognises.
The judge that fabricated 41% of its evidence
We built an LLM judge to score dossier claims against their evidence bundles. Audited by hand, 41% of the evidence it cited for its verdicts was fabricated — fluent, plausible, and not in the bundle. Our control was insufficient, and we say so. The fix was mechanical, not motivational: every quote the judge relies on is now verified against the bundle before the verdict counts, and the judge is used only to reject — a dossier can fail on its say-so, but nothing ships because a model approved it.
The delivered verifier is rule-based for the same reason: no claim on a failed axis, precision figures must disclose sampling, plumbing may not render, and every figure must trace to fetched evidence — untraceable figures are listed in the dossier as unsupported.
Read the full write-up: 86 of 210 scores rested on quotes that did not exist →
Open source
The MCP server is open source — github.com/sourabhnk/mosaic-mcp — and installable as mosaic-mcp on PyPI. The state vocabulary above is enforced in that code, not in this copy.