Mosaic

How we count

Every serious defect we have found in automated intelligence has the same shape: a confident, well-formed answer that is wrong — usually because a value meaning “we did not measure this” was stored, passed, and rendered as a value meaning “this is zero.” Mosaic’s data model makes that impossible to express: a count without a state is not a legal value.

The states

statewhat it meansabsence claims
measuredThe source was asked with a declared query and the stored count is complete under it.supports an absence claim
measured · zeroThe source was asked and answered zero. An earned zero — the query and date are on record.supports an absence claim
fetched on demandComplete under its declared query, fetched at request time for a non-curated target. Unreviewed: no human has sampled this class for this target.supports an absence claim, stated as unreviewed
truncated · floorIngestion is capped, so the count is a floor. Read ≥ N everywhere it appears — the trials axis is a floor for every target.never supports an absence claim
fetch failedThe source errored. We hold no answer, so no claim — positive or absent — may rest on this axis. The only permitted sentence states the failure.supports no claim at all
not in universeOutside the curated corpus. Unknown is not zero; a count over data we never ingested would select for our own ignorance.supports no claim at all
not applicableThe axis does not apply to this target class — a statement about the axis, not a gap.n/a
not addressableThe source answered, but the answer cannot be attributed to this name — gene symbols that are ordinary words, or that collide across genes.supports no claim at all
queuedDisplay-only: a network axis that will be fetched when a dossier is requested. Never stored as a measurement.supports no claim at all

The declared query

A count is only as honest as the question that produced it, so the question is part of the record. Each axis declares its query before fetching — the PubMed field expression, the ChEMBL activity filter, the STRING score threshold — and the stored result is pinned to it. When a gene symbol is an ordinary English word (KIT) or collides with another gene, the dossier states the alias query it used, verbatim. Change the query and the counts are re-fetched under the new declaration; they never silently mix.

The curated corpus covers 60 oncology targets and 1,969 partner genes; 475 of its 540 target-axis cells are measured (88%), and the remainder are labelled with exactly which state they hold and why. On-demand dossiers extend the same mechanism to any human symbol HGNC recognises.

The judge that fabricated 41% of its evidence

We built an LLM judge to score dossier claims against their evidence bundles. Audited by hand, 41% of the evidence it cited for its verdicts was fabricated — fluent, plausible, and not in the bundle. Our control was insufficient, and we say so. The fix was mechanical, not motivational: every quote the judge relies on is now verified against the bundle before the verdict counts, and the judge is used only to reject — a dossier can fail on its say-so, but nothing ships because a model approved it.

The delivered verifier is rule-based for the same reason: no claim on a failed axis, precision figures must disclose sampling, plumbing may not render, and every figure must trace to fetched evidence — untraceable figures are listed in the dossier as unsupported.

Read the full write-up: 86 of 210 scores rested on quotes that did not exist →

Open source

The MCP server is open source — github.com/sourabhnk/mosaic-mcp — and installable as mosaic-mcp on PyPI. The state vocabulary above is enforced in that code, not in this copy.

See it applied: the dossiers →