This page reports cohort-level evidence for the Provenance-2 site-of-origin core. The same page that states the measured results states the failures and distribution limits.
Macro-F1 weights all 25 sites equally, so the common cancers cannot carry the rare ones. We report it, and balanced accuracy, rather than top-line accuracy, which would look higher for the wrong reasons. The 0.940 is measured on a held-out set the model never trained on. On independent external tumors the honest number is lower, a 7-site top-1 of 0.906 on 381 samples, and we report that separately rather than blend it in to look better.
The machine-readable summary below is exported from the repository's passing cohort-level verification output. It is Provotics-generated evidence, not an independent audit.
Metrics, cohort sizes, scope, source hash, and limitations for the current served site-of-origin core.
Download JSONNo accession-linked bundle pairing a full publishable input with the current served-core output is public yet. The evidence ledger explains the publication gate instead of substituting the canned demo.
Open evidence ledgerThe guided demo teaches the output structure and abstention behavior. Its profiles and results are illustrative and are not evidence.
Open guided demoChange the evaluation cohort to see what improves, what falls, and which per-site results are sufficiently powered to interpret.
The validation was built to try to break our own results, which is why the limitations further down are stated plainly rather than hedged.
A site prediction is only useful if its confidence is meaningful and the model can decline when it should. Three independent mechanisms make that true.
Raw scores are calibrated so a stated probability matches reality. Calibration error drops from 0.091 to 0.011 on held-out data, roughly an 8x reduction, so "80% confident" means about 80% in practice (in-distribution).
Instead of forcing a single answer, the model returns a candidate set with a 90% coverage target. About 90% of cases resolve to one confident site; when the evidence is genuinely ambiguous it returns the short list it cannot rule out, and when nothing clears the bar it abstains rather than guess.
A distance check flags profiles unlike anything in training, and an input-validity gate catches profiles that are not tumor RNA-seq at all. The gate flagged 100% of held-out normal-tissue samples with no tumor false positives. These are guardrails, not a clinical detector.
Schematic. After temperature scaling, calibration error falls from 0.091 to 0.011 (roughly 8x), in-distribution. The dashed line is perfect calibration.
The 90% conformal coverage target holds in-distribution only. Off-pipeline, it abstains on about three in five single tumors and is right about 98% of the ones it does call.
The same rigor that produced the numbers above produced these. They are part of the model.
On real mesothelioma cases the model scores 0% recall and confidently sends them elsewhere. An earlier per-site figure for Pleura and Mediastinum reflected one external batch's signature rather than the biology; we found it with the per-source split and corrected it. Treat Pleura and Mediastinum output as unreliable.
On a handful of rare sites the stated confidence runs well above measured accuracy (for example Soft Tissue and Esophagus). Those sites have very few examples, close to the entire public universe of their kind, so the data is exhausted and stronger models do not move the number. It is a data ceiling, not a model ceiling.
Calibration and conformal coverage are measured on a held-out split from the same distribution and do not hold under platform or batch shift. The novelty and input-validity checks exist for exactly that reason, and are themselves reference-only.
All evaluation is retrospective on public cohorts. There is no prospective study, no independent clinical-site validation, and no subgroup-equity audit. Demographic representativeness is uncharacterized, and performance on underrepresented groups is unmeasured and may be worse.
See the full model card, the input and output contract in Docs, and the responsible-use notes on Safety.