What the model is, what it is for, what it is not, how it is measured, and the places it is weak. Written to be checked, not admired.
Provenance-2 reads one bulk tumor RNA-seq expression profile and estimates the body site the tumor came from, across 25 anatomical sites. It is built to be honest about uncertainty: every site call carries a calibrated probability, a candidate set that lets the model abstain when the evidence is ambiguous, and an out-of-distribution check that flags inputs unlike anything it was trained on.
Reflects the served core verified in the repository in July 2026. Numbers trace to the internal model card and reproducibility harness; see Validation for methodology and the public evidence ledger for what can and cannot currently be downloaded.
Research and education: exploring what a tumor's transcriptome reveals about its tissue of origin, studying calibrated uncertainty, and generating hypotheses on cohorts you already have. It reads expression values only, never identifiable patient data.
Not for clinical, diagnostic, prognostic, or treatment-selection use. It predicts an anatomical site, not a histological diagnosis, stage, or grade. Inputs from other assays (single-cell, microarray, targeted panels), other normalizations, or non-tumor tissue are out of distribution and should not be trusted.
A profile is harmonized to a common reference, mapped onto the fixed 3,882-gene panel, and scored by the calibrated ensemble. The output is a site probability for research interpretation, a conformal candidate set, and a novelty check. Cross-pipeline comparability remains limited and the in-distribution uncertainty guarantees do not transfer under platform shift. Model internals, panel contents, and ensemble details are not publicly disclosed.
We lead with the number measured on held-out patients the model never trained on, not the prettiest one. On independent external tumors the honest figure is lower, a 7-site top-1 of 0.906 on 381 samples, and we report it separately rather than blend it in.
A pool of 17,410 retrospective tumor RNA-seq profiles drawn from open, public genomic cohorts. All are public and retrospective and contain expression values, not protected health information. The exact composition of the training pool is not published; the figure is the training pool, not a single held-out test set.
These are stated with confidence because the validation was deliberately adversarial. Read them as part of the model, not a disclaimer.
On real mesothelioma cases the model scores 0% recall and confidently misroutes them. An earlier per-site figure for Pleura and Mediastinum turned out to reflect one external batch's signature, not the biology, and we corrected it. Treat any Pleura and Mediastinum output as unreliable.
Pleura and Mediastinum, Thymus, Esophagus, Skin, and Eye have very few examples (on the order of seventeen each in the relevant held-out evaluation), so their per-site metrics are statistically fragile. These tissues are close to the entire public universe of their kind, so the data is exhausted at source, and stronger models do not move the number.
A single tumor sequenced on a different pipeline is the hardest case. Rather than guess, the model abstains on most of those and commits only when it is confident. Inputs that skip the harmonization step will be misclassified.
Calibration and conformal coverage are measured on a held-out split from the same distribution. They do not hold under platform or batch shift. The novelty and input-validity checks exist precisely because of this, and are themselves reference-only, not a validated clinical detector.
All evaluation is retrospective on public cohorts. There is no prospective study, no independent clinical-site validation, and no subgroup-equity audit. The training cohorts are not characterized here for demographic balance, and performance on underrepresented groups is unmeasured and may be worse.
Provenance-2 is invite-only during the research phase and granted under a confidentiality agreement. See how the flagship core is validated, check the public evidence status, read Safety, then apply for access.