15% of all profit is donated to Heal Palestine and the Palestine Children's Relief Fund (PCRF)
Platform Research Validation Safety Pricing Company Demo Request access
Flagship model card

Provenance-2, on the record.

What the model is, what it is for, what it is not, how it is measured, and the places it is weak. Written to be checked, not admired.

Research and educational use only. Provenance-2 is not a medical device. It has not been reviewed or cleared by any regulator, is not CLIA or CAP validated, and must not be used to make decisions about a real patient. Outputs are research artifacts from a model evaluated retrospectively on public cohorts, not medical findings. A confident site prediction is not a cancer diagnosis.
Overview

What Provenance-2 is

Provenance-2 reads one bulk tumor RNA-seq expression profile and estimates the body site the tumor came from, across 25 anatomical sites. It is built to be honest about uncertainty: every site call carries a calibrated probability, a candidate set that lets the model abstain when the evidence is ambiguous, and an out-of-distribution check that flags inputs unlike anything it was trained on.

At a glance

The model in one table

Model
Provenance-2 (Provotics' flagship 25-site site-of-origin core)
Task
Predict 1 of 25 anatomical sites of origin from tumor RNA-seq
Input
One bulk RNA-seq expression profile (gene-level)
Output
A calibrated site of origin, a 90% conformal candidate set, and a novelty flag
Gene panel
A fixed 3,882-gene panel
Training data
17,410 retrospective public tumor profiles (see Training data)
Held-out accuracy
macro-F1 0.940 on the held-out test set (balanced accuracy 0.942)
Status
Research phase, invite-only access
Use
Research and educational only, not a medical device

Reflects the served core verified in the repository in July 2026. Numbers trace to the internal model card and reproducibility harness; see Validation for methodology and the public evidence ledger for what can and cannot currently be downloaded.

Intended use

Intended use, and what is out of scope

What it is for

Research and education: exploring what a tumor's transcriptome reveals about its tissue of origin, studying calibrated uncertainty, and generating hypotheses on cohorts you already have. It reads expression values only, never identifiable patient data.

Out of scope

Not for clinical, diagnostic, prognostic, or treatment-selection use. It predicts an anatomical site, not a histological diagnosis, stage, or grade. Inputs from other assays (single-cell, microarray, targeted panels), other normalizations, or non-tumor tissue are out of distribution and should not be trusted.

How it reads a profile

From input to a site-of-origin result

A profile is harmonized to a common reference, mapped onto the fixed 3,882-gene panel, and scored by the calibrated ensemble. The output is a site probability for research interpretation, a conformal candidate set, and a novelty check. Cross-pipeline comparability remains limited and the in-distribution uncertainty guarantees do not transfer under platform shift. Model internals, panel contents, and ensemble details are not publicly disclosed.

Evaluation

Measured on held-out tumors first

We lead with the number measured on held-out patients the model never trained on, not the prettiest one. On independent external tumors the honest figure is lower, a 7-site top-1 of 0.906 on 381 samples, and we report it separately rather than blend it in.

0.940macro-F1 on the held-out test set (balanced accuracy 0.942), scored evenly across all 25 sites
3 in 5off-pipeline single tumors it abstains on rather than guess; on the ones it does call, it is right about 98% of the time
0.091 → 0.011calibration error, cut about 8x, so a stated confidence means what it says
90%conformal coverage on in-distribution input; about 90% of cases resolve to a single confident call, the rest return a candidate set or abstain. Coverage does not hold off-pipeline
Training data

Public, retrospective, expression only

A pool of 17,410 retrospective tumor RNA-seq profiles drawn from open, public genomic cohorts. All are public and retrospective and contain expression values, not protected health information. The exact composition of the training pool is not published; the figure is the training pool, not a single held-out test set.

Limitations

Where it is weak

These are stated with confidence because the validation was deliberately adversarial. Read them as part of the model, not a disclaimer.

It does not recognize real mesothelioma

On real mesothelioma cases the model scores 0% recall and confidently misroutes them. An earlier per-site figure for Pleura and Mediastinum turned out to reflect one external batch's signature, not the biology, and we corrected it. Treat any Pleura and Mediastinum output as unreliable.

Rare sites are data-starved (a data ceiling, not a model ceiling)

Pleura and Mediastinum, Thymus, Esophagus, Skin, and Eye have very few examples (on the order of seventeen each in the relevant held-out evaluation), so their per-site metrics are statistically fragile. These tissues are close to the entire public universe of their kind, so the data is exhausted at source, and stronger models do not move the number.

Cross-platform inputs are harder

A single tumor sequenced on a different pipeline is the hardest case. Rather than guess, the model abstains on most of those and commits only when it is confident. Inputs that skip the harmonization step will be misclassified.

The uncertainty guarantees are in-distribution only

Calibration and conformal coverage are measured on a held-out split from the same distribution. They do not hold under platform or batch shift. The novelty and input-validity checks exist precisely because of this, and are themselves reference-only, not a validated clinical detector.

No fairness audit, no clinical validation

All evaluation is retrospective on public cohorts. There is no prospective study, no independent clinical-site validation, and no subgroup-equity audit. The training cohorts are not characterized here for demographic balance, and performance on underrepresented groups is unmeasured and may be worse.

Scope

The 25 sites

Adrenal GlandBladder & UrinaryBlood & Bone MarrowBrain & CNSBreastCervixColorectalEsophagusEyeHead & NeckKidneyLiver & BiliaryLungLymph NodeOvaryPancreasPleura & MediastinumProstateSkinSoft Tissue & BoneStomachTestisThymusThyroidUterus
Access

Invite-only, under agreement

Provenance-2 is invite-only during the research phase and granted under a confidentiality agreement. See how the flagship core is validated, check the public evidence status, read Safety, then apply for access.