15% of all profit is donated to Heal Palestine and the Palestine Children's Relief Fund (PCRF)
Platform Research Validation Safety Pricing Company Demo Request access
Provenance-2 · flagship research system

A calibrated site-of-origin read from one RNA-seq profile.

The Provenance-2 core estimates anatomical site across 25 sites, reports uncertainty, and abstains when the evidence is insufficient. Separate research-preview signals can add context, but they do not have equivalent validation and do not replace orthogonal assays.

Capabilities and maturity

One measured core, with research layers kept separate

Site of origin is the flagship task with published cohort-level evidence. Confidence, abstention, and attribution support that task. Subtype, immune, and target signals remain exploratory and carry their own limits.

Site of origin flagship core

Estimates one of 25 anatomical sites (macro-F1 0.940 on the 2,700-sample internal held-out test, in-distribution). Under platform shift, the served decision policy abstains on roughly three in five profiles rather than force a single-site call.

Molecular subtype exploratory

Provides a within-site transcriptomic subtype signal only where a research head exists, including breast, colorectal, brain, and stomach. Evidence varies by site; it is not an assay result.

Target context hypothesis only

Links over-expressed genes to ChEMBL target and compound context. Over-expression is not dependency or predicted efficacy; these are research leads, never treatment recommendations. ChEMBL-derived context remains free research data, not a paid capability.

Immune context exploratory

Summarizes literature-signature expression patterns from immune-cold to inflamed. It is a descriptive transcriptomic estimate, not IHC, flow cytometry, or a validated clinical biomarker.

Confidence and abstention core guardrail

The site-of-origin call carries calibrated confidence and a conformal candidate set in-distribution. Low-confidence or unfamiliar inputs can be flagged, widened to a set, or abstained on.

Attribution interpretability aid

Gene attributions show which inputs most influenced a site call. They support review, but they are not causal drivers, biomarkers, or evidence about any single gene.

The readout

What a single call looks like

Not a score in isolation. A ranked, calibrated read you can interrogate, with the alternatives it considered and the confidence behind the call.

Site-of-origin readoutillustrative
Prostate 93% Bladder 4% Kidney 2% Adrenal gland 1%

Illustrative example, not a live model output. On the held-out test set the model reads the right site at macro-F1 0.940 (in-distribution), and it abstains rather than force a call when no site is clearly ahead.

Calibrated confidencereliability
predicted confidence → observed accuracy →

Schematic reliability diagram. Temperature scaling cut the model's calibration error roughly 8x (ECE 0.091 to 0.011), so a stated 80% means about 80% in practice, in-distribution.

Workflow

From raw counts to a bounded result

01

Bring your profile

Start from a single tumor expression profile, gene-level counts or TPM.

02

QC and harmonize

The pipeline maps to a common gene space, normalizes, and screens quality before any prediction.

03

Read the result

The flagship core returns a site estimate, alternatives, uncertainty, and attributions. Any exploratory research layer is labeled separately with its own caveat.

Understand the interface, then inspect the evidence

The guided demo is illustrative, not a live model run. The evidence ledger publishes cohort-level validation and states why a reproducible accession-linked sample bundle is not yet public.