15% of profit is pledged to Heal Palestine and the Palestine Children's Relief Fund (PCRF)
Products Platform Research Validation Safety Pricing Company Demo Request access
Notes

Estimating tumor site of origin from bulk RNA-seq.

A research framing for why gene expression can carry anatomical origin signal, what breaks that signal, and why a useful system must know when not to answer.

Tumor site of origin is one of the most durable labels in cancer research. It organizes cohorts, frames comparative biology, and often anchors how a sample is interpreted in a study. Metastases and incomplete clinical context can hide that label. Bulk RNA-seq does not restore a clinical diagnosis. What it can do, under careful research conditions, is recover a statistical estimate of anatomical origin from the expression programs a tissue still carries.

This note is written as a methods-and-limits explainer for researchers evaluating transcriptomic site-of-origin tools. It is not a diagnostic guide. Provotics systems are research-use-only and are not medical devices. Nothing here recommends clinical use, patient management, or treatment decisions.

Why expression can encode origin

Cells keep lineage and tissue programs even after transformation. Prostate tumors tend to keep prostate-associated expression; lung tumors keep lung-associated programs; breast tumors keep breast programs. Those programs are not a single marker gene. They are joint patterns across thousands of correlated measurements, mixed with tumor subtype, purity, stroma, inflammation, and batch effects.

A bulk RNA-seq profile is therefore both rich and noisy. Famous markers help intuition (for example KLK3 in prostate or NKX2-1 in lung contexts), but robust research classifiers usually lean on a panel of genes chosen for signal stability rather than on one gene alone. Provenance-2, Provotics' flagship research system, reads a fixed 3,882-gene panel and returns a calibrated site-of-origin estimate across 25 anatomical sites. Details live on the platform and model card pages.

What the research task actually is

In research terms, site-of-origin estimation from expression is a multi-class prediction problem with severe class imbalance, biological adjacency between sites, and strong sensitivity to how the input was processed. The useful output is rarely a naked label. A useful research output usually includes:

  • a top estimate and alternative candidates,
  • a confidence number that means what it says under defined conditions,
  • an explicit option to abstain when the input looks unlike the training world,
  • and a clear statement of which cohorts and pipelines the numbers come from.

Without those pieces, a high accuracy on one held-out table is easy to misread. With them, a researcher can decide whether a call is worth following into literature, orthogonal assays, or further computational work.

Where the signal breaks

Site-of-origin from RNA-seq is hardest exactly where biology or measurement drifts away from the training distribution:

  • Cross-platform and off-pipeline inputs. Counts produced by a different quantification stack can look like a different statistical world even when the biology is the same tissue.
  • Rare or adjacent sites. Small-support tissues and biologically neighboring sites create residual confusions that more model capacity often fails to erase.
  • Low purity or heavily stromal samples. The measured profile can be dominated by non-tumor programs.
  • Incomplete panels or wrong identifiers. Missing genes, wrong gene namespaces, or unnormalized values invalidate the contract before the model runs.

This is why Provotics publishes both in-distribution held-out figures and weaker external or off-pipeline results on the validation page, and why the evidence ledger states what is downloadable versus still unpublished. If you only read the strongest table, you will over-trust the system.

Research framing versus diagnostic framing

Literature sometimes discusses cancers of unknown primary (CUP) and pathology workflows in the same breath as transcriptomic classifiers. Provotics does not target diagnostic CUP SEO or clinical pathology product claims. The honest research framing is narrower:

  • estimate anatomical origin for research samples with known or hypothesized labels,
  • study failure modes when labels are uncertain,
  • and keep every public claim inside research-use-only boundaries.

A research estimate can still be wrong. Independent verification remains the researcher's job. Provotics would rather return an abstention or a novelty flag than invent certainty.

How Provenance-2 approaches the problem

Provenance-2 is designed around those limits rather than around a single headline score:

  • one bulk expression profile in, a site-of-origin research estimate out,
  • temperature-scaled confidence so probabilities are usable for triage,
  • conformal candidate sets and abstention when nothing clears the bar,
  • novelty checks for inputs that sit outside what the model has learned,
  • and exploratory heads (subtype, target hypotheses, immune context) labeled separately so they are never confused with the paid site-of-origin core.

Exact served-core numbers, cohort sizes, and external gaps are maintained on Validation and Model card. Prefer those pages over repeating figures out of context. For a shorter companion note on why confidence must be honest, see What a calibrated probability actually buys you and the longer calibration and abstention note.

Practical checks before you trust a call

If you are evaluating any site-of-origin research tool, including Provenance-2, ask for:

  1. the exact input contract (gene panel, normalization, accepted file shapes),
  2. which cohorts the headline metrics come from,
  3. what happens on off-pipeline or external data,
  4. whether the system can abstain, and
  5. whether exploratory outputs are labeled as exploratory.

Provotics also ships a free file readiness checker so you can inspect format fit before an access application. That utility is intentionally narrow: readiness is not performance.

What this note is not claiming

It does not claim that RNA-seq replaces histology, imaging, or clinical workup. It does not claim that Provotics diagnoses cancer of unknown primary. It does not invent uplift, sensitivity, or clinical utility statistics. It argues only that bulk expression can carry origin signal for research, that the signal is fragile under distribution shift, and that calibrated confidence plus abstention are part of making that signal usable.

If you want the product boundary next, start at Platform. If you want the honesty stack, start at Validation and Evidence.

Provotics is a research and educational project. It is not a medical device and is not intended for clinical diagnosis or treatment decisions.
← All notes