15% of all profit is donated to Heal Palestine and the Palestine Children's Relief Fund (PCRF)
Platform Research Validation Safety Pricing Company Demo Request access
Safety & responsible use

Built to say “I don't know”.

In oncology, a confident wrong answer is worse than no answer. Provotics is designed around calibration, abstention, and transparency so its outputs stay honest and reviewable.

Not a medical device

Provotics is a research and educational project. It is not a medical device, has not been cleared or approved by any regulator, and must not be used for clinical diagnosis, treatment decisions, or patient care. Its outputs are research hypotheses.

Guardrails

How we keep it honest

Calibrated confidence

Predictions carry probabilities that mean what they say, so users can weigh a call instead of trusting it blindly. This calibration is measured in-distribution only, on GDC FPKM-UQ data of the kind the model was trained on, and degrades on inputs from a different sequencing platform or pipeline.

Abstention

Low-confidence and out-of-distribution profiles are flagged or abstained on, not forced into a label. The conformal coverage guarantee behind this holds in-distribution (GDC FPKM-UQ) only and degrades off-platform, where the model leans on abstention rather than a coverage promise.

Explainability

Each call exposes the driver genes behind it, so an expert can check it against known biology.

Data privacy

No identifiable patient data should ever be submitted. Inputs are expression profiles, not patient records. See our Privacy Policy.

Limitations

Where the model is weakest

We publish our failure modes on purpose. A model that hides them is more dangerous than one that names them. These come straight from our model card.

  • External performance is below the internal number. Macro-F1 is 0.940 on the held-out test set (n=2,700, GDC, 25 sites), but on independent external tumors the honest figure is lower: a 7-site top-1 of 0.906 on 381 cBioPortal samples. We report the external number separately rather than blend it in to look better.
  • Cross-platform inputs are harder. Calibration and abstention guarantees hold in-distribution and degrade on data from a different sequencing platform or pipeline. The model abstains on inputs it does not recognize rather than guess.
  • Some sites are genuinely weak. Rare sites with little training data (Thymus, Pleura and Mediastinum, Esophagus, Skin) are less reliable, and a few sites are overconfident on real patients (for example Soft Tissue and Esophagus, where stated confidence runs well above measured accuracy).
  • Mesothelioma is a known miss, and it fails silently. An earlier pleural-cancer result turned out to be a batch artifact: on real TCGA mesothelioma the model currently has 0% recall. Worse, it does not abstain on these cases. It confidently mislabels real mesothelioma as Thymus (at roughly 0.76 confidence), so the error looks like a normal, sure-footed call rather than an uncertain one. We do not claim mesothelioma detection, and a confident Thymus call should not be trusted to rule mesothelioma out.
  • Training data skews toward certain populations. The training cohorts (GDC, predominantly TCGA-derived) are not characterized for race or ethnicity, ancestry, age, sex, or geographic balance, and are known in the literature to skew toward certain populations. Performance across underrepresented groups is unmeasured and may be materially worse.
  • No subgroup fairness audit yet. We have not completed a fairness audit across sex or genetic ancestry, so we make no subgroup performance claims.
  • Accuracy depends heavily on the source cohort. On raw external cohorts from different sequencing centres (before our harmonization step), measured per-source accuracy ranged from about 6% to 100% across eight independent cBioPortal studies. This is why the model harmonizes inputs and is designed to abstain on profiles it does not recognize rather than guess, though that safeguard is not absolute and a confident call on an unfamiliar cohort is not a guarantee of accuracy. Do not assume uniform accuracy across cohorts or platforms.
  • Everything downstream is a hypothesis. Subtype, druggable-target, and drug-response outputs are research hypotheses for expert review, not conclusions, and never treatment guidance.
Data and attribution

What the model is built on

Provotics is built on open and public research data. Drug and mechanism annotations in the treatment-lead outputs are derived from ChEMBL, made available by EMBL-EBI (the European Bioinformatics Institute) under the Creative Commons Attribution-ShareAlike 3.0 licence (CC BY-SA 3.0). We attribute ChEMBL as the source; any redistribution of ChEMBL-derived data carries the same share-alike obligation. Training and benchmarking also use open genomics data including GDC / TCGA, GTEx, CPTAC, and DepMap, each under its own terms. Full data-source detail is in our Access Agreement.

Questions about responsible use?

We're glad to talk through appropriate use, validation, and limitations before you apply.