Provotics reads cancer's molecular signature from gene expression. The work is about doing that accurately, calibrating the uncertainty honestly, and keeping every call traceable to biology.
Each tumor carries roughly 18,000 gene measurements. From that full profile, Provotics reads a curated 3,882-gene panel and learns the patterns that distinguish tissues of origin and molecular states.
Each profile is scored and turned into one calibrated result, built to stay robust on noisy, correlated, high-dimensional data.
Probabilities are calibrated so a stated 80% means roughly 80% in practice, the difference between a number you can act on and one you can't.
Out-of-distribution profiles are flagged and abstained on rather than forced into a label, so a wrong call is caught instead of reported.
Numbers from a held-out test set, not the training set.
Fully independent, out-of-distribution cohorts show the generalization gap that honest external validation always reveals. Some rare sites remain weak, and the model is deliberately built to abstain on samples it doesn't recognize rather than guess.
We write all of this down. The model card documents what Provenance-2 is, what it is for, and where it fails, and the validation page shows how every number was measured, including the sites it gets wrong.
Provotics is a research and educational project. It is not a medical device, and its outputs are research hypotheses, not clinical decisions. Read more on the Safety page.
Open to research-lab and biotech collaborations. Request access to evaluate the model under our access agreement.