Short pieces on methodology, validation, and the choices behind reading cancer from gene expression.
A tumor profile carries roughly 18,000 genes. We read a curated panel of 3,882, because more columns are not more signal, and a focused panel travels better across labs and abstains when a site is too rare to call.
A calibrated confidence means 90% happens to be right about 90% of the time. Calibration error fell from 0.091 to 0.011, and at 90% conformal coverage the model abstains when it is unsure. That is worth more than a bigger accuracy number.
A held-out macro-F1 of 0.940 and a raw single-tumor number that drops to about 0.69 on cross-platform data are both true. Hiding the second one is how models earn false trust.
Where a cancer started is written in which genes it expresses. The hard part is reading that signal out of 18,000 noisy, correlated measurements, and knowing when not to.