Getting a Cancer AI Pipeline From Notebook to Something Useful
Most people think building a diagnostic or therapeutic AI for oncology means training a model on some public dataset and calling it a day. It does not work like that. The gap between a research script and a system that a pathologist would actually look at is enormous. I spent roughly two years building a model that could segment tumor regions from whole-slide H&E images and predict microsatellite instability status. Here is what that process actually looks like, including the parts that do not make it into publications.Artificial Intelligence In Cancer Research Diagnosis And Therapy
The phrase gets thrown around a lot. At its core, this work falls into three buckets that overlap but are fundamentally different problems. Diagnostic AI covers classification and segmentation tasks — detecting tumors, grading tissue, predicting biomarkers from images or omics data. Therapeutic AI is closer to predictive pharmacology and treatment response modeling. Research AI is the infrastructure layer: data curation, augmentation strategies, and validation pipelines that most people skip until their model collapses in production. If you are just starting, pick one bucket and do not try to build all three at once. A system that tries to do diagnosis and drug response prediction simultaneously without proper data will fail on both tasks. The models will overfit to whatever signal is strongest in your training set, which is usually batch effects or scanner artifacts, not biology. I work primarily with PyTorch and MONAI for the imaging side. MONAI is not glamorous but it handles DICOM, NIfTI, and multi-GPU distributed training without the headaches you get trying to bolt those capabilities onto a generic framework. For tabular omics data, I usually strip the pipeline down to something simpler — a lightly modified PyTorch implementation with early stopping and a custom loss function that accounts for class imbalance. You do not need a heavy library for that.
Data Acquisition and Cleaning
This is where the real work lives. I cannot stress this enough. Most teams spend maybe ten percent of their time on data and the rest on model architecture. That is backwards. In my experience, data quality determines whether you get a model that generalizes or a model that memorizes the hospital it was trained at. I pull data from publicly available sources when possible — The Cancer Genome Atlas, TCIA, the PAIP challenges for pancreatic histopathology, CPTAC for proteomics. But public data is never enough on its own. The distributions are too narrow. I supplement with institutional data, which requires IRB approval and a data use agreement. That process alone takes three to six months depending on your institution's review board. Do not underestimate that timeline. Once you have the data, cleaning is relentless. You will encounter missing labels, inconsistent slide formats, batch effects from different scanners, and annotations that were done by three different pathologists using three different criteria. I wrote a preprocessing pipeline that first checks for DICOM tag consistency, normalizes stain colors using Macenko normalization for H&E slides, and then runs a quick quality filter that rejects tiles with excessive blur or out-of-bounds annotations. The pipeline catches roughly fifteen percent of tiles that would otherwise poison the training set.
Here is a specific example of a problem I ran into that almost broke the entire project. I was training a model to predict mismatch repair deficiency from colon cancer histology slides. The model achieved ninety-four percent AUC on the internal test set. I was excited. Then I ran it on a small held-out set from a different hospital and the AUC dropped to sixty-one percent. The model had learned to detect scanner-specific artifacts, not biological signals. I had to go back and add a domain adversarial training component and retrain with stain augmentation that covered the full range of staining variability across institutions. The AUC on the external set went from sixty-one to eighty-three. Still not great, but usable.
Get the Full Details

Model Architecture Choices
For image-based tasks, convolutional architectures are still the default, but vision transformers are gaining ground. I use a hybrid approach — a ResNet or ConvNeXt backbone with attention pooling for whole-slide images. The reason is practical. WSI tiles are gigapixel images. You cannot feed the whole slide into a transformer. Patch-based processing with a memory-efficient attention mechanism is the only approach that does not require a cluster of A100s and three days of training per epoch. For tabular data from genomics or proteomics, I use a straightforward multilayer perceptron with dropout and batch normalization. The counter-intuitive part is that more complex architectures do not help here. I tested a transformer on a gene expression prediction task and it performed worse than the MLP. The dataset was too small and the signal too sparse. Transformers need thousands of samples to justify their parameter count. With a few hundred patient samples, a simple MLP with careful regularization outperforms everything else. Another thing people get wrong is the loss function. Standard cross-entropy assumes balanced classes. Cancer datasets are rarely balanced. I use a combination of focal loss and weighted cross-entropy, with the weights computed from the inverse frequency of each class in the training set. This shifts the model's focus toward the minority class without completely ignoring the majority. The trade-off is slightly lower overall accuracy but dramatically better sensitivity on the classes that matter clinically.
Validation and Clinical Deployment
Academic validation and clinical validation are two different things. A model can achieve state-of-the-art results on a held-out test set and still be useless in a real hospital. I learned this the hard way. My metastatic prostate cancer detection model worked perfectly on cross-validated data but failed when deployed on slides from a different staining lab. The staining variability was outside the range of what the model had seen during training. The workaround was to build a simple domain adaptation layer that adjusted the feature distribution to match a reference staining protocol. I used a small set of anchor slides from the target lab and trained a lightweight adapter module that mapped the model's internal representations into the new domain. This improved performance by about twelve percentage points on the target domain. It is not a perfect solution. You still need prospective validation before any model touches patient care, but it gets you past the most common failure mode in early deployment. For therapeutic AI, the challenge is different. Predicting drug response from cell line data like GDSC or CTPAC does not translate well to patients. The in vitro environment is nothing like a human tumor. I use a transfer learning approach where I fine-tune on patient-derived data when available, or I rely on multi-omics integration to capture patient-specific signals. Even then, the predictions are probabilistic at best. The model might tell you that a patient has a sixty percent probability of responding to a certain kinase inhibitor, but that does not replace a clinical trial or a liquid biopsy for actionable mutations.
Practical Implementation Steps
Start with a clear question. "Can AI help with cancer?" is not a question you can answer with a model. "Can a model predict PD-L1 expression from H&E whole-slide images in non-small cell lung cancer?" is a question you can test. Define your input modality, your output label, and your evaluation metric before you write a single line of code. Set up your environment with MONAI, PyTorch, and a GPU with at least twenty-four gigabytes of VRAM if you are working with whole-slide images. Smaller GPUs work for tabular data and lower-resolution images. Install cuDNN and make sure your CUDA version matches your PyTorch build. Getting this wrong wastes half a day. Build a minimal reproducible pipeline before attempting anything complex. Load a single slide, preprocess it, run inference, and visualize the output. If you cannot do that with ten lines of code, your architecture is too complicated for the data you have. Simplify until it works, then add complexity incrementally.
Track everything. I use Weights and Biases for experiment tracking because it handles large-scale hyperparameter sweeps without requiring you to write custom logging code. Every run should include the dataset version, preprocessing parameters, model architecture, training hyperparameters, and evaluation metrics. When your model breaks three months later and you need to reproduce it, you will be glad you did this.
Common Pitfalls and What I Do Instead
Overfitting to batch effects is the number one problem. I address it with domain randomization during training — randomly adjusting brightness, contrast, and hue to simulate different staining protocols. I also use test-time augmentation where I run the model on multiple augmented versions of the same slide and average the predictions. This adds about thirty seconds per slide but significantly improves robustness. Data leakage is another silent killer. I have seen it happen when patient samples appear in both training and test sets because the split was done at the slide level instead of the patient level. A single patient can contribute multiple slides. If those slides end up in different sets, the model learns patient-specific features rather than disease features. Always split at the patient level and verify that no patient ID appears in more than one set. The interpretability question is unavoidable in clinical AI. A model that predicts cancer but cannot explain why will not be used. I use Grad-CAM for image-based models and SHAP values for tabular data. These are not perfect — they show correlation, not causation — but they give clinicians enough signal to decide whether to trust the output. A heatmap that highlights the tumor region and aligns with the pathologist's own assessment is more persuasive than a black-box score, even if the score is slightly more accurate.
There are scenarios where AI simply does not work and nobody wants to admit it. Rare cancers with fewer than one hundred training samples will produce unreliable models regardless of architecture. Multi-class classification with overlapping histological features often hits a performance ceiling around seventy percent accuracy no matter how much data you add. In those cases, the honest recommendation is to collect more data or narrow the problem scope. Building a model that performs barely above chance and then presenting it as a breakthrough is bad for the field and worse for patients. If you want to start experimenting, the MONAI repository on GitHub has a collections of tutorials and pre-trained models that cover most common oncology tasks. The nifti datasets provided by TCIA can be downloaded directly through MONAI's built-in data loaders. Start with a classification task on a public dataset like TCGA-LUAD or TCGA-COAD before attempting anything involving multiple modalities or proprietary data. The field moves fast. Models that were state of the art eighteen months ago are already obsolete. What has not changed is the importance of clean data, rigorous validation, and understanding what your model actually learned versus what it happened to correlate with. Those basics determine whether your work is useful or just another paper with an impressive AUC that dies in translation.
