FIG. 02.6 — Project notes
- Independent
- Finished
Breast Cancer Tumor Classification with Explainable AI
An end-to-end, reproducible comparison of logistic regression and random forest, with leakage-safe model selection and an untouched holdout.
Educational only — not a clinical tool
Built for learning on public data. It is not a medical device and must not be used for diagnosis.
- Holdout accuracy
- 96.49%
- Holdout ROC-AUC
- 0.9960
- Malignant recall
- 92.86%
- CV fits
- 105

01Focus
An end-to-end, reproducible comparison of logistic regression and random forest, with leakage-safe model selection and an untouched holdout.
02Data
UCI Wisconsin Diagnostic Breast Cancer dataset, loaded via scikit-learn: 569 samples, 30 features describing cell nuclei from digitized fine-needle aspirate images (212 malignant / 357 benign).
03Methods
- Stratified 80/20 split (random_state 42); exploratory analysis on the training split only.
- Pipelines with median imputation and scaling (logistic regression) or random forest, with preprocessing learned within each fold.
- The same 5-fold stratified cross-validation for both models: 21 configurations, 105 CV fits.
- A prespecified selection rule — highest mean training CV ROC-AUC, ties favor logistic regression — recorded before the holdout was touched.
- Fixed 0.5 threshold, with no tuning on the holdout.
- SHAP values for every holdout sample, checked for additivity (max error 3.6e-15).
- One-command reproduction that regenerates the figures and README; continuous integration.
04Tools
- Python
- scikit-learn
- SHAP
- Streamlit
- pytest
- GitHub Actions
05Visualizations


06Findings
- Logistic regression was selected.
- Holdout (114 samples): 96.49% accuracy, 0.9960 ROC-AUC, 92.86% malignant recall.
07Limitations
- Small historical sample, not a screening cohort.
- No external or prospective validation.
- One fixed holdout, with sampling uncertainty.
- SHAP shows model associations, not biological causation.
- Educational only — not a clinical tool.