Bioinformatics
Ovarian cancer prediction (M.Sc. research)
Comparative study of baseline and ensemble machine-learning models for ovarian cancer malignancy prediction from routine blood biomarkers.
Problem
Ovarian cancer is often detected late because early malignancy is hard to distinguish from benign ovarian tumours using routine tests alone. The research question was whether standard blood work and tumour markers already collected in clinical practice can separate malignant from benign cases well enough to support screening.
The study used a public dataset from the Third Affiliated Hospital of Soochow University: 349 individuals collected between July 2011 and July 2018, split into 178 patients with benign ovarian tumours and 171 with ovarian cancer. Each record carried 49 predictor factors covering age and menopause status, 22 general chemistry tests, 19 routine blood tests, and 6 tumour markers, with the histological diagnosis as the target.
How it works
The work is organised as a set of Jupyter notebooks that run the full pipeline: data description, exploratory analysis, preliminary statistics, feature engineering, model training, and per-model analysis. To test how the train/validation/test split affects the comparison, the same pipeline was run under three ratios: 60-20-20, 70-15-15, and 80-10-10.
Feature selection used the mRMR (minimum redundancy, maximum relevance) algorithm to reduce the 49 predictors to the top 20, keeping features such as age, CA125, menopause status, albumin, and neutrophil and lymphocyte ratios. The comparison spanned baseline classifiers (Logistic Regression, K-Nearest Neighbours, Decision Trees) against ensemble methods (Voting, Stacking, and Boosting with XGBoost and Gradient Boosting). Models were evaluated with cross-validation, and runs were tracked with MLflow.
- Dataset: 349 patients, 49 pathology-derived features, binary malignant vs benign target.
- Feature selection: mRMR reduction to the top 20 features.
- Baseline models: Logistic Regression, K-Nearest Neighbours, Decision Trees.
- Ensemble models: Voting, Stacking, Boosting (XGBoost and Gradient Boosting), plus stacked ensembles.
- Validation: cross-validation across three split ratios, tracked with MLflow.
Hard parts
The dataset is small (349 rows) and close to balanced, so the comparison had to guard against results that would not hold on held-out data. Running the same pipeline under three split ratios and using cross-validation was the response, showing how stable each model’s ranking was across splits.
The 49 clinical features carry redundancy (correlated blood counts and chemistry panels), which can inflate apparent performance. mRMR feature selection addressed this by picking features individually relevant to the target while minimising redundancy among themselves, cutting the input to 20 features before modelling.
Results
The study delivered a like-for-like comparison of baseline and ensemble classifiers on the same 20 mRMR-selected features, evaluated with cross-validation under three split ratios. Model performance was compared on several standard classification metrics (not only accuracy), and the ensemble methods were analysed against the baseline classifiers to identify which family generalised better on this cohort.
A follow-on engineering track later took the trained models toward a deployable screening API, and in doing so flagged and corrected preprocessing issues from the original research pipeline (data leakage, naive imputation, and missing feature scaling), which is a concrete outcome of validating the research end to end.
Artifacts
- Three end-to-end notebooks (with cross-validation) for the 60-20-20, 70-15-15, and 80-10-10 splits, exported as both notebooks and PDF.
- A LaTeX export of the 60-20-20 notebook for write-up.
- A pinned modelling stack (scikit-learn, XGBoost, CatBoost, mRMR selection, MLflow).
- A downstream spec for an ovarian cancer risk screening API built on the research models.