Unsupervised Machine Learning Clustering of Prostate Cancer Cases in Soweto Using Clinical, Demographic, and Socioeconomic Variables (PSA, GS, TNM)
Loading...
Date
Authors
Journal Title
Journal ISSN
Volume Title
Publisher
University of the Witwatersrand, Johannesburg
Abstract
Introduction: Prostate cancer (PCa) is one of the leading cause of morbidity among men worldwide, and men of African descent experience significantly higher incidence, prevalence, and mortality. Traditional diagnostic methods, for PCa typically depend on specific biomarkers, such as, prostate specific antigen (PSA). These are gathered through procedures, such as the digital rectal exam (DRE). DRE is non-invasive, dependent on clinician’s expertise, thus highly subjective, often leading to vari- ability in detection rates. Furthermore, the excessive cost and time associated with the procedures often create barriers to timely diagnosis, particularly in low income countries (LIC). This highlights the need for innovative, cost-effective, and scalable approaches to PCa diagnosis. Traditional di- agnosis methods often overlook the wealth of information in high-dimensional data collected from cases over time, such as PSA, Gleason scores (GS), and socio-demographic data. The primary objec- tive of this study was to investigate factors associated with PCa morbidity among men of African descent using unsupervised machine learning (UML) techniques. UML offers a promising approach by identifying correlations, patterns, and risk profiles in complex datasets, which facilitates faster, precise and cost effective diagnostics. Methods: This study utilized UML techniques to address these problems, using a robust dataset from the Men of African Descent and Carcinoma of the Prostate (MADCaP) study, conducted in Soweto, South Africa; during the survey period 2016 to 2020. Stratification of cases into clinically meaningful groups was performed using a combination of key prognostic variables: GS, P SA level, and T N M stage. The T N M criteria included primary Tumour extent (T), regional lymph node involvement (N), and the presence of distant metastasis (M). This multi-factorial approach was de- signed to elucidate distinct risk profiles and generate novel insights into PCa heterogeneity. Key methodologies included data pre-processing, clustering using advanced UML algorithms, and eval- uation using the silhouette score to validate the robustness of the clusters. Applying clustering algorithms including hierarchical, K-means, and spectral clustering, this research sought to uncover hidden correlations and stratify patients into meaningful risk groups. This approach not only com- plements traditional diagnostic methods, but also provides a pathway for more efficient, accessible, and personalized PCa management, particularly in populations with limited access to advanced healthcare resources. ii Results: The dataset comprised clinical, behavioural, and demographic information for (n = 1, 096) PCa cases. Case ages ranged from 40 − 94 years, with a median age of 67 (IQR : 62 − 94). Based on GS, 36.6% of cases were classified as high-risk, 52.9% as intermediate-risk, and 10.8% as low-risk. Reported cigarette consumption varied widely, reaching up to 60 cigarettes per day, with a median of 8 (IQR : 5 − 8). Clustering analysis revealed clinically significant subgroups, underscoring the combined influence of GS, PSA levels, and TNM staging on PCa severity. These findings underscore the utility of UML in case stratification and highlight its potential to in- form risk-adaptive clinical guidelines. To validate the clustering models applied to the PCa dataset, the Silhouette score was employed. This metric evaluates both intra-cluster cohesion and inter- cluster separation, with values ranging from –1 to +1. Scores approaching +1 indicate well-defined and clearly separated clusters, values near 0 suggest overlapping clusters, and values approaching –1 reflect poor or incorrect clustering assignments. Conclusion:The results of this study demonstrate the value of applying UML to large-scale clinical datasets, providing a pathway to uncover underlying subgroups that may enhance clinical decision- making, support personalized PCa management strategies, and inform policymaking. Future re- search should aim to validate these PCa clusters using supervised machine learning approaches and incorporate additional dataset features to further explore their implications for treatment outcomes.
Description
A research report submitted in fulfillment of the requirements for the Master of Science , in the Faculty of Health Sciences, School of Public Health, University of the Witwatersrand, Johannesburg, 2025
Citation
Rwafa, Tanaka. (2025). Unsupervised Machine Learning Clustering of Prostate Cancer Cases in Soweto Using Clinical, Demographic, and Socioeconomic Variables (PSA, GS, TNM) [Master’s dissertation, University of the Witwatersrand, Johannesburg]. WIReDSpace. https://hdl.handle.net/10539/49776