Comparative Analysis Ensemble Learning Models in BRCA Dataset
Tarih
Yazarlar
Dergi Başlığı
Dergi ISSN
Cilt Başlığı
Yayıncı
Erişim Hakkı
Özet
Machine learning (ML) has become a transformative tool in medical applications, particularly in the diagnosis of diseases such as breast cancer. This study evaluates the performance of various ML models, including single classifiers (Logistic Regression-LR, K-Nearest Neighbors-KNN, Random Forest-RF, and Support Vector Machine-SVM) and ensemble learning (EL) models (Extreme Gradient Boosting-XGBoost, Soft Voting, Hard Voting, Bagging, and Stacking), for breast cancer diagnosis using the Breast Cancer Wisconsin (Diagnostic) dataset. LR emerged as the most effective single classifier, achieving the highest accuracy (98.25%). Among EL models, Soft Voting showed superior performance with an accuracy of 97.37% and competitive Area Under the Curve (AUC) values. Pairwise statistical comparison using McNemar's test revealed significant differences between XGBoost and other EL models, while Soft Voting, Hard Voting, Bagging, and Stacking did not show statistically significant differences. These results underscore the importance of EL techniques in achieving robust and reliable predictions. The results highlight the potential of ML to improve diagnostic accuracy and support clinical decision making in breast cancer detection, paving the way for further advances in AI-driven healthcare solutions. © 2025 Elsevier B.V., All rights reserved.










