A Comparative Study of Machine Learning and Deep Ensemble Architectures for Cardiovascular Disease Risk Prediction on Imbalanced Health-Indicator Data

Cardiovascular disease (CVD) remains the single largest contributor to global mortality, yetearly screening is often out of reach in resource-constrained settings where clinical testing iscostly or unavailable. This paper investigates whether inexpensive, self-reported healthindicators can support reliable automated CVD risk screening. Fifteen classifiers spanningclassical machine learning, deep neural architectures, and two purpose-built deep ensembles—EnsCVDD-Net and BlCVDD-Net—are evaluated on the publicly available Heart Disease HealthIndicators dataset. To counter the strong class skew inherent in population-level survey data,the pipeline couples point-biserial correlation-driven feature selection with Adaptive Synthetic(ADASYN) minority oversampling. Across accuracy, precision, recall, and F1-score, a soft-votingensemble of heterogeneous base learners delivers the strongest overall performance (91.7%accuracy, 0.918 F1), outperforming every individual model as well as the deep architectures by awide margin, while remaining light enough for real-time inference. The selected model ispackaged into a browser-based screening application that returns instant risk estimates fromuser-entered indicators. The results suggest that, for tabular health-indicator data of this kind,carefully balanced classical ensembles remain a stronger and more deployable choice thandeeper networks, and that accessible pre-clinical screening tools can be built entirely fromopen-source components.

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC