Evaluation of Machine Learning Models for Heart Disease Prediction under Progressive Feature Reduction: An Exploratory Study
Abstract
Cardiovascular disease (CVD) remains the main cause of death worldwide, which makes timely and accurate risk identification very important for reducing related risks. In constrained screening situations, there is need for predictive models that can work with fewer inputs as a decision support system. Reduced features count, however, does not necessarily means lower input features acquisition burden. This study presents a comparative evaluation of eleven commonly used machine learning classifiers on progressively reduced subsets of the commonly used UCI Cleveland Heart Disease dataset. From the full 13-feature dataset, we extracted three subsets containing 10, 8, and 6 features. This was done using Chi-Square-based feature selection technique. The aim was to examine how performance of classifier changes under progressive feature reduction and to identify models that remain robust for situations when fewer inputs are available. This paper should be read as an exploratory baseline study. It does not propose a new prediction model. Instead, it compares commonly used models under the same feature-reduction setting to understand their behaviour. Results of the study suggest that Logistic Regression demonstrated consistent performance across all feature sets, including the most reduced 6-feature dataset with little loss in classification metrics. Random Forest and Naïve Bayes also gave competitive performance. However, the reduced subsets were selected statistically and retained some clinically burdensome variables. Therefore, fewer features should not be treated as direct evidence of lower cost or reduced screening time. Cost/time implications are discussed qualitatively and were not measured. The study contributes comparative baseline evidence on classifier robustness under feature reduction and highlights the need for feasibility-aware feature selection in future work












