Analisis Pengaruh Seleksi Atribut Relief-F dan Gain Ratio Terhadap Performa Naïve Bayes Classifier
Keywords:
Attribute SelectionAbstract
Naïve Bayes Classifier (NBC) is one of the most popular probabilistic classification algorithms in
data mining, known for its simplicity and efficiency. However, NBC performance tends to
degrade when datasets contain irrelevant or noisy attributes. This study analyzes the effect of
attribute selection using Relief-F and Gain Ratio methods on the performance improvement of
NBC. Two benchmark datasets from the UCI Machine Learning Repository were selected to
represent contrasting data characteristics: the House Vote dataset (435 records, symbolic
attributes with balanced class distribution) and the Bank Marketing dataset (45,211 records,
numeric and categorical attributes with severe class imbalance, approximately 88% majority
class). All experiments were implemented in Google Colaboratory using Python. Three
experimental scenarios were applied to each dataset: (1) NBC without attribute selection as
baseline, (2) NBC with Relief-F attribute selection, and (3) NBC with Gain Ratio attribute
selection. Performance evaluation used 10-fold cross-validation with metrics including
accuracy, precision, recall, F1-score, and confusion matrix. Results show that on the House Vote
dataset, Relief-F increased NBC accuracy from 90.11% to 93.79% (+3.68%), while Gain Ratio
reduced accuracy to 89.43%. On the Bank Marketing dataset, Relief-F improved accuracy to
89.36% and improved minority class recall from 29.34% to 35.71%, while Gain Ratio yielded
only marginal improvement. Overall, Relief-F proved more effective than Gain Ratio in enhancing
NBC performance, particularly on datasets with clear classification patterns and imbalanced
class distribution.
References
Ramadandi, S. J. (2020). Klasifikasi gaya belajar mahasiswa menggunakan metode Naïve Bayes
Classifier. Jurnal Teknologi Dan Informasi (JATI), 10(September), 170–179.
https://doi.org/10.34010/jati.v10i2
Riany, A. F., & Testiana, G. (2023). Penerapan data mining untuk klasifikasi penyakit. Prosiding
Seminar Nasional, 297–305.
Saelan, M. R. R., Sahputra, D. A., Widiastuti, W., & Gata, W. (2020). Komparasi algoritma
klasifikasi untuk prediksi minat sekolah tinggi pelajar pada Students Alcohol
Consumption.
Jurnal
Sains
Dan
Informatika,
6(2),
120–129.
https://doi.org/10.34128/jsi.v6i2.236
Sari, S. I. P., Pranoto, W. J., & Verdikha, N. A. (2023). Analisis pengaruh Gain Ratio untuk
algoritma K-Nearest Neighbor pada klasifikasi data banjir di Kota Samarinda. Jurnal
Sains
Komputer
Dan
Teknologi
Informasi,
6(1),
54–59.
https://doi.org/10.33084/jsakti.v6i1.5472
Setiyorini, T., & Asmono, R. T. (2020). Implementation of Gain Ratio and K-Nearest Neighbor for
classification of student performance. Jurnal Pilar Nusa Mandiri, 16(1), 19–24.
https://doi.org/10.33480/pilar.v16i1.813
Xi, J., Jiang, Q., Liu, H., & Gao, X. (2023). Lithological mapping research based on feature
selection model of Relief-F-RF. Applied Sciences.
Yusra, R. N., Sitompul, O. S., & Sawaluddin. (2021). Kombinasi K-Nearest Neighbor (KNN) dan
Relief-F untuk meningkatkan akurasi pada klasifikasi data. InfoTekJar: Jurnal Nasional
Informatika Dan Teknologi Jaringan, 1, 0–5.
Zhao, M., & Ye, N. (2024). High-dimensional ensemble learning classification: An ensemble
learning classification algorithm based on high-dimensional feature space
reconstruction.
Applied
https://doi.org/10.3390/app14051956
Sc



