Evaluasi SMOTE dan SMOTE+Tomek untuk Mengatasi Ketidakseimbangan Kelas pada Prediksi Stunting Balita Berbasis Pembelajaran Mesin
DOI:
https://doi.org/10.53564/sajy3k06Keywords:
Stunting, Class Imbalance, K-Nearest Neighbours, Random Forest, SMOTEAbstract
Class imbalance is a recurring obstacle in machine learning based screening of child nutritional status. This study evaluates and compares the effect of SMOTE and SMOTE+Tomek Links on the classification of toddler nutritional status using K-Nearest Neighbours (KNN) and Random Forest (RF). The data consist of 9,426 anthropometric records with four predictors, labelled into three classes: normal (7,212; 76.51%), moderate stunting (1,612; 17.10%) and severe stunting (602; 6.39%), a majority to minority ratio of about 12:1. Min-Max scaling and resampling were fitted on training data only, inside a pipeline, and the models were assessed on an independent 20% test set (n = 1,886) with stratified 5-fold cross validation. At baseline, RF reached 0.975 accuracy and 0.932 macro-F1, while KNN displayed an accuracy paradox: 0.913 accuracy but only 0.392 recall on severe stunting (47 of 120 cases). SMOTE raised KNN recall to 0.733 and macro-F1 from 0.753 to 0.816. For RF, SMOTE improved both criteria at once: recall rose from 0.825 to 0.908 (99 to 109 cases) and macro-F1 from 0.932 to 0.949. SMOTE+Tomek performed almost identically, removing only 30 of 17,307 training samples. RF with SMOTE is therefore recommended
Downloads
References
[1] WHO, “Malnutrition,” World Health Organization. [Online]. Available: https://www.who.int/news-room/fact-sheets/detail/malnutrition
[2] M. Y. E. Soekatri, S. Sandjaja, and A. Syauqy, “Stunting Was Associated with Reported Morbidity, Parental Education and Socioeconomic Status in 0.5–12-Year-Old Indonesian Children,” IJERPH, vol. 17, no. 17, p. 6204, Aug. 2020, doi: 10.3390/ijerph17176204.
[3] S. Zaleha and H. Idris, “IMPLEMENTATION OF STUNTING PROGRAM IN INDONESIA: A NARRATIVE REVIEW,” JAKI, vol. 10, no. 1, pp. 143–151, Jun. 2022, doi: 10.20473/jaki.v10i1.2022.143-151.
[4] J. H. Rah, S. Sukotjo, N. Badgaiyan, A. A. Cronin, and H. Torlesse, “Improved sanitation is associated with reduced child stunting amongst Indonesian children under 3 years of age,” Maternal & Child Nutrition, vol. 16, no. S2, p. e12741, Oct. 2020, doi: 10.1111/mcn.12741.
[5] S. Ndagijimana, I. H. Kabano, E. Masabo, and J. M. Ntaganda, “Prediction of Stunting Among Under-5 Children in Rwanda Using Machine Learning Techniques,” J Prev Med Public Health, vol. 56, no. 1, pp. 41–49, Jan. 2023, doi: 10.3961/jpmph.22.388.
[6] E. K. Anku and H. O. Duah, “Predicting and identifying factors associated with undernutrition among children under five years in Ghana using machine learning algorithms,” PLoS ONE, vol. 19, no. 2, p. e0296625, Feb. 2024, doi: 10.1371/journal.pone.0296625.
[7] S. M. J. Rahman et al., “Investigate the risk factors of stunting, wasting, and underweight among under-five Bangladeshi children and its prediction based on machine learning approach,” PLoS ONE, vol. 16, no. 6, p. e0253172, Jun. 2021, doi: 10.1371/journal.pone.0253172.
[8] KEMENTERIAN KESEHATAN RI, Buku Saku SSGI 2022. KEMENTERIAN KESEHATAN RI, 2022.
[9] A. Aziz, F. Insani, J. Jasril, and F. Syafria, “Implementasi Metode Learning Vector Quantization (LVQ) Untuk Klasifikasi Keluarga Beresiko Stunting,” bits, vol. 5, no. 1, Jun. 2023, doi: 10.47065/bits.v5i1.3478.
[10] S. Y. Andriyani, M. S. Lydia, and S. Efendi, “Optimization of Support Vector Machine Algorithm Using Stunting Data Classification,” J. Prisma. Sains, vol. 11, no. 1, p. 164, Jan. 2023, doi: 10.33394/j-ps.v11i1.6619.
[11] A. Iriany, W. Ngabu, D. Arianto, and A. Putra, “CLASSIFICATION OF STUNTING USING GEOGRAPHICALLY WEIGHTED REGRESSION-KRIGING CASE STUDY: STUNTING IN EAST JAVA,” BAREKENG: J. Math. & App., vol. 17, no. 1, pp. 0495–0504, Apr. 2023, doi: 10.30598/barekengvol17iss1pp0495-0504.
[12] Y. S. Asgedom et al., “Levels of stunting associated factors among under-five children in Ethiopia: A multi-level ordinal logistic regression analysis,” PLoS ONE, vol. 19, no. 1, p. e0296451, Jan. 2024, doi: 10.1371/journal.pone.0296451.
[13] R. Wajgi and D. Wajgi, “Malnutrition detection in infants using machine learning approach,” presented at the PROCEEDINGS OF THE INTERNATIONAL CONFERENCE ON COMPUTATIONAL INTELLIGENCE AND COMPUTING APPLICATIONS-21 (ICCICA-21), Nagpur, India, 2022, p. 040006. doi: 10.1063/5.0076876.
[14] V. Patel and H. Bhavsar, “Review on Data Balancing Approaches for Skewed Data Classification,” in 2024 7th International Conference on Contemporary Computing and Informatics (IC3I), Greater Noida, India: IEEE, Sep. 2024, pp. 235–241. doi: 10.1109/IC3I61595.2024.10828620.
[15] K. U. Apu and M. Ali, “A Systematic Literature Review on AI Approaches To Address Data Imbalance In Machine Learning,” FAET, vol. 2, no. 01, pp. 58–77, Jan. 2025, doi: 10.70937/faet.v2i01.57.
[16] J. Wang, H. Gu, and H. Gu, “Optimizing Stroke Prediction in Machine Learning by Addressing Data Imbalance,” in 2023 3rd International Signal Processing, Communications and Engineering Management Conference (ISPCEM), Montreal, QC, Canada: IEEE, Nov. 2023, pp. 665–669. doi: 10.1109/ISPCEM60569.2023.00125.
[17] S. He, “Addressing data imbalance in neural network spam detection with insights from SMS spam collection,” TNS, vol. 39, no. 1, pp. 212–218, Jul. 2024, doi: 10.54254/2753-8818/39/20240636.
[18] N. Faulina, K. Nisa, and W. Warsono, “Enhancing Tuberculosis Diagnosis: Effective Naive Bayes Classification using SMOTE and Tomek Links for Imbalanced Data,” InPrime:Ind.Jour.Pure.Applied.Math, vol. 6, no. 2, pp. 98–111, Oct. 2024, doi: 10.15408/inprime.v6i2.41463.
[19] M. M. Hasan Bhuiyan, S. A. Poly, and A. Kumar Acharyan, “Comparative Study on Brainstroke Prediction Using Data Balancing Techniques,” in 2024 International Conference on Advancements in Smart, Secure and Intelligent Computing (ASSIC), Bhubaneswar, India: IEEE, Jan. 2024, pp. 1–5. doi: 10.1109/ASSIC60049.2024.10507962.
[20] W. Rahayu et al., “Synthetic Minority Oversampling Technique (SMOTE) for Boosting the Accuracy of C4.5 Algorithm Model,” j. of artif. intell. and eng. appl., vol. 3, no. 3, pp. 624–630, Jun. 2024, doi: 10.59934/jaiea.v3i3.469.
[21] F. Y. A’la, N. Firdaus, Hartatik, and M. A. Safi’Ie, “SMOTE on Numeric Breast Cancer Dataset to Overcome Imbalance Class,” in 2023 6th International Conference of Computer and Informatics Engineering (IC2IE), Lombok, Indonesia: IEEE, Sep. 2023, pp. 335–339. doi: 10.1109/IC2IE60547.2023.10331221.
[22] C. El Morr, M. Jammal, H. Ali-Hassan, and W. El-Hallak, “Data Preprocessing,” in Machine Learning for Practical Decision Making, vol. 334, in International Series in Operations Research & Management Science, vol. 334. , Cham: Springer International Publishing, 2022, pp. 117–163. doi: 10.1007/978-3-031-16990-8_4.
[23] G. Y. Lee, L. Alzamil, B. Doskenov, and A. Termehchy, “A Survey on Data Cleaning Methods for Improved Machine Learning Model Performance,” Sep. 15, 2021, arXiv: arXiv:2109.07127. doi: 10.48550/arXiv.2109.07127.
[24] M. D. V. Prasad and S. T, “Enhancing K-Means Clustering Accuracy Through Modified Robust Scaling Technique,” Nov. 18, 2024, Computer Science and Mathematics. doi: 10.20944/preprints202411.1245.v1.
[25] H. A. Ahmed, P. J. Muhammad Ali, A. K. Faeq, and S. M. Abdullah, “An Investigation on Disparity Responds of Machine Learning Algorithms to Data Normalization Method,” ARO, vol. 10, no. 2, pp. 29–37, Sep. 2022, doi: 10.14500/aro.10970.
[26] A. Pajankar and A. Joshi, “Preparing Data for Machine Learning,” in Hands-on Machine Learning with Python, Berkeley, CA: Apress, 2022, pp. 79–97. doi: 10.1007/978-1-4842-7921-2_6.
[27] V. V. Starovoitov and Yu. I. Golub, “Data normalization in machine learning,” Informatika (Minsk), vol. 18, no. 3, pp. 83–96, Sep. 2021, doi: 10.37661/1816-0301-2021-18-3-83-96.
[28] A. Asesh, “Normalization and Bias in Time Series Data,” in Digital Interaction and Machine Intelligence, vol. 440, C. Biele, J. Kacprzyk, W. Kopeć, J. W. Owsiński, A. Romanowski, and M. Sikorski, Eds., in Lecture Notes in Networks and Systems, vol. 440. , Cham: Springer International Publishing, 2022, pp. 88–97. doi: 10.1007/978-3-031-11432-8_8.
[29] H. Bichri, A. Chergui, and M. Hain, “Investigating the Impact of Train / Test Split Ratio on the Performance of Pre-Trained Models with Custom Datasets,” IJACSA, vol. 15, no. 2, 2024, doi: 10.14569/IJACSA.2024.0150235.
[30] B. Supri, Rudianto, Abdurohim, Badriatul Mawadah, and Helmi Ali, “Asian Stock Index Price Prediction Analysis Using Comparison of Split Data Training and Data Testing,” jemsi, vol. 9, no. 4, pp. 1403–1408, Aug. 2023, doi: 10.35870/jemsi.v9i4.1339.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 JURNAL AKADEMIKA

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.








