Optimization of the Decision Tree Algorithm Using SMOTE for the Classification of Diabetes Mellitus Types Based on Clinical Medical Record Data

Authors

  • Ia Riham Nurrahma Universitas Cipasung Tasikmalaya
  • Nuk Ghurroh Setyoningrum Universitias Cipasung Tasikmalaya

DOI:

https://doi.org/10.69916/jkbti.v5i3.531

Keywords:

Decision Tree, SMOTE, Diabetes Mellitus Classification, Medical Record Data, Class Imbalance

Abstract

Diabetes mellitus is a chronic metabolic disorder characterized by elevated blood glucose levels. Accurate early classification of diabetes types is crucial for determining appropriate clinical interventions. However, clinical datasets often suffer from severe class imbalance, leading to model bias toward the majority class. This study proposes the optimization of the Decision Tree classification algorithm by integrating the Synthetic Minority Over-sampling Technique (SMOTE) using real-world clinical medical record data from the Singaparna Community Health Center. The dataset comprises 225 patient records with 38 health attributes. Preprocessing and attribute transformation techniques were applied, followed by stratified random sampling (80:20 split). To resolve the extreme class imbalance (1 Type 1 DM vs 224 Type 2 DM samples), SMOTE was applied exclusively to the training set, balancing the dataset to 358 instances (179 samples per class). The Decision Tree model was trained and evaluated using validation curves to prevent overfitting. The experimental results demonstrate that the SMOTE-optimized Decision Tree model with an optimal max depth of 3 achieved robust classification performance, yielding an accuracy of 86.67%, precision of 88.54%, recall of 84.12%, F1-score of 85.20%, and an ROC-AUC score of 0.9333. Furthermore, the decision tree visualization revealed transparent, age-based rule splits matching clinical domain knowledge. This study confirms that SMOTE effectively mitigates class imbalance in clinical decision-tree modeling, providing a reliable and interpretable decision-support tool for diabetes screening in primary healthcare settings.

Downloads

Download data is not yet available.

References

B. A. C. Permana and I. K. Dewi, “Komparasi Metode Klasifikasi Data Mining Decision Tree dan Naïve Bayes Untuk Prediksi Penyakit Diabetes,” Infotek J. Inform. dan Teknol., vol. 4, no. 1, pp. 63–69, 2021, doi: 10.29408/jit.v4i1.2994.

Kementerian Kesehatan RI, “Keputusan Menteri Kesehatan Republik Indonesia Nomor HK.01.07/MENKES/2009/2024 tentang Pedoman Nasional Pelayanan Klinis Tata Laksana Diabetes Melitus pada Anak,” 2024.

M. W. Nadeem, H. G. Goh, M. Hussain, S.-Y. Liew, I. Andonovic, and M. A. Khan, “Deep Learning for Diabetic Retinopathy Analysis: A Review, Research Challenges, and Future Directions,” Sensors, vol. 22, no. 18, Sep. 2022, doi: 10.3390/s22186780.

M. Vadduri and P. Kuppusamy, “Enhancing Ocular Healthcare: Deep Learning-Based Multi-Class Diabetic Eye Disease Segmentation and Classification,” IEEE Access, vol. 11, pp. 137881–137898, 2023, doi: 10.1109/ACCESS.2023.3339574.

R. N. Ramadhon, A. Ogi, A. P. Agung, R. Putra, S. S. Febrihartina, and U. Firdaus, “Implementasi Algoritma Decision Tree untuk Klasifikasi Pelanggan Aktif atau Tidak Aktif pada Data Bank,” Karimah Tauhid, vol. 3, no. 2, pp. 1860–1874, 2024.

Shelly, “Tinjauan Statistik dan Non Statistik pada Financial Distress,” 2021, Semarang, Indonesia.

B. G. Sudarsono, M. I. Leo, A. Santoso, and F. Hendrawan, “Analisis Data Mining Data Netflix Menggunakan Aplikasi Rapid Miner,” JBASE - J. Bus. Audit Inf. Syst., vol. 4, no. 1, pp. 13–21, 2021, doi: 10.30813/jbase.v4i1.2729.

A. R. Farrasy, A. N. Al-kharits, M. Bustomi, N. Maulana, and T. A. Rahman, “Prediksi Diabetes Berdasarkan Faktor Medis Pasien Menggunakan Algoritma Decision Tree,” J. Inf. Technol. Informatics Eng., vol. 1, no. 1, pp. 30–36, 2025.

A. P. Sabrina and Z. Fatah, “Implementasi Algoritma Decision Tree Untuk Klasifikasi Resiko Diabetes,” J.

Mhs. Tek. Inform., vol. 4, no. 2, pp. 273–279, 2025, doi: 10.35473/jamastika.v4i2.4543.

O. Siboro, Y. P. Banjarnahor, A. Gultom, N. A. Siagian, and P. D. P. Silitonga, “Penanganan Data Ketidakseimbangan dalam Pendekatan SMOTE Guna Meningkatkan Akurasi Algoritma K-NN,” in Sem. Nas. Inov. Sains Teknol. Inf. Komput. (SNISTIK), 2024, pp. 473–478.

M. Ghahremani, H. A. Otieno, M. Kazreen, and S. Shiaeles, “Machine Learning-Based Loan Approval Automation: Enhancing Efficiency, Accuracy and Fairness in Credit Decision-Making,” Expert Syst., vol. 43, no. 7, p. e70307, 2026, doi: 10.1111/exsy.70307.

N. S. Thomas and S. Kaliraj, “An Improved and Optimized Random Forest Based Approach to Predict the Software Faults,” SN Comput. Sci., vol. 5, no. 5, 2024, doi: 10.1007/s42979-024-02764-x.

X. Zhang, L. Wang, and Y. Liu, “An improved SMOTE algorithm for enhanced imbalanced data classification in medical informatics,” Sci. Rep., vol. 15, no. 1, pp. 1120–1132, 2025.

S. E. Saqila, I. P. Ferina, and A. Iskandar, “Analisis Perbandingan Kinerja Clustering Data Mining Untuk Normalisasi Dataset,” J. Sist. Komput. dan Inform., vol. 5, no. 2, pp. 356–365, 2023, doi: 10.30865/json.v5i2.6919.

D. Supriadi et al., “Prediction of Cooperative Loan Feasibility Using the K-Nearest Neighbor Algorithm,” J. Pilar Nusa Mandiri, vol. 17, no. 1, pp. 39–46, 2021, doi: 10.33480/pilar.v17i1.2183.

J. N. Oruh and B. C. Iwuji, “A Hybrid Prediction Model for Classifying Student’S Academic Performance Using Voting Ensemble Method,” Fudma J. Sci., vol. 10, no. 3, pp. 364–376, 2026, doi: 10.33003/fjs-20261003-4751.

M. P. Pulungan, A. Purnomo, and A. Kurniasih, “Penerapan SMOTE untuk Mengatasi Imbalance Class dalam Klasifikasi Kepribadian MBTI Menggunakan Naive Bayes Classifier,” J. Teknol. Inf. dan Ilmu Komput., vol. 11, no. 5, pp. 1033–1042, 2024, doi: 10.25126/jtiik.2024117989.

A. Firizkiansah, A. Muhammad, and I. R. Maulana, “Optimasi Klasifikasi Data Teks Menggunakan Algoritma Logistic Regression dengan TF-IDF dan SMOTE,” JIKOMTI J. Ilm. Ilmu Komput. dan Teknol. Inf., vol. 2, no. 1, pp. 29–36, 2025.

M. Ozkan, “Toward Automated Quality Control: An End-to-End Image Processing and Machine Learning Pipeline for Bullet Case Mouth Defect Detection,” IEEE Access, vol. 13, pp. 162764–162778, 2025, doi: 10.1109/ACCESS.2025.3610358.

R. Kurniawan et al., “Hypertension prediction using machine learning algorithm among Indonesian adults,” IAES Int. J. Artif. Intell., vol. 12, no. 2, pp. 776–784, 2023, doi: 10.11591/ijai.v12.i2.pp776-784.

M. Jawahar et al., “Computer-aided diagnosis of COVID-19 from chest X-ray images using histogramoriented gradient features and Random Forest classifier,” Multimed. Tools Appl., vol. 81, no. 28, pp. 40451–40468, 2022, doi: 10.1007/s11042-022-13183-6.

Downloads

Published

2026-09-01

PlumX Metrics

Scite Metrics

Altmetric

How to Cite

[1]
I. Riham Nurrahma and Nuk Ghurroh Setyoningrum, “Optimization of the Decision Tree Algorithm Using SMOTE for the Classification of Diabetes Mellitus Types Based on Clinical Medical Record Data”, JKBTI, vol. 5, no. 3, pp. 405–416, Sep. 2026.