A Dual-Pipeline Imbalance-Robust Framework for SMS Spam Detection: Achieving Flawless Precision via SMOTE-Augmented Ensembles with Rigorous Statistical Validation
DOI:
https://doi.org/10.69916/comtechno.v4i1.510Keywords:
SMS Spam Detection, Machine Learning, TF-IDF Vectorization, SMOTE Oversampling, Statistical ValidationAbstract
The rapid proliferation of digital communication has exponentially increased the volume of Short Message Service (SMS) spam, exposing mobile users to systemic convenience disruptions, productivity drops, and severe financial losses through sophisticated fraudulent schemes. To construct a highly dependable filtering mechanism, this study presents a rigorous dual-pipeline machine learning framework that systematically addresses the challenges of class imbalance in statistical text mining. Utilizing a verified dataset of 5,572 Indonesian-context short messages, the raw textual corpus is subjected to uniform case normalization, structural URL extraction, and character filtering before feature projection via Term Frequency–Inverse Document Frequency (TF-IDF) vectorization. To overcome the inherent accuracy paradox of skewed class distributions, the experimental design evaluates a baseline pipeline (imbalanced data) against a synthetic data augmentation pipeline leveraging the Synthetic Minority Oversampling Technique (SMOTE) across four distinct classifiers: Logistic Regression, Naive Bayes, Linear Support Vector Machine (Linear SVM), and Random Forest. Empirical results demonstrate that while the baseline Linear SVM serves as the optimal standalone model for overall balance, achieving a peak accuracy of 98.11% and a dominant F1-Score of 92.83%, the SMOTE-augmented Random Forest configuration yields an exceptional high-security alternative by securing a flawless 100.00% precision envelope alongside an 83.89% recall rate. Advanced post-hoc evaluations including McNemar's statistical significance tests (, for Random Forest), qualitative error analyses of semantic edge cases, and runtime profiling confirm that the developed architecture establishes a highly scalable, mathematically verified, and low-latency solution suitable for integration into real-time telecom filtering gateways.
References
N. K. Nagwani and A. Sharaff, “SMS spam filtering and thread identification using bi-level text classification and clustering techniques,” J. Inf. Sci., vol. 43, no. 1, pp. 75–87, Feb. 2017, doi: 10.1177/0165551515616310.
N. N. Amir Sjarif, Y. Yahya, S. Chuprat, and N. H. F. Mohd Azmi, “Support Vector Machine Algorithm for SMS Spam Classification in The Telecommunication Industry,” Int. J. Adv. Sci. Eng. Inf. Technol., vol. 10, no. 2, pp. 635–639, Apr. 2020, doi: 10.18517/ijaseit.10.2.10175.
K. Shaukat, S. Luo, V. Varadharajan, I. A. Hameed, and M. Xu, “A Survey on Machine Learning Techniques for Cyber Security in the Last Decade,” IEEE Access, vol. 8, pp. 222310–222354, 2020, doi: 10.1109/ACCESS.2020.3041951.
H. AbouGrad, S. Chakhar, and A. Abubahia, “Decision Making by Applying Machine Learning Techniques to Mitigate Spam SMS Attacks,” 2023, pp. 154–166. doi: 10.1007/978-3-031-30396-8_14.
P. A. Raharja, M. F. Sidiq, and D. C. Fransisca, “Comparative Analysis of Multinomial Naïve Bayes and Logistic Regression Models for Prediction of SMS Spam,” JURNAL MEDIA INFORMATIKA BUDIDARMA, vol. 6, no. 3, p. 1290, Jul. 2022, doi: 10.30865/mib.v6i3.4019.
N. Al Moubayed, T. Breckon, P. Matthews, and A. S. McGough, “SMS Spam Filtering Using Probabilistic Topic Modelling and Stacked Denoising Autoencoder,” 2016, pp. 423–430. doi: 10.1007/978-3-319-44781-0_50.
W. A. Prabowo and F. Azizah, “Sentiment Analysis for Detecting Cyberbullying Using TF-IDF and SVM,” Jurnal RESTI (Rekayasa Sistem dan Teknologi Informasi), vol. 4, no. 6, Dec. 2020, doi: 10.29207/resti.v4i6.2753.
Dedy Sugiarto, Ema Utami, and Ainul Yaqin, “Perbandingan Kinerja Model TF-IDF dan BOW untuk Klasifikasi Opini Publik Tentang Kebijakan BLT Minyak Goreng,” Jurnal Teknik Industri, vol. 12, no. 3, pp. 272–277, 2022, doi: 10.25105/jti.v12i3.15669.
J. Piskorski and G. Jacquet, “TF-IDF Character N-grams versus Word Embedding-based Models for Fine-grained Event Classification: A Preliminary Study,” Proceedings of the Workshop on Automated Extraction of Socio-political Events from News 2020, no. May, pp. 26–34, 2020.
N. Latifah, R. Dwiyansaputra, and G. S. Nugraha, “Multiclass Text Classification of Indonesian Short Message Service (SMS) Spam using Deep Learning Method and Easy Data Augmentation,” MATRIK : Jurnal Manajemen, Teknik Informatika dan Rekayasa Komputer, vol. 23, no. 3, pp. 663–676, Jun. 2024, doi: 10.30812/matrik.v23i3.3835.
O. Abayomi‐Alli, S. Misra, and A. Abayomi‐Alli, “A deep learning method for automatic SMS spam classification: Performance of learning algorithms on indigenous dataset,” Concurr. Comput., vol. 34, no. 17, Aug. 2022, doi: 10.1002/cpe.6989.
A. TEKEREK, “Support Vector Machine Based Spam SMS Detection,” Politeknik Dergisi, vol. 22, no. 3, pp. 779–784, Sep. 2019, doi: 10.2339/politeknik.429707.
Y. HaCohen-Kerner, D. Miller, and Y. Yigal, “The influence of preprocessing on text classification using a bag-of-words representation,” PLoS One, vol. 15, no. 5, p. e0232525, May 2020, doi: 10.1371/journal.pone.0232525.
U. Maqsood, S. Ur Rehman, T. Ali, K. Mahmood, T. Alsaedi, and M. Kundi, “An Intelligent Framework Based on Deep Learning for SMS and e-mail Spam Detection,” Applied Computational Intelligence and Soft Computing, vol. 2023, pp. 1–16, Sep. 2023, doi: 10.1155/2023/6648970.
M. S. Al Ghofany, R. Dwiyansaputra, F. Bimantoro, and Khairunnas, “Indonesian SMS Spam Detection Using TF-RF Feature Weighting Method and Support Vector Machine Classifier,” in Proceedings of the First Mandalika International Multi-Conference on Science and Engineering 2022, MIMSE 2022 (Informatics and Computer Science), Dordrecht: Atlantis Press International BV, 2022, pp. 117–129. doi: 10.2991/978-94-6463-084-8_12.
M. Mujahid et al., “Data oversampling and imbalanced datasets: an investigation of performance for machine learning and feature engineering,” J. Big Data, vol. 11, no. 1, p. 87, Jun. 2024, doi: 10.1186/s40537-024-00943-4.
F. M. Rizky, J. Jondri, and K. M. Lhaksmana, “Twitter Sentiment Analysis of Kanjuruhan Disaster using Word2Vec and Support Vector Machine,” Building of Informatics, Technology and Science (BITS), vol. 5, no. 1, Jun. 2023, doi: 10.47065/bits.v5i1.3612.
Y. Wang, “An XGBoost-Based Cyber Threat Detection Framework for Enhancing Security in University E-Government Systems,” SECURITY AND PRIVACY, vol. 8, no. 5, Sep. 2025, doi: 10.1002/spy2.70089.
P. Joseph and S. Y. Yerima, “A comparative study of word embedding techniques for SMS spam detection,” in 2022 14th International Conference on Computational Intelligence and Communication Networks (CICN), IEEE, Dec. 2022, pp. 149–155. doi: 10.1109/CICN56167.2022.10008245.
G. Waja, G. Patil, C. Mehta, and S. Patil, “How AI Can be Used for Governance of Messaging Services: A Study on Spam Classification Leveraging Multi-Channel Convolutional Neural Network,” International Journal of Information Management Data Insights, vol. 3, no. 1, p. 100147, Apr. 2023, doi: 10.1016/j.jjimei.2022.100147.
O. Abayomi‐Alli, S. Misra, and A. Abayomi‐Alli, “A deep learning method for automatic SMS spam classification: Performance of learning algorithms on indigenous dataset,” Concurr. Comput., vol. 34, no. 17, Aug. 2022, doi: 10.1002/cpe.6989.
U. Srinivasarao and A. Sharaff, “Machine intelligence based hybrid classifier for spam detection and sentiment analysis of SMS messages,” Multimed. Tools Appl., vol. 82, no. 20, pp. 31069–31099, Aug. 2023, doi: 10.1007/s11042-023-14641-5.
M. Sharabov, G. Tsochev, V. Gancheva, and A. Tasheva, “Filtering and Detection of Real-Time Spam Mail Based on a Bayesian Approach in University Networks,” Electronics (Basel)., vol. 13, no. 2, p. 374, Jan. 2024, doi: 10.3390/electronics13020374.
Downloads
Published
Scite Metrics
Altmetric
How to Cite
Issue
Section
License
Copyright (c) 2026 Zulpan Hadi, Selamet Riadi, Supardianto, Aulia Riswanti Naya, Liana Trihardianingsih

This work is licensed under a Creative Commons Attribution 4.0 International License.











