ANALYSIS ON MACHINE LEARNING MODELS ROBUSTNESS AGAINST NOISE AND CONCEPT DRIFT IN THE CICIDS2017 DATASET

Authors

Abstract

Intrusion Detection Systems (IDS) based on machine learning face significant challenges in real-world deployment due to noise in network data and concept drift caused by evolving attack patterns. This study analyzes the robustness of three machine learning models Logistic Regression, Decision Tree, and Random Forest against noise and concept drift using the CICIDS2017 dataset. The experimental design includes baseline testing, noise robustness testing with three intensity levels (5%, 10%, 20%), and concept drift evaluation through temporal data splitting. Performance is measured using accuracy, precision, recall, and F1-score metrics, with robustness score calculated as a weighted average (40% baseline, 30% noise, 30% drift). Results show Logistic Regression achieves the highest robustness score (60.07%) due to excellent noise tolerance (1.09% F1-score degradation), followed by Random Forest (59.65%) and Decision Tree (57.60%). However, all models are categorized as NOT ROBUST against concept drift with degradation ranging from 29.64% to 30.34%, indicating sudden drift between Thursday and Friday data. The findings reveal that noise robustness and concept drift robustness are independent characteristics that do not correlate. Logistic Regression is recommended for practical deployment due to its optimal combination of robustness, interpretability, and computational efficiency. This research contributes to understanding model stability under non-ideal data conditions and emphasizes the necessity of implementing adaptive mechanisms such as periodic retraining and drift detection in operational IDS architecture.

Downloads

Download data is not yet available.

Author Biography

Azhar Bintang Pramudyanto, Amikom Purwokerto University

Azhar Bintang Pramudyanto is a student at the Faculty of Computer Science, Universitas Amikom Purwokerto. His research interests include machine learning, network security, and intrusion detection systems.

References

Imanuel Toding Bua and Nur Isdah Idris, “Analisis Kebijakan Keamanan Siber di Indonesia: Studi Kasus Kebocoran Data Nasional pada Tahun 2024,” Desentralisasi J. Hukum, Kebijak. Publik, dan Pemerintah., vol. 2, no. 2, pp. 100–114, May 2025, doi: 10.62383/desentralisasi.v2i2.653.

D. M. Manias, A. Chouman, and A. Shami, “Model Drift in Dynamic Networks,” IEEE Commun. Mag., vol. 61, no. 10, pp. 78–84, Oct. 2023, doi: 10.1109/MCOM.003.2200306.

J. Smithie, “Intrusion detection in cybersecurity,” Adv. Eng. Innov., vol. 2, no. 1, pp. 13–16, Oct. 2023, doi: 10.54254/2977-3903/2/2023014.

R. Morzelona and R. R. Mirajkar, “Implementation and Evaluation of Intrusion Detection Systems using Machine Learning Classifiers on Network Traffic Data,” Res. J. Comput. Syst. Eng., vol. 4, no. 2, pp. 103–116, Dec. 2023, doi: 10.52710/rjcse.81.

I. Maltseva, Y. Chernysh, and Y. Protsyuk, “Development of algorithms for early detection of cyberattacks on networks using machine learning,” Commun. Informatiz. cybersecurity Syst. Technol., vol. 1, no. 6, pp. 105–115, Dec. 2024, doi: 10.58254/viti.6.2024.08.105.

E. Mousavipour and A. Dimanchev, “Adaptive Anomaly Detection in Evolving Network Environments”.

E. Victor, L. Barboza, P. Ricardo, and L. De Almeida, “Challenges on Classifying Data Streams with Concept Drift,” no. September 2022, pp. 126–132, 2024.

W. Yang and J. Guo, A Concept Drift Detection Approach Based on Jensen-Shannon Divergence for Network Traffic Classification, vol. 1, no. 1. Association for Computing Machinery. doi: 10.1145/3573942.3573979.

M. Catillo, A. Del Vecchio, A. Pecchia, and U. Villano, “A Case Study with CICIDS2017 on the Robustness of Machine Learning against Adversarial Attacks in Intrusion Detection,” ACM Int. Conf. Proceeding Ser., 2023, doi: 10.1145/3600160.3605031.

P. Neirz, H. Allende, and C. Saavedra, “Attribute Relevance Score: A Novel Measure for Identifying Attribute Importance,” Algorithms, vol. 17, no. 11, p. 518, Nov. 2024, doi: 10.3390/a17110518.

S. Gaïffas, I. Merad, and Y. Yu, “WildWood: A New Random Forest Algorithm,” IEEE Trans. Inf. Theory, vol. 69, no. 10, pp. 6586–6604, Oct. 2023, doi: 10.1109/TIT.2023.3287432.

F. Acito, “Logistic Regression,” in Predictive Analytics with KNIME, Cham: Springer Nature Switzerland, 2023, pp. 125–167. doi: 10.1007/978-3-031-45630-5_7.

O. Graham and J. Lloris, “Comparative Study of Supervised Learning Algorithms for Intrusion Detection with a Focus on Logistic Regression.” May 07, 2025. doi: 10.20944/preprints202505.0420.v1.

Koushik Paul, Sayandeep Paik, Siddhartha Kuri, Soumyadip Majumder, and Avijit Kumar Chaudhuri, “A Novel Intrusion Detection System Using Multiple Linear Regression,” Int. J. Eng. Technol. Manag. Sci., vol. 7, no. 2, pp. 75–86, 2023, doi: 10.46647/ijetms.2023.v07i02.010.

R. Kimanzi, P. Kimanga, D. Cherori, and P. K. Gikunda, “Deep Learning Algorithms Used in Intrusion Detection Systems -- A Review,” 2024, [Online]. Available: http://arxiv.org/abs/2402.17020

E. Gilmore, V. Estivill-Castro, and R. Hexel, “More Interpretable Decision Trees,” 2021, pp. 280–292. doi: 10.1007/978-3-030-86271-8_24.

Galuh Nurvinda K, “Mengenal Algoritma Terpopuler dalam Data Science 2024,” DQLAB, 2024. https://dqlab.id/mengenal-algoritma-terpopuler-dalam-data-science-2024#:~:text=3. Algoritma Decision Tree Decision Tree adalah,dan hasilnya disajikan dalam bentuk pohon keputusan.

A. A. A. Hadi and A. M. Hadi, “Improving cybersecurity with random forest algorithm-based big data intrusion detection system: A performance analysis,” 2024, p. 040012. doi: 10.1063/5.0191707.

A. T Devi, L. Tukaram, and R. Purad, “Enhancing Intrusion Detection with Principal Component Analysis and Random Forest,” Int. J. Sci. Eng. Res., pp. 10–13, Sep. 2025, doi: 10.70729/SE25904191409.

O. H. Ram, N. Gorea, A. Reif, L. Bonasera, and S. Augustin, “The Role of Noisy Data in Improving CNN Robustness for Image Classification,” vol. 13606, pp. 1–16, 2025, doi: 10.1117/12.3063563.

S. Greco, B. Vacchetti, D. Apiletti, and T. Cerquitelli, “Unsupervised Concept Drift Detection from Deep Learning Representations in Real-time,” pp. 1–23, 2025.

D. Krstinić, A. K. Skelin, I. Slapničar, and M. Braović, “Multi-Label Confusion Tensor,” IEEE Access, vol. 12, pp. 9860–9870, 2024, doi: 10.1109/ACCESS.2024.3353050.

S. Sathyanarayanan, “Confusion Matrix-Based Performance Evaluation Metrics,” African J. Biomed. Res., pp. 4023–4031, Nov. 2024, doi: 10.53555/AJBR.v27i4S.4345.

Published

2026-07-22

Issue

Section

Artikel