info@ijrretas.com
+91 77710 84928 Support
ijrretas Logo
  • About
    • About The Journal
    • Aim & Scope
    • Privacy Statement
    • Journal Policies
    • Disclaimer
    • Abstracting and Indexing
    • FAQ
  • Current
  • Archive
  • For Author
    • Submit Paper Online
    • Article Processing Charges
    • Submission Guidelines
    • Manuscript Types
    • Download Article Template
    • Download Copyright Form
  • Editorial Board
    • Editorial Board
    • Editors Responsibilities
  • More
    • Conference
    • Contact
    • Pay Online
    • Blogs
  • Submit Paper

Recent Papers

Dedicated to advancing knowledge through rigorous research and scholarly publication

  1. Home
  2. Recent Papers

Authors: Rohit Chandra, Dr. Krishna Murari

Certificate: View Certificate

Abstract

The growing volume and complexity of network traffic have made manual and signature-based monitoring insufficient for detecting malicious activity. This paper presents a robust machine learning framework, based on the Light Gradient Boosting Machine (LightGBM), for detecting outliers in network traffic data. The framework integrates data cleaning, one-hot encoding, logarithmic transformation, mutual-information-based feature selection, standardisation and a regularised LightGBM classifier trained with early stopping. It was evaluated on the NSL-KDD benchmark and compared with Logistic Regression, Random Forest and a Support Vector Machine under identical conditions. On a held-out test set of 25,193 connections, LightGBM achieved an accuracy of 99.933%, a precision of 99.957%, a recall of 99.898%, an F1-score of 99.927% and a ROC-AUC of 0.99999, making only 17 errors. It outperformed all baseline models while training about four times faster than Random Forest. SHAP-based interpretation showed that byte volume, connection rate, service concentration and SYN-error features drive detection. Robustness experiments confirmed stable performance across data splits and random seeds, and resilience to severe class imbalance, with an F1-score of 97.74% when outliers formed only 1% of traffic. On an independent test set containing unseen attack types, accuracy fell to 78.54%, identifying generalisation to novel attacks as the main limitation. The results show that LightGBM offers an accurate, efficient and interpretable foundation for network outlier detection.

Introduction

1.1 Background of Network Security

Computer networks underpin banking, healthcare, government, commerce and personal communication, and the continuous growth of cloud services, mobile access and connected devices has made them larger and more complex than ever. This openness creates value but also exposes every connected system to potential attack. Denial of Service attacks exhaust the resources of targets, probing attacks map networks for weaknesses, and intrusion attacks attempt to gain unauthorised access or escalate privileges. Because traditional perimeter defences such as firewalls only enforce rules about permitted traffic, organisations increasingly rely on continuous monitoring to identify malicious activity hidden within legitimate flows (Khraisat et al., 2019).

The scale of this challenge continues to grow. Encrypted protocols now protect most web traffic, which prevents inspection of payloads and forces detection to rely on observable metadata such as byte counts, connection timing and service usage. At the same time, automated attack tools and botnets allow even unskilled actors to generate large volumes of hostile traffic. Monitoring systems must therefore examine every connection quickly, identify suspicious behaviour from connection-level characteristics and present analysts with a manageable number of reliable alerts.

1.2 Network Outliers and Anomalous Traffic

A network outlier is a connection or traffic pattern that departs from the behaviour produced by normal network use. Such departures may be caused by attacks, misconfigurations or unusual but benign events. Outliers are rarely extreme in a single attribute; more often they are defined by unusual combinations of attributes, such as a connection that transfers no data while forming part of a burst of connections that end in errors. In this paper, every connection labelled as an attack is treated as an outlier and every legitimate connection as normal, which allows outlier detection to be framed as a supervised binary classification problem.

Network outliers take several forms. Point anomalies are individual connections that are unusual in themselves, contextual anomalies are unusual only in a particular setting, and collective anomalies are groups of connections that are abnormal together even when each appears ordinary. Flooding and scanning attacks typically appear as collective anomalies, which is why the time-based and host-based traffic statistics that summarise recent connections to the same host or service are especially informative. Intrusions into user sessions, by contrast, often appear as subtle point anomalies that closely resemble normal activity.

1.3 Intrusion Detection Approaches

Intrusion detection systems monitor traffic and raise alerts when suspicious behaviour is observed. Signature-based systems compare traffic with known attack patterns and are precise, but they cannot recognise attacks for which no signature exists. Anomaly-based systems model the difference between normal and abnormal behaviour and can generalise beyond known patterns, but they have historically suffered from high false alarm rates (Khraisat et al., 2019; Liu & Lang, 2019). Modern traffic volumes also require detection to be fast and automated, since analysts cannot inspect millions of connections manually, and alerts must be accurate enough to avoid overwhelming them.

The effectiveness of any detection approach is ultimately judged by two competing error rates. A missed attack may allow an intruder to establish persistent access, while a false alarm consumes analyst time and, if frequent, leads to alert fatigue in which genuine alerts are ignored. An effective detector must keep both error rates low simultaneously, and it must continue to do so when attacks are rare, when traffic changes over time and when new attack types appear.

1.4 Machine Learning for Anomaly Detection

Machine learning enables detectors to learn complex relationships among many traffic features directly from labelled data and to classify new connections in microseconds. A wide range of algorithms has been applied to the task, from linear models and support vector machines to random forests, gradient boosting and deep neural networks (Ahmad et al., 2021). However, recent critical work has shown that many evaluations rely on single random splits of benchmark data, omit computational cost and seldom test generalisation, which can overstate real-world performance (Arp et al., 2022). A credible detection framework must therefore be assessed for efficiency, interpretability and robustness as well as accuracy.

Machine learning approaches differ in the kinds of relationships they can represent and in their computational cost. Linear models are fast and transparent but cannot capture the non-linear and interacting patterns typical of attacks. Kernel methods can learn non-linear boundaries but scale poorly to large datasets. Deep neural networks are highly expressive but demand substantial data, tuning and computing resources, and their decisions are difficult to explain. Tree-based ensembles combine flexibility with efficiency and offer natural measures of feature importance, which makes them attractive for tabular connection records.

1.5 LightGBM and Gradient Boosting

LightGBM is a gradient boosting decision tree framework designed for speed and memory efficiency on large datasets (Ke et al., 2017). It builds an ensemble of trees sequentially, with each tree correcting the errors of the previous ones, and it introduces histogram-based split finding, gradient-based one-side sampling, exclusive feature bundling and leaf-wise tree growth. These properties suit network traffic, which is voluminous, heterogeneous, skewed and sparse after categorical encoding. LightGBM also supports exact SHAP explanations of individual predictions (Lundberg et al., 2020). This paper develops and systematically evaluates a LightGBM-based framework for network outlier detection, comparing it with conventional classifiers and assessing its robustness under varied conditions.

Leaf-wise growth is particularly relevant to network data. Instead of expanding every node at a given depth, LightGBM splits the leaf that yields the greatest reduction in loss, which allows it to concentrate model capacity on the difficult boundary between normal sessions and subtle attacks while separating distinctive flooding traffic with only a few splits. Because this strategy can overfit, the framework in this paper controls tree complexity through leaf limits, minimum leaf sizes, subsampling, regularisation and validation-based early stopping.

Conclusion

This paper developed and evaluated a robust LightGBM-based framework for network outlier detection. By combining leakage-free pre-processing, mutual-information-based feature selection and a regularised LightGBM model with early stopping, the framework achieved 99.933?curacy with only 5 false alarms and 12 missed attacks among 25,193 test connections. It outperformed Logistic Regression, an SVM and Random Forest on every metric while training about four times faster than Random Forest. SHAP analysis showed that its decisions rest on interpretable traffic behaviours, namely little or no data transfer, bursts of connections, SYN errors and scattering across services. The framework was stable across splits, seeds and folds and resilient to severe class imbalance. Its main limitation is reduced recall on attack types absent from training, as shown on KDDTest+. Future work should validate the framework on modern datasets, recalibrate thresholds under distribution shift, combine it with unsupervised detectors in a layered architecture, and extend it to multi-class detection. Overall, LightGBM provides an accurate, efficient and interpretable foundation for practical network anomaly detection. The findings have practical implications for network defenders. The very low false alarm rate and microsecond-level latency make real-time deployment feasible, while SHAP explanations allow analysts to see why each connection was flagged, which supports faster triage and greater trust. The small number of features needed for high accuracy also allows lightweight models to be deployed on resource-constrained devices such as edge gateways. Researchers, for their part, should evaluate detectors on independent data containing unseen attacks and report computational cost alongside accuracy, so that published results better reflect operational conditions.

References

1. Ahmad, Z., Shahid Khan, A., Wai Shiang, C., Abdullah, J., & Ahmad, F. (2021). Network intrusion detection system: A systematic study of machine learning and deep learning approaches. Transactions on Emerging Telecommunications Technologies, 32(1), e4150. 2. Arp, D., Quiring, E., Pendlebury, F., Warnecke, A., Pierazzi, F., Wressnegger, C., Cavallaro, L., & Rieck, K. (2022). Dos and don'ts of machine learning in computer security. In Proceedings of the 31st USENIX Security Symposium (pp. 3971–3988). USENIX Association. 3. Chicco, D., & Jurman, G. (2020). The advantages of the Matthews correlation coefficient (MCC) over F1 score and accuracy in binary classification evaluation. BMC Genomics, 21, Article 6. 4. Engelen, G., Rimmer, V., & Joosen, W. (2021). Troubleshooting an intrusion detection dataset: The CICIDS2017 case study. In 2021 IEEE Security and Privacy Workshops (SPW) (pp. 7–12). IEEE. 5. Grinsztajn, L., Oyallon, E., & Varoquaux, G. (2022). Why do tree-based models still outperform deep learning on typical tabular data? In Advances in Neural Information Processing Systems 35 (pp. 507–520). 6. Jin, D., Lu, Y., Qin, J., Cheng, Z., & Mao, Z. (2020). SwiftIDS: Real-time intrusion detection system based on LightGBM and parallel intrusion detection mechanism. Computers & Security, 97, 101984. https://doi.org/10.1016/j.cose.2020.101984 7. Ke, G., Meng, Q., Finley, T., Wang, T., Chen, W., Ma, W., Ye, Q., & Liu, T.-Y. (2017). LightGBM: A highly efficient gradient boosting decision tree. In Advances in Neural Information Processing Systems 30 (pp. 3146–3154). 8. Khraisat, A., Gondal, I., Vamplew, P., & Kamruzzaman, J. (2019). Survey of intrusion detection systems: Techniques, datasets and challenges. Cybersecurity, 2, Article 20. 9. Kilincer, I. F., Ertam, F., & Sengur, A. (2021). Machine learning methods for cyber security intrusion detection: Datasets and comparative study. Computer Networks, 188, 107840. 10. Liu, H., & Lang, B. (2019). Machine learning and deep learning methods for intrusion detection systems: A survey. Applied Sciences, 9(20), 4396. 11. Liu, J., Gao, Y., & Hu, F. (2021). A fast network intrusion detection system using adaptive synthetic oversampling and LightGBM. Computers & Security, 106, 102289. https://doi.org/10.1016/j.cose.2021.102289 12. Lundberg, S. M., Erion, G., Chen, H., DeGrave, A., Prutkin, J. M., Nair, B., Katz, R., Himmelfarb, J., Bansal, N., & Lee, S.-I. (2020). From local explanations to global understanding with explainable AI for trees. Nature Machine Intelligence, 2(1), 56–67. 13. Lundberg, S. M., & Lee, S.-I. (2017). A unified approach to interpreting model predictions. In Advances in Neural Information Processing Systems 30 (pp. 4765–4774). 14. Resende, P. A. A., & Drummond, A. C. (2018). A survey of random forest based methods for intrusion detection systems. ACM Computing Surveys, 51(3), Article 48. 15. Ring, M., Wunderlich, S., Scheuring, D., Landes, D., & Hotho, A. (2019). A survey of network-based intrusion detection data sets. Computers & Security, 86, 147–167. 16. Seth, S., Singh, G., & Kaur Chahal, K. (2021). A novel time efficient learning-based approach for smart intrusion detection system. Journal of Big Data, 8, Article 111. https://doi.org/10.1186/s40537-021-00498-8 17. Sharafaldin, I., Habibi Lashkari, A., & Ghorbani, A. A. (2018). Toward generating a new intrusion detection dataset and intrusion traffic characterization. In Proceedings of the 4th International Conference on Information Systems Security and Privacy (ICISSP) (pp. 108–116). SciTePress. 18. Talukder, M. A., Islam, M. M., Uddin, M. A., Hasan, K. F., Sharmin, S., Alyami, S. A., & Moni, M. A. (2024). Machine learning-based network intrusion detection for big and imbalanced data using oversampling, stacking feature embedding and feature extraction. Journal of Big Data, 11, Article 33. https://doi.org/10.1186/s40537-024-00886-w 19. Vinayakumar, R., Alazab, M., Soman, K. P., Poornachandran, P., Al-Nemrat, A., & Venkatraman, S. (2019). Deep learning approach for intelligent intrusion detection system. IEEE Access, 7, 41525–41550. 20. Yin, C., Zhu, Y., Fei, J., & He, X. (2017). A deep learning approach for intrusion detection using recurrent neural networks. IEEE Access, 5, 21954–21961.

Copyright

2026

Download Paper

Paper Id: IJRRETAS291

Publish Date: 2026-07-03

ISSN: 2455-4723

Publisher Name: ijrretas

About ijrretas

ijrretas is a leading open-access, peer-reviewed journal dedicated to advancing research in applied sciences and engineering. We provide a global platform for researchers to disseminate innovative findings and technological breakthroughs.

ISSN
2455-4723
Established
2015

Quick Links

Home Submit Paper Author Guidelines Editorial Board Past Issues Topics
Fees Structure Scope & Topics Terms & Conditions Privacy Policy Refund and Cancellation Policy

Contact Us

304 Siver Mall RNT Marg, Indore (M.P) - India

info@ijrretas.com

+91 77710 84928

www.ijrretas.com

Indexed In
Google Scholar Crossref DOAJ ResearchGate CiteFactor
© 2026 ijrretas. All Rights Reserved.
Privacy Policy Terms of Service