TARU PUBLICATIONS
Journal of Discrete Mathematical Sciences and Cryptography cover
Open Access ·Peer-reviewed·ISSN (Online): 2169-0065·ISSN (Print): 0972-0529

Monthly Journal: Publishes theoretical and applied research in all areas of Discrete Mathematical Sciences, Cryptography, Combinatorics, Elliptic Curves and Information Security.

Issues up to 2022 co-published with and available at:Taylor & Francis Online
submissions@tarupublications.com
Open Access Research Article

Robust cybersecurity through discrete and large language models for effective phishing attack detection

, , , , *

* Corresponding author · click or hover a name for details

pp. 1621–1638Vol. 28Issue 5-AAugust 2025DOI: 10.47974/JDMSC-2162 Crossmark XML
Received:
06 Nov 2024
Published Online:
30 Aug 2025
Article type:
Research Article
Language:
EN
Article no.:
JDMSC-2162
Pages:
1621–1638

Abstract

This study explores the efficacy of four Large Language Models (LLMs)—BERT, DistilBERT, RoBERTa, and DeBERTa—in classifying URLs as either legitimate or phishing. The research methodology is structured into three phases: dataset processing, model fine-tuning, and performance evaluation. Each LLM is fine-tuned to distinguish between phishing and legitimate URLs. The models are evaluated using both a primary dataset with extensive features and an external dataset with minimal features to rigorously assess their robustness. The models consistently achieved high performance on the primary dataset, with AUC scores of 0.99, indicating near-perfect discrimination between phishing and legitimate URLs. DistilBERT excels in F1-score (99.992%), accuracy (99.991%), and precision (99.985%), showcasing its efficiency in real-world applications. BERT and DeBERTa also demonstrate excellent results, while RoBERTa, though slightly lower in precision, remains competitive. The model’s performance was also evaluated on an external test dataset containing 450,176 labeled URLs, the dataset helped access the model’s performance under extremely constrained conditions. BERT and DeBERTa showed low F1-scores (1.14% and 2.78%, respectively) despite high precision, indicating poor recall. DistilBERT performs moderately better with a 42.44% F1-score, while RoBERTa achieves the highest F1-score of 68.03%, suggesting superior balance between precision and recall under severe feature constraints. Overall, while all models exhibit strong performance with rich features, their ability to maintain efficacy under limited feature conditions varies. This study underscores the importance of developing models with robust performance across diverse scenarios, highlighting RoBERTa’s superior recall in feature-scarce environments and DistilBERT’s overall efficiency.

Keywords

Subject Classifications

68T0768P25

References

[1] P. Yang, G. Zhao, and P. Zeng, “Phishing website detection based on multidimensional features driven by deep learning,” IEEE Access, vol. 7, pp. 15196–15209 (2019).
[2] P. Yi, Y. Guan, F. Zou, Y. Yao, W. Wang, and T. Zhu, “Web phishing detection using a deep learning framework,” Wireless Communications and Mobile Computing, vol. 2018, Article ID 4678746 (2018).
[3] M. A. A. Siddiq, M. Arifuzzaman, and M. S. Islam, “Phishing website detection using deep learning,” in Proceedings of the 2nd International Conference on Computing Advancements (ICCA), pp. 83–88 (2022).
[4] F. Greco, G. Desolda, A. Esposito, and A. Carelli, “David versus Goliath: Can machine learning detect LLM-generated text? A case study in the detection of phishing emails,” in Proceedings of The Italian Conference on CyberSecurity (ITASEC) (2024).
[5] F. Trad and A. Chehab, “Prompt engineering or fine-tuning? A case study on phishing detection with large language models,” Machine Learning and Knowledge Extraction, vol. 6, no. 1, pp. 367–384 (2024).
[6] M. Bethany, A. Galiopoulos, E. Bethany, M. B. Karkevandi, N. Vishwamitra, and P. Najafirad, “Large language model lateral spear phishing: A comparative study in large-scale organizational settings,” arXiv preprint arXiv:2401.09727 (2024).
[7] A. Prasad and S. Chandra, “PhiUSIIL Phishing URL (Website),” UCI Machine Learning Repository (2024). [Online]. Available: https://doi.org/10.1016/j.cose.2023.103545
[8] M. M. Uddin, N. Jahan, F. M. Shahin, M. R. Hashem, and M. S. Kaiser, “A comparative analysis of machine learning-based website phishing detection using URL information,” in Proceedings of the 5th International Conference on Pattern Recognition and Artificial Intelligence (PRAI), pp. 69–74 (2022).
[9] J. Devlin, M. W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” arXiv preprint arXiv:1810.04805 (2018).
[10] V. Sanh, L. Debut, J. Chaumond, and T. Wolf, “DistilBERT, a distilled version of BERT: Smaller, faster, cheaper and lighter,” arXiv preprint arXiv:1910.01108 (2019).
[11] Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, and V. Stoyanov, “RoBERTa: A robustly optimized BERT pretraining approach,” arXiv preprint arXiv:1907.11692 (2019).
[12] P. He, X. Liu, J. Gao, and W. Chen, “DeBERTa: Decoding-enhanced BERT with disentangled attention,” arXiv preprint arXiv:2006.03654 (2020).
[13] T. Wolf, L. Debut, V. Sanh, J. Chaumond, C. Delangue, A. Moi, P. Cistac, T. Rault, R. Louf, M. Funtowicz, and A. M. Rush, “HuggingFace’s transformers: State-of-the-art natural language processing,” arXiv preprint arXiv:1910.03771 (2019).
[14] I. Loshchilov and F. Hutter, “Fixing weight decay regularization in Adam,” arXiv preprint arXiv:1711.05101 (2017).
[15] D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” arXiv preprint arXiv:1412.6980 (2014).
[16] D. Goyal, F. Sheth, P. Mathur, and A. K. Gupta, “Discrete mathematical models for enhancing cybersecurity: A mathematical and statistical analysis of machine learning approaches in phishing attack detection,” Journal of Discrete Mathematical Sciences and Cryptography, vol. 27, no. 2-B, pp. 569–599 (2024).
[17] P. Mathur, F. Sheth, D. Goyal, and A. K. Gupta, “Deep insight: Mathematical modeling and statistical analysis for mango leaf disease classification using advanced deep learning models,” Journal of Interdisciplinary Mathematics, vol. 27, no. 2, pp. 317–342 (2024).
[18] A. Kumar, A. K. Gupta, D. Panwar, S. Chaurasia, and D. Goyal, “Operating system security with discrete mathematical structure for secure round robin scheduling method with intelligent time quantum,” (2023)
[19] J. K. S. Kaitholikkal and B. Arthi, “Phishing URL dataset,” Mendeley Data, V1 (2024). [Online]. Available: https://doi.org/10.17632/vfszbj9b36.1
[20] T. Teubner, C. M. Flath, C. Weinhardt, W. van der Aalst, and O. Hinz, “Welcome to the era of ChatGPT et al.: The prospects of large language models,” Business and Information Systems Engineering, vol. 65, no. 2, pp. 95–101 (2023).
[21] P. Wang, L. Li, L. Chen, F. Song, B. Lin, Y. Cao, J. Li, Q. He, J. Wang, and Z. Sui, “Making large language models better reasoners with alignment,” arXiv preprint arXiv:2309.02144 (2023).
[22] C. Song, T. Ristenpart, and V. Shmatikov, “Machine learning models that remember too much,” in Proceedings of the 2017 ACM SIGSAC Conference on Computer and Communications Security, pp. 587–601 (2017).

Views: 220Downloads: 10Citations: 0