TARU PUBLICATIONS
Journal of Discrete Mathematical Sciences and Cryptography cover
Open Access ·Peer-reviewed·ISSN (Online): 2169-0065·ISSN (Print): 0972-0529

Monthly Journal: Publishes theoretical and applied research in all areas of Discrete Mathematical Sciences, Cryptography, Combinatorics, Elliptic Curves and Information Security.

Issues up to 2022 co-published with and available at:Taylor & Francis Online
submissions@tarupublications.com
Open Access Research Article

Phishing email detection using MBFT-Lite : A multi-branch federated transformer framework

, , , *

* Corresponding author · click or hover a name for details

pp. 1473–1483Vol. 29Issue 3March 2026DOI: 10.47974/JDMSC-2606 Crossmark XML
Received:
01 Jul 2025
Published Online:
11 Mar 2026
Article type:
Research Article
Language:
EN
Article no.:
JDMSC-2606
Pages:
1473–1483

Abstract

Phishing attacks continue to pose a significant cybersecurity threat by exploiting users through deceptive emails. While traditional centralized phishing detection methods achieve high accuracy, they face critical limitations, including privacy concerns, scalability challenges, and susceptibility to data distribution shifts. To address these issues, MBFT-Lite (Multi-Branch Federated Transformer-Lite) is proposed—a novel, lightweight, privacy-preserving phishing detection framework based on federated learning. MBFT-Lite employs a dual-branch architecture that combines semantic analysis via a TinyBERT model with structural analysis through XGBoost models on metadata features. The TinyBERT models are aggregated using the FedAvg a lgorithm, while the XGBoost models remain decentralized to enhance privacy and reduce communication overhead. Predictions from both branches are fused using strategies such as weighted averaging, max-confidence voting, and a meta-classifier. Extensive evaluations across four diverse email datasets—CEAS 08, SpamAssassin, Nigerian Fraud Emails, and Nazario—demonstrate that MBFT-Lite achieves over 99.4% accuracy across clients and 97.9% accuracy on unseen data, while preserving data privacy and ensuring scalability.

Keywords

Subject Classifications

Primary 93A30Secondary 49K15

References

[1] E. G. Dada, J. S. Bassi, H. Chiroma, S. I. M. Abdulhamid, A. O. Adetunmbi, and O. E. Ajibuwa, “Machine learning for email spam filtering: Review, approaches and open research problems,” Heliyon, vol. 5, no. 6 (Jun. 2019).
[2] Y. Fang, C. Zhang, C. Huang, L. Liu, and Y. Yang, “Phishing email detection using improved RCNN model with multilevel vectors and attention mechanism,” IEEE Access, vol. 7, pp. 56, 56329 - 56340 (2019).
[3] A. Das, S. Baki, A. El Aassal, R. Verma, and A. Dunbar, “SoK: A comprehensive reexamination of phishing research from the security perspective,” IEEE Commun. Surveys Tuts., vol. 22, no. 1, pp. 671–708 (2019).
[4] B. Y. Lin, C. He, Z. Ze, H. Wang, Y. Hua, C. Dupuy, and S. Avestimehr., “FedNLP: Benchmarking federated learning methods for natural language processing tasks,” in Findings of the Assoc. Comput. Linguistics: NAACL 2022, pp. 157–175 (Jul. 2022).
[5] L. U. Khan, S. R. Pandey, N. H. Tran, W. Saad, Z. Han, M. N. Nguyen, and C. S. Hong, “Federated learning for edge networks: Resource optimization and incentive mechanism,” IEEE Commun. Mag., vol. 58, no. 10, pp. 88–93 (Oct. 2020).
[6] K. Han, A. Xiao, E. Wu, J. Guo, C. Xu, and Y. Wang, “Transformer in transformer,” in Advances in Neural Information Processing Systems (NeurIPS), vol. 34, pp. 15908-15919 (2021).
[7] C. Thapa, J. W. Tang, A. Abuadbba, Y. Gao, S. Camtepe, S. Nepal, M. Almashor, and Y. Zheng, “Evaluation of federated learning in phishing email detection,” Sensors, vol. 23, no. 9, Art. no. 4346 (2023).
[8] M. Alazab, S. P. R. Rm, P. K. R. Maddikunta, T. R. Gadekallu, and Q. V. Pham, “Federated learning for cybersecurity: Concepts, challenges, and future directions,” IEEE Trans. Ind. Informatics, vol. 18, no. 5, pp. 3501–3509 (May 2021).
[9] M. Revathi, “Exploring the efficacy of federated-continual learning nodes with attention-based classifier for robust web phishing detection: An empirical investigation,” arXiv preprint arXiv:2405.03537 (2024).
[10] J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” in Proc. 2019 Conf. North American Chapter of the Assoc. Comput. Linguistics: Human Language Technologies (NAACL-HLT), pp. 4171–4186 (Jun. 2019).
[11] P. Dadheech, A. Kumar, and V. Singh, “An optimization of bitmap index compression technique in bulk data movement infrastructure,” TARU J. Sustain. Technol. Comput., vol. 1, no. 2, pp. 81–91 (2019), doi: 10.47974/2019.TJSTC.003.
[12] M. Vasti and A. Dev, “Ensemble learning based predictive modelling on a highly imbalanced multiclass data,” J. Inf. Optim. Sci., vol. 45, no. 8, pp. 2141–2164 (2024), doi: 10.47974/JIOS-1778.
[13] S. V. Devika, B. Unhelkar, S. S. Shankar, P. Chakrabarti, and S. Arvind, “Managerial perspectives: Utilizing MLP-CNN for predicting and classifying DDoS attacks,” J. Inf. Optim. Sci., vol. 45, no. 8, pp. 2071–2079 (2024), doi: 10.47974/JIOS-1651.

Views: 108Downloads: 11Citations: 0