TARU PUBLICATIONS
Journal of Information and Optimization Sciences cover
Open Access ·Peer-reviewed·ISSN (Online): 2169-0103·ISSN (Print): 0252-2667

WoS  JIF 2026 : 0.4 (Q4)

Powered by:Powered by

Monthly Journal: Publishes theoretical and applied research on topics in information and optimization sciences.

Issues up to 2022 co-published with and available at:Taylor & Francis
submissions@tarupublications.com
Open Access Research Article

Ensemble learning based predictive modelling on a highly imbalanced multiclass data

* ,

* Corresponding author · click or hover a name for details

pp. 2141–2164Vol. 45Issue 8November 2024DOI: 10.47974/JIOS-1778XML
Received:
10 Jan 2024
Published Online:
18 Dec 2024
Article type:
Research Article
Language:
EN
Article no.:
JIOS-1778
Pages:
2141–2164

Abstract

Class imbalance in the real-world datasets is a big challenge and the domains such as fraud detection, calamity occurrences, bankruptcy prediction etc. are prone to class imbalance due to the nature of occurrences of the events. In this paper, the detailed research using six ensemble machine learning techniques is applied to the undersampled, oversampled and the original dataset and the results are compared. The results of the research study indicates that amongst the applied six ensemble learners, the best learner is Random Forest algorithm (with entropy gain) implemented using ten-fold cross validation on the SMOTE oversampled dataset. 0.95 AUC and 0.8689 accuracy i.e. an increase of 4% in accuracy and substantial increase in other performance indicators is observed as compared to the remaining five ensemble classifiers.

Keywords

Subject Classifications

(2010) 86A15 (Seismology)68T05 (Learning and Adaptive Systems)62H30 (Classification and Discrimination)62H30 (Cluster Analysis)68T10 (Pattern recognition)

References

[1] D. Velusamy and K. Ramasamy, “Ensemble of heterogeneous classifiers for diagnosis and prediction of coronary artery disease with reduced feature subset”, Computer Methods and Programs in Biomedicine, vol. 198, (2021). doi: https://doi.org/10.1016/j.cmpb.2020.105770.
[2] D.A. Cieslak, N.V. Chawla, A. Striegel, “Combating Imbalance in Network Intrusion Datasets”, in Proceedings of the International Conference on Granular Computing, Atlanta, USA, 10-12, pp. 732-737, May (2006). doi: 10.1109/GRC.2006.1635905.
[3] F. Deng, J. Huang, X. Yuan, C. Cheng and L. Zheng, “Performance and efficiency of machine learning algorithms for analyzing rectangular biomedical data”, Laboratory Investigation, vol. 101, pp. 430–441 (2021), doi: https://doi.org/10.1038/s41374-020-00525-x.
[4] F. Kamalov, S. Moussa and J. A. Reyes, “KDE-Based Ensemble Learning for Imbalanced Data”, Electronics 2022, vol. 11, Issue 17 (2022), doi: https://doi.org/10.3390/electronics11172703.
[5] F.S. Hernández, J.C. Ballesteros‐Herráez, M.S. Kraiem, M. Sánchez‐Barba, M.N. Moreno‐García, “Predictive Modeling of ICU Healthcare‐Associated Infections from Imbalanced Data. Using Ensembles and a Clustering‐Based Undersampling Approach”, Applied Sciences, vol. 9 (2019), doi:10.3390/app9245287.
[6] J.A. Sáez, J. Luengo, J. Stefanowski and F.Herrera, “SMOTE‐IPF: Addressing the noisy and borderline examples problem in imbalanced classification by a resampling method with filtering”, Information Sciences, vol. 291, pp. 184–203 (2015), doi:https://doi.org/10.1016/j.ins.2014.08.051.
[7] J.M. Johnson and T.M. Khoshgoftaar, “Data-Centric AI for Healthcare Fraud Detection”, SN Computer Science, vol. 4, no. 4, (2023), doi: 10.1007/s42979-023-01809-x.
[8] J.M. Johnson and T.M. Khoshgoftaar, “Deep Learning and Thresholding with Class-Imbalanced Big Data”, in Proceedings of the 18th IEEE International Conference on Machine Learning and Applications (ICMLA), pp. 755-762 (2019).
[9] J.M. Johnson and T.M. Khoshgoftaar, “Survey on deep learning with class imbalance”, Journal Big Data, vol. 6, no. 1, (2019), doi:https://doi.org/10.1186/s40537-019-0192-5.
[10] J.P. Zhang and I. Mani, “KNN Approach to Unbalanced Data Distributions: A Case Study Involving Information Extraction”, in Proceedings of the International Conference on Machine Learning (ICML 2003), Workshop on Learning from Imbalanced Data Sets, Washington, DC, USA, 21 August (2003).
[11] K. Taehoon and A. Hyunchul, “A hybrid under-sampling approach for better bankruptcy prediction”, Journal of Intelligent Information Systems, vol. 21, no. 2, pp.173–190 (2015), doi: http://dx.doi.org/10.13088/jiis.2015.21.2.173.
[12] L. Liu, X. Wu and S. Li, “Solving the class imbalance problem using ensemble algorithm: application of screening for aortic dissection”, BMC Medical Informatics and Decision Making, vol. 82. https://doi.org/10.1186/s12911-022-01821-w
[13] M. H. A. Hamid, M. Yusoff and A. H. Mohamed, “Survey on Highly Imbalanced Multi-class Data”, International Journal of Advanced Computer Science and Applications (2022), doi:10.14569/ijacsa.2022.0130627.
[14] N. Jain and D. Virmani, “Ensemble learning using fast rule based fuzzy K-means pre clustering and classification for aquatic behavior-extracted tsunami prediction,” Journal of Information and Optimization Sciences, vol. 40, no. 2, pp. 441-453 (2019). doi: 10.1080/02522667.2019.1580884.
[15] P. Chujai, K. Chomboon, P. Teerarassamee, N. Kerdprasop and K. Kerdprasop, “Ensemble Learning For Imbalanced Data Classification Problem”, in proceedings of International Conference on Industrial Application Engineering, January (2015), DOI: 10.12792/iciae2015.079.
[16] P. Jain, M. S. Bajpai and R. Pamula, “A Modified DBSCAN Algorithm for Anomaly Detection in Time-series Data with Seasonality”, The International Arab Journal of Information Technology (IAJIT), vol. 19, no. 1, pp. 23 -28, January (2022), doi: 10.34028/iajit/19/1/3.
[17] P. Kang, S. Cho and D.L. MacLachlan, “Improved response modeling based on clustering, under‐sampling, and ensemble”, Expert Systems Applications, vol. 39, pp.6738–6753 (2012), doi:10.1016/j.eswa.2011.12.028.
[18] P. Kumar, R. Bhatnagar, K. Gaur, and A. Bhatnagar, “Classification of Imbalanced Data: Review of Methods and Applications”, IOP Conf. Series: Materials Science and Engineering, vol. 1099, no.1, March (2021), IOP Publishing, doi:10.1088/1757-899X/1099/1/012077.
[19] P. Shi and Z. Wang, “An Ensemble Tree Classifier for Highly Imbalanced Data Classification”, Journal of Systems Science and Complexity, vol. 34, no. 6, pp.2250–2266, 2021. doi:https://doi.org/10.1007/s11424-021-1038-8
[20] R. Kulkarni, S. Revathy, and S. Patil, “A Novel Approach to Maximize G-mean in Nonstationary Data with Recurrent Imbalance Shifts”, The International Arab Journal of Information Technology (IAJIT), vol. 18, no. 1, pp. 103 - 113, January (2021), doi: 10.34028/iajit/18/1/12.
[21] R.A. Bauder and T.M.Khoshgoftaar, “The Effects of Varying Class Distribution on Learner Behavior for Medicare Fraud Detection with Imbalanced Big Data”, Health Information Science and Systems, vol. 6, no. 1, (2018), doi: 10.1007/s13755-018-0051-3.
[22] R.A. Bauder, T.M. Khoshgoftaar and T. Hasanin, “An Empirical Study on Class Rarity in Big Data”, in Proceedings of the 17th IEEE international conference on machine learning and applications (ICMLA),  pp. 785–790 (2018), https ://doi.org/10.1109/ICMLA.2018.00125.
[23] S. M. Mousavi, Y. Sheng, W. Zhu and G. C. Beroza, “STanford EArthquake Dataset (STEAD): A Global Data Set of Seismic Signals for AI,” in IEEE Access, vol. 7, pp. 179464-179476 (2019), doi: 10.1109/ACCESS.2019.2947848.
[24] S. Wang, Y. Dai, J. Shen, L. Zhang, Y. Wang, and C. Zhang, “Research on expansion and classification of imbalanced data based on SMOTE algorithm,” Sci. Rep., vol. 11, no. 24039 (2021), doi: 10.1038/s41598-021-03430-5.
[25] S. Yen and Y. Lee, “Cluster-based under-sampling approaches for imbalanced data distributions”, Expert Systems with Applications, vol. 36, no. 3, Part 1, pp.5718-5727 (2009), ISSN 0957-4174, https://doi.org/10.1016/j.eswa.2008.06.108.
[26] Stanford Earthquake Dataset (STEAD)- A Global Data Set of Seismic Signals for AI, IEEE Access, September 2021.[online]. Available: https://www.kaggle.com/datasets/isevilla/stanford-earthquake-dataset-stead/data.
[27] T. Maciejewski and J. Stefanowski, “Local neighbourhood extension of SMOTE for mining imbalanced data,” in Proceedings of the IEEE Symposium on Computational Intelligence and Data Mining, Paris, France, pp. 104–111 (2011), doi: 10.1109/CIDM.2011.5949434.
[28] U. R. Salunkhe and S. N. Mali, “Classifier Ensemble Design for Imbalanced Data Classification: A Hybrid Approach”, Procedia Computer Science, vol. 85, pp. 725-732 (2016), ISSN 1877-0509,doi: https://doi.org/10.1016/j.procs.2016.05.259.
[29] V. J. Kadam and S. M. Jadhav, “Performance analysis of hyperparameter optimization methods for ensemble learning with small and medium sized medical datasets,” Journal of Discrete Mathematical Sciences and Cryptography, vol. 23, no. 1, pp. 115-123 (2020). doi: 10.1080/09720529.2020.1721871.
[30] V. Rupapara, F. Rustam, W. Aljedaani, H. F. Shahzad, and E. Lee, “Blood cancer prediction using leukemia microarray gene data and hybrid logistic vector trees model,” Scientific Reports, vol. 12, no. 1000 (2022), doi: https://doi.org/10.1038/s41598-022-04835-6.
[31] W. Wei, J. Li, L. Cao, Y. Ou and J. Chen, “Effective detection of sophisticated online banking fraud on extremely imbalanced data”, World Wide Web, vol. 16, no. 4, pp.449–75 (2013), doi:https://doi.org/10.1007/s11280-012-0178-0.
[32] W. Yang, C. Pan, and Y. Zhang, “An oversampling method for imbalanced data based on spatial distribution of minority samples SD-KMSMOTE,” Scientific Reports, vol. 12, p. 16820 (2022), doi: https://doi.org/10.1038/s41598-022-21046-1.
[33] X. Guo, Y. Yin, C. Dong, G. Yang and G. Zhou, “On the Class Imbalance Problem”, presented at the Fourth International Conference on Natural Computation, vol. 4, pp.192-201, 18-20 Oct. (2008), ISSN 978-0-7695-3304-9/08, DOI 10.1109/ICNC.2008.871.
[34] Y. Mi, “Imbalanced classification based on active learning SMOTE”, Research Journal of Applied Sciences, Engineering and Technology, vol. 5, no. 3, pp. 944–949 (2013).

Views: 181Downloads: 17Citations: 0