Ensemble learning based predictive modelling on a highly imbalanced multiclass data
*Manka VastiCorresponding authormankavasti@gmail.comAffiliation 1University School of Information, Communication and TechnologyGuru Gobind Singh Indraprastha UniversityNew Delhi, 110078, IndiaAffiliation 2Department of Computer Science and EngineeringSchool of Engineering and SciencesGD Goenka UniversityGurugram, Haryana, 122103, IndiaView full profile → , Amita Devamita_dev@hotmail.comDirectorate of Training and Technical EducationDelhi, 110034, IndiaView full profile →
* Corresponding author · click or hover a name for details
- Received:
- 10 Jan 2024
- Published Online:
- 18 Dec 2024
- Article type:
- Research Article
- Language:
- EN
- Article no.:
- JIOS-1778
- Pages:
- 2141–2164
Abstract
Keywords
Subject Classifications
References
[1] D. Velusamy and K. Ramasamy, “Ensemble of heterogeneous classifiers for diagnosis and prediction of coronary artery disease with reduced feature subset”, Computer Methods and Programs in Biomedicine, vol. 198, (2021). doi: https://doi.org/10.1016/j.cmpb.2020.105770.
[2] D.A. Cieslak, N.V. Chawla, A. Striegel, “Combating Imbalance in Network Intrusion Datasets”, in Proceedings of the International Conference on Granular Computing, Atlanta, USA, 10-12, pp. 732-737, May (2006). doi: 10.1109/GRC.2006.1635905.
[3] F. Deng, J. Huang, X. Yuan, C. Cheng and L. Zheng, “Performance and efficiency of machine learning algorithms for analyzing rectangular biomedical data”, Laboratory Investigation, vol. 101, pp. 430–441 (2021), doi: https://doi.org/10.1038/s41374-020-00525-x.
[4] F. Kamalov, S. Moussa and J. A. Reyes, “KDE-Based Ensemble Learning for Imbalanced Data”, Electronics 2022, vol. 11, Issue 17 (2022), doi: https://doi.org/10.3390/electronics11172703.
[5] F.S. Hernández, J.C. Ballesteros‐Herráez, M.S. Kraiem, M. Sánchez‐Barba, M.N. Moreno‐García, “Predictive Modeling of ICU Healthcare‐Associated Infections from Imbalanced Data. Using Ensembles and a Clustering‐Based Undersampling Approach”, Applied Sciences, vol. 9 (2019), doi:10.3390/app9245287.
[6] J.A. Sáez, J. Luengo, J. Stefanowski and F.Herrera, “SMOTE‐IPF: Addressing the noisy and borderline examples problem in imbalanced classification by a resampling method with filtering”, Information Sciences, vol. 291, pp. 184–203 (2015), doi:https://doi.org/10.1016/j.ins.2014.08.051.
[7] J.M. Johnson and T.M. Khoshgoftaar, “Data-Centric AI for Healthcare Fraud Detection”, SN Computer Science, vol. 4, no. 4, (2023), doi: 10.1007/s42979-023-01809-x.
[8] J.M. Johnson and T.M. Khoshgoftaar, “Deep Learning and Thresholding with Class-Imbalanced Big Data”, in Proceedings of the 18th IEEE International Conference on Machine Learning and Applications (ICMLA), pp. 755-762 (2019).
[9] J.M. Johnson and T.M. Khoshgoftaar, “Survey on deep learning with class imbalance”, Journal Big Data, vol. 6, no. 1, (2019), doi:https://doi.org/10.1186/s40537-019-0192-5.
[10] J.P. Zhang and I. Mani, “KNN Approach to Unbalanced Data Distributions: A Case Study Involving Information Extraction”, in Proceedings of the International Conference on Machine Learning (ICML 2003), Workshop on Learning from Imbalanced Data Sets, Washington, DC, USA, 21 August (2003).
[11] K. Taehoon and A. Hyunchul, “A hybrid under-sampling approach for better bankruptcy prediction”, Journal of Intelligent Information Systems, vol. 21, no. 2, pp.173–190 (2015), doi: http://dx.doi.org/10.13088/jiis.2015.21.2.173.
[12] L. Liu, X. Wu and S. Li, “Solving the class imbalance problem using ensemble algorithm: application of screening for aortic dissection”, BMC Medical Informatics and Decision Making, vol. 82. https://doi.org/10.1186/s12911-022-01821-w
[13] M. H. A. Hamid, M. Yusoff and A. H. Mohamed, “Survey on Highly Imbalanced Multi-class Data”, International Journal of Advanced Computer Science and Applications (2022), doi:10.14569/ijacsa.2022.0130627.
[14] N. Jain and D. Virmani, “Ensemble learning using fast rule based fuzzy K-means pre clustering and classification for aquatic behavior-extracted tsunami prediction,” Journal of Information and Optimization Sciences, vol. 40, no. 2, pp. 441-453 (2019). doi: 10.1080/02522667.2019.1580884.
[15] P. Chujai, K. Chomboon, P. Teerarassamee, N. Kerdprasop and K. Kerdprasop, “Ensemble Learning For Imbalanced Data Classification Problem”, in proceedings of International Conference on Industrial Application Engineering, January (2015), DOI: 10.12792/iciae2015.079.
[16] P. Jain, M. S. Bajpai and R. Pamula, “A Modified DBSCAN Algorithm for Anomaly Detection in Time-series Data with Seasonality”, The International Arab Journal of Information Technology (IAJIT), vol. 19, no. 1, pp. 23 -28, January (2022), doi: 10.34028/iajit/19/1/3.
[17] P. Kang, S. Cho and D.L. MacLachlan, “Improved response modeling based on clustering, under‐sampling, and ensemble”, Expert Systems Applications, vol. 39, pp.6738–6753 (2012), doi:10.1016/j.eswa.2011.12.028.
[18] P. Kumar, R. Bhatnagar, K. Gaur, and A. Bhatnagar, “Classification of Imbalanced Data: Review of Methods and Applications”, IOP Conf. Series: Materials Science and Engineering, vol. 1099, no.1, March (2021), IOP Publishing, doi:10.1088/1757-899X/1099/1/012077.
[19] P. Shi and Z. Wang, “An Ensemble Tree Classifier for Highly Imbalanced Data Classification”, Journal of Systems Science and Complexity, vol. 34, no. 6, pp.2250–2266, 2021. doi:https://doi.org/10.1007/s11424-021-1038-8
[20] R. Kulkarni, S. Revathy, and S. Patil, “A Novel Approach to Maximize G-mean in Nonstationary Data with Recurrent Imbalance Shifts”, The International Arab Journal of Information Technology (IAJIT), vol. 18, no. 1, pp. 103 - 113, January (2021), doi: 10.34028/iajit/18/1/12.
[21] R.A. Bauder and T.M.Khoshgoftaar, “The Effects of Varying Class Distribution on Learner Behavior for Medicare Fraud Detection with Imbalanced Big Data”, Health Information Science and Systems, vol. 6, no. 1, (2018), doi: 10.1007/s13755-018-0051-3.
[22] R.A. Bauder, T.M. Khoshgoftaar and T. Hasanin, “An Empirical Study on Class Rarity in Big Data”, in Proceedings of the 17th IEEE international conference on machine learning and applications (ICMLA), pp. 785–790 (2018), https ://doi.org/10.1109/ICMLA.2018.00125.
[23] S. M. Mousavi, Y. Sheng, W. Zhu and G. C. Beroza, “STanford EArthquake Dataset (STEAD): A Global Data Set of Seismic Signals for AI,” in IEEE Access, vol. 7, pp. 179464-179476 (2019), doi: 10.1109/ACCESS.2019.2947848.
[24] S. Wang, Y. Dai, J. Shen, L. Zhang, Y. Wang, and C. Zhang, “Research on expansion and classification of imbalanced data based on SMOTE algorithm,” Sci. Rep., vol. 11, no. 24039 (2021), doi: 10.1038/s41598-021-03430-5.
[25] S. Yen and Y. Lee, “Cluster-based under-sampling approaches for imbalanced data distributions”, Expert Systems with Applications, vol. 36, no. 3, Part 1, pp.5718-5727 (2009), ISSN 0957-4174, https://doi.org/10.1016/j.eswa.2008.06.108.
[26] Stanford Earthquake Dataset (STEAD)- A Global Data Set of Seismic Signals for AI, IEEE Access, September 2021.[online]. Available: https://www.kaggle.com/datasets/isevilla/stanford-earthquake-dataset-stead/data.
[27] T. Maciejewski and J. Stefanowski, “Local neighbourhood extension of SMOTE for mining imbalanced data,” in Proceedings of the IEEE Symposium on Computational Intelligence and Data Mining, Paris, France, pp. 104–111 (2011), doi: 10.1109/CIDM.2011.5949434.
[28] U. R. Salunkhe and S. N. Mali, “Classifier Ensemble Design for Imbalanced Data Classification: A Hybrid Approach”, Procedia Computer Science, vol. 85, pp. 725-732 (2016), ISSN 1877-0509,doi: https://doi.org/10.1016/j.procs.2016.05.259.
[29] V. J. Kadam and S. M. Jadhav, “Performance analysis of hyperparameter optimization methods for ensemble learning with small and medium sized medical datasets,” Journal of Discrete Mathematical Sciences and Cryptography, vol. 23, no. 1, pp. 115-123 (2020). doi: 10.1080/09720529.2020.1721871.
[30] V. Rupapara, F. Rustam, W. Aljedaani, H. F. Shahzad, and E. Lee, “Blood cancer prediction using leukemia microarray gene data and hybrid logistic vector trees model,” Scientific Reports, vol. 12, no. 1000 (2022), doi: https://doi.org/10.1038/s41598-022-04835-6.
[31] W. Wei, J. Li, L. Cao, Y. Ou and J. Chen, “Effective detection of sophisticated online banking fraud on extremely imbalanced data”, World Wide Web, vol. 16, no. 4, pp.449–75 (2013), doi:https://doi.org/10.1007/s11280-012-0178-0.
[32] W. Yang, C. Pan, and Y. Zhang, “An oversampling method for imbalanced data based on spatial distribution of minority samples SD-KMSMOTE,” Scientific Reports, vol. 12, p. 16820 (2022), doi: https://doi.org/10.1038/s41598-022-21046-1.
[33] X. Guo, Y. Yin, C. Dong, G. Yang and G. Zhou, “On the Class Imbalance Problem”, presented at the Fourth International Conference on Natural Computation, vol. 4, pp.192-201, 18-20 Oct. (2008), ISSN 978-0-7695-3304-9/08, DOI 10.1109/ICNC.2008.871.
[34] Y. Mi, “Imbalanced classification based on active learning SMOTE”, Research Journal of Applied Sciences, Engineering and Technology, vol. 5, no. 3, pp. 944–949 (2013).




