Boosting inspired overlap reduction : A novel approach for reducing overlap in imbalanced classification problems
*Pooja TyagiCorresponding authorpooja.tyagi94@gmail.comUniversity School of Information, Communication and TechnologyGuru Gobind Singh Indraprastha UniversityDwarka, New Delhi, 110078, IndiaView full profile → , Jaspreeti Singhjaspreeti_singh@ipu.ac.inUniversity School of Information, Communication and TechnologyGuru Gobind Singh Indraprastha UniversityDwarka, New Delhi, 110078, IndiaView full profile → , Anjana GosainAnjana_gosain@ipu.ac.inUniversity School of Information, Communication and TechnologyGuru Gobind Singh Indraprastha UniversityDwarka, New Delhi, 110078, IndiaView full profile →
* Corresponding author · click or hover a name for details
- Received:
- 01 Mar 2025
- Published Online:
- 09 May 2026
- Article type:
- Research Article
- Language:
- EN
- Article no.:
- JIOS-2141
- Pages:
- 2319–2343
Abstract
Keywords
Subject Classifications
References
[1] H. Guo, Y. Li, J. Shang, M. Gu, Y. Huang, and B. Gong, “Learning from class-imbalanced data: Review of methods and applications,” Expert Syst. Appl., vol. 73, pp. 220–239 (2017).
[2] V. García, R. Alejo, J. S. Sánchez, J. M. Sotoca, and R. A. Mollineda, “Combined effects of class imbalance and class overlap on instance-based classification,” in Proc. 7th Int. Conf. Intell. Data Eng. Autom. Learn. (IDEAL), Burgos, Spain, pp. 371–378 (2006).
[3] A. A. Khan, O. Chaudhari, and R. Chandra, “A review of ensemble learning and data augmentation models for class imbalanced problems: Combination, implementation and evaluation,” Expert Syst. Appl., vol. 244, p. 122778 (2024).
[4] Z. Xu, D. Shen, T. Nie, and Y. Kou, “A hybrid sampling algorithm combining M-SMOTE and ENN based on random forest for medical imbalanced data,” J. Biomed. Inform., vol. 107, p. 103465 (2020).
[5] P. Tyagi, J. Singh, and A. Gosain, “Handling class imbalance problem using feature selection techniques: A review,” in Proc. Int. Conf. Innovations Comput. Intell. Comput. Vis., Singapore (2022).
[6] K. Matloob, M. A. Khan, S. Abbas, M. T. Khan, and S. Khan, “A comparative performance analysis of data resampling methods on imbalance medical data,” IEEE Access, vol. 9, pp. 109960–109975 (2021).
[7] A. More, “Survey of resampling techniques for improving classification performance in unbalanced datasets,” arXiv preprint, arXiv:1608.06048 (2016).
[8] R. Ghorbani and R. Ghousi, “Comparing different resampling methods in predicting students’ performance using machine learning techniques,” IEEE Access, vol. 8, pp. 67899–67911 (2020).
[9] F. Feng, Y. Xu, Q. Li, Z. Huang, and X. Zheng, “Using cost-sensitive learning and feature selection algorithms to improve the performance of imbalanced classification,” IEEE Access, vol. 8, pp. 69979–69996 (2020).
[10] B. Pes, “Learning from high-dimensional biomedical datasets: The issue of class imbalance,” IEEE Access, vol. 8, pp. 13527–13540 (2020).
[11] M. Bach and A. Werner, “Cost-sensitive feature selection for class imbalance problem,” in Information Systems Architecture and Technology, Springer (2018).
[12] K. M. Hasib, M. M. Rahman, M. S. Hossain, and A. K. M. Masum, “A survey of methods for managing the classification and solution of data imbalance problem,” arXiv preprint, arXiv:2012.11870 (2020).
[13] M. S. Santos, P. H. Abreu, N. Japkowicz, A. Fernández, C. Soares, S. Wilk, and J. Santos, “On the joint-effect of class imbalance and overlap: A critical review,” Artif. Intell. Rev., vol. 55, no. 8, pp. 6207–6275 (2022).
[14] M. S. Santos, P. H. Abreu, N. Japkowicz, A. Fernández, and J. Santos, “A unifying view of class overlap and imbalance: Key concepts, multi-view panorama, and open avenues for research,” Inf. Fusion, vol. 89, pp. 228–253 (2023).
[15] A. Kumar, D. Singh, and R. S. Yadav, “Class overlap handling methods in imbalanced domain: A comprehensive survey,” Multimedia Tools Appl., vol. 83, no. 23, pp. 63243–63290 (2024).
[16] S. Wojciechowski and S. Wilk, “Difficulty factors and preprocessing in imbalanced data sets: An experimental study on artificial data,” Found. Comput. Decis. Sci., vol. 42, no. 2, pp. 149–176 (2017).
[17] P. Vuttipittayamongkol, E. Elyan, and A. Petrovski, “On the class overlap problem in imbalanced data classification,” Knowl.-Based Syst., vol. 212, p. 106631 (2021).
[18] G. H. Fu, Y. J. Wu, M. J. Zong, and L. Z. Yi, “Feature selection and classification by minimizing overlap degree for class-imbalanced data in metabolomics,” Chemom. Intell. Lab. Syst., vol. 196, p. 103906 (2020).
[19] A. Kumar, D. Singh, and R. S. Yadav, “Entropy and improved k-nearest neighbor search-based undersampling method to handle class overlap in imbalanced datasets,” Concurrency Comput.: Pract. Exp., vol. 36, no. 2, p. e7894 (2024).
[20] L. Gong, H. Zhang, J. Zhang, M. Wei, and Z. Huang, “A comprehensive investigation of the impact of class overlap on software defect prediction,” IEEE Trans. Softw. Eng., vol. 49, no. 4, pp. 2440–2458 (2022).
[21] C. Liu, Y. Jin, H. Chen, and X. Zhang, “Constrained oversampling: An oversampling approach to reduce noise generation in imbalanced datasets with class overlapping,” IEEE Access, vol. 10, pp. 91452–91465 (2022).
[22] H. K. Lee and S. B. Kim, “An overlap-sensitive margin classifier for imbalanced and overlapping data,” Expert Syst. Appl., vol. 98, pp. 72–83 (2018).
[23] D. Devi, S. K. Biswas, and B. Purkayastha, “Learning in presence of class imbalance and class overlapping using one-class SVM and undersampling technique,” Connection Sci., vol. 31, no. 2, pp. 105–142 (2019).
[24] F. Grina, Z. Elouedi, and E. Lefevre, “Evidential undersampling approach for imbalanced datasets with class-overlapping and noise,” in Proc. 18th Int. Conf. Modeling Decisions Artif. Intell. (MDAI), pp. 181–192 (2021).
[25] E. Alogogianni and M. Virvou, “Handling class imbalance and class overlap in machine learning applications for undeclared work prediction,” Electronics, vol. 12, no. 4, p. 913 (2023).
[26] P. Vuttipittayamongkol, E. Elyan, A. Petrovski, and C. Jayne, “Overlap-based undersampling for improving imbalanced data classification,” in Proc. IDEAL, Madrid, Spain, pp. 689–697 (2018).
[27] P. Vuttipittayamongkol and E. Elyan, “Neighbourhood-based undersampling approach for handling imbalanced and overlapped data,” Inf. Sci., vol. 509, pp. 47–70 (2020).
[28] R. Zhang, Z. Zhang, and D. Wang, “RFCL: A new under-sampling method of reducing the degree of imbalance and overlap,” Pattern Anal. Appl., vol. 24, pp. 641–654 (2021).
[29] W. Wang and D. Sun, “The improved AdaBoost algorithms for imbalanced data classification,” Inf. Sci., vol. 563, pp. 358–374 (2021).
[30] R. Patil and K. Vanjerkhede, “Random forest machine learning approach for efficient optimization of Sierpinski carpet fractal antenna,” J. Inf. Optim. Sci., vol. 45, no. 4, pp. 851–861 (2024).
[31] A. Upadhyay, M. Singh, and V. K. Yadav, “Improvised number identification using SVM and random forest classifiers,” J. Inf. Optim. Sci., vol. 41, no. 2, pp. 387–394 (2020).
[32] V. García, R. A. Mollineda, and J. S. Sánchez, “On the k-NN performance in a challenging scenario of imbalance and overlapping,” Pattern Anal. Appl., vol. 11, pp. 269–280 (2008).
[33] Z. Borsos, C. Lemnaru, and R. Potolea, “Dealing with overlap and imbalance: A new metric and approach,” Pattern Anal. Appl., vol. 21, pp. 381–395 (2018).
[34] J. A. Sáez, M. Galar, and B. Krawczyk, “Addressing the overlapping data problem in classification using the one-vs-one decomposition strategy,” IEEE Access, vol. 7, pp. 83396–83411 (2019).
[35] A. Guzmán-Ponce, R. M. Valdovinos, J. S. Sánchez, and J. R. Marcial-Romero, “A new under-sampling method to face class overlap and imbalance,” Appl. Sci., vol. 10, no. 15, p. 5164 (2020).
[36] P. Vuttipittayamongkol and E. Elyan, “Improved overlap-based undersampling for imbalanced dataset classification with application to epilepsy and Parkinson’s disease,” Int. J. Neural Syst., vol. 30, no. 8, p. 2050043 (2020).
[37] M. M. Nwe and K. T. Lynn, “KNN-based overlapping samples filter approach for classification of imbalanced data,” in Software Engineering Research, Management and Applications, pp. 55–73 (2020).
[38] B. W. Yuan, Z. L. Zhang, X. G. Luo, Y. Yu, X. H. Zou, and X. D. Zou, “OIS-RF: A novel overlap and imbalance-sensitive random forest,” Eng. Appl. Artif. Intell., vol. 104, p. 104355 (2021).
[39] Y. Zhang and F. Han, “An improved ensemble classification algorithm for imbalanced data with sample overlap,” in Proc. Int. Conf. Neural Comput. Adv. Appl., pp. 454–468 (2022).
[40] M. A. Islam, M. A. Uddin, S. Aryal, and G. Stea, “An ensemble learning approach for anomaly detection in credit card data with imbalanced and overlapped classes,” J. Inf. Secur. Appl., vol. 78, p. 103618 (2023).
[41] S. Jubair, J. Yang, and B. Ali, “Overlap to equilibrium: Oversampling imbalanced datasets using overlapping degree,” Inf. Process. Manage., vol. 62, no. 2, p. 103975 (2025).
[42] A. Natekin and A. Knoll, “Gradient boosting machines, a tutorial,” Front. Neurorobot., vol. 7, p. 21 (2013).
[43] S. J. Yen and Y. S. Lee, “Cluster-based under-sampling approaches for imbalanced data distributions,” Expert Syst. Appl., vol. 36, no. 3, pp. 5718–5727 (2009).
[44] M. A. U. H. Tahir, S. Asghar, A. Manzoor, and M. A. Noor, “A classification model for class imbalance dataset using genetic programming,” IEEE Access, vol. 7, pp. 71013–71037 (2019).
[45] G. E. Batista, R. C. Prati, and M. C. Monard, “A study of the behavior of several methods for balancing machine learning training data,” ACM SIGKDD Explor. Newsl., vol. 6, no. 1, pp. 20–29 (2004).
[46] P. Tyagi, J. Singh, and A. Gosain, “Comprehensive empirical investigation for prioritizing the pipeline of using feature selection and data resampling techniques,” J. Intell. Fuzzy Syst., vol. 46, no. 3, pp. 6019–6040 (2024).
[47] X. Tao, Q. Li, H. Zhang, and J. Li, “Real-value negative selection oversampling for imbalanced dataset learning,” Expert Syst. Appl., vol. 129, pp. 118–134 (2019).
[48] M. Kubat and S. Matwin, “Addressing the curse of imbalanced training sets: One-sided selection,” in Proc. Int. Conf. Mach. Learn. (ICML), Nashville, TN, USA, p. 179 (1997).
[49] S. Fotouhi, S. Asadi, and M. W. Kattan, “A comprehensive data-level analysis for cancer diagnosis on imbalanced data,” J. Biomed. Inform., vol. 90, p. 103089 (2019).
[50] V. H. Barella, L. P. Garcia, M. C. de Souto, A. C. Lorena, and A. C. de Carvalho, “Assessing the data complexity of imbalanced datasets,” Inf. Sci., vol. 553, pp. 83–109 (2021).
[51] J. Laurikkala, “Improving identification of difficult small classes by balancing class distribution,” in Proc. 8th Conf. Artif. Intell. Med. Eur. (AIME), Cascais, Portugal, pp. 63–66 (2001).
[52] I. Mani and I. Zhang, “kNN approach to unbalanced data distributions: A case study involving information extraction,” in Proc. Workshop Learn. Imbalanced Datasets, Washington, DC, USA, pp. 1–7 (2003).
[53] J. Demšar, “Statistical comparisons of classifiers over multiple data sets,” J. Mach. Learn. Res., vol. 7, pp. 1–30 (2006).
[54] J. Derrac, S. García, D. Molina, and F. Herrera, “A practical tutorial on the use of non-parametric statistical tests as a methodology for comparing evolutionary and swarm intelligence algorithms,” Swarm Evol. Comput., vol. 1, no. 1, pp. 3–18 (2011).




