TARU PUBLICATIONS
Journal of Information and Optimization Sciences cover
Hybrid ·Peer-reviewed·ISSN (Online): 2169-0103·ISSN (Print): 0252-2667

WoS  JIF 2026 : 0.4 (Q4)

Powered by:Powered by

Monthly Journal: Publishes theoretical and applied research on topics in information and optimization sciences.

Issues up to 2022 co-published with and available at:Taylor & Francis
submissions@tarupublications.com
Open Access Research Article

Boosting inspired overlap reduction : A novel approach for reducing overlap in imbalanced classification problems

* , ,

* Corresponding author · click or hover a name for details

pp. 2319–2343Vol. 47Issue 6June 2026DOI: 10.47974/JIOS-2141XML
Received:
01 Mar 2025
Published Online:
09 May 2026
Article type:
Research Article
Language:
EN
Article no.:
JIOS-2141
Pages:
2319–2343

Abstract

The classification of imbalanced datasets has garnered a significant research attention over the past decades. Although, numerous solutions have been proposed to address imbalanced classification problems, recent research suggests that class imbalance alone does not significantly degrade learning performance. Rather, class overlap, which has often been overlooked in previous studies has an even greater negative impact on the performance of learning algorithms. This paper proposes a Boosting-Inspired Overlap Reduction (BiOR) algorithm to handle the problem of class overlap in imbalanced datasets. BiOR employs a variance-based adaptive threshold mechanism, where the overlap threshold is dynamically determined as a function of the variance of margins in both majority and minority classes. By iteratively adjusting weights of samples, BiOR  ensures selective removal of highly overlapping majority instances, thereby enhancing class separability without excessive data loss. The effectiveness of BiOR is evaluated on 20 benchmark real world datasets as well as 16 synthetic datasets, demonstrating its superior classification performance over existing resampling techniques.

Keywords

Subject Classifications

68T01

References

[1] H. Guo, Y. Li, J. Shang, M. Gu, Y. Huang, and B. Gong, “Learning from class-imbalanced data: Review of methods and applications,” Expert Syst. Appl., vol. 73, pp. 220–239 (2017).
[2] V. García, R. Alejo, J. S. Sánchez, J. M. Sotoca, and R. A. Mollineda, “Combined effects of class imbalance and class overlap on instance-based classification,” in Proc. 7th Int. Conf. Intell. Data Eng. Autom. Learn. (IDEAL), Burgos, Spain, pp. 371–378 (2006).
[3] A. A. Khan, O. Chaudhari, and R. Chandra, “A review of ensemble learning and data augmentation models for class imbalanced problems: Combination, implementation and evaluation,” Expert Syst. Appl., vol. 244, p. 122778 (2024).
[4] Z. Xu, D. Shen, T. Nie, and Y. Kou, “A hybrid sampling algorithm combining M-SMOTE and ENN based on random forest for medical imbalanced data,” J. Biomed. Inform., vol. 107, p. 103465 (2020).
[5] P. Tyagi, J. Singh, and A. Gosain, “Handling class imbalance problem using feature selection techniques: A review,” in Proc. Int. Conf. Innovations Comput. Intell. Comput. Vis., Singapore (2022).
[6] K. Matloob, M. A. Khan, S. Abbas, M. T. Khan, and S. Khan, “A comparative performance analysis of data resampling methods on imbalance medical data,” IEEE Access, vol. 9, pp. 109960–109975 (2021).
[7] A. More, “Survey of resampling techniques for improving classification performance in unbalanced datasets,” arXiv preprint, arXiv:1608.06048 (2016).
[8] R. Ghorbani and R. Ghousi, “Comparing different resampling methods in predicting students’ performance using machine learning techniques,” IEEE Access, vol. 8, pp. 67899–67911 (2020).
[9] F. Feng, Y. Xu, Q. Li, Z. Huang, and X. Zheng, “Using cost-sensitive learning and feature selection algorithms to improve the performance of imbalanced classification,” IEEE Access, vol. 8, pp. 69979–69996 (2020).
[10] B. Pes, “Learning from high-dimensional biomedical datasets: The issue of class imbalance,” IEEE Access, vol. 8, pp. 13527–13540 (2020).
[11] M. Bach and A. Werner, “Cost-sensitive feature selection for class imbalance problem,” in Information Systems Architecture and Technology, Springer (2018).
[12] K. M. Hasib, M. M. Rahman, M. S. Hossain, and A. K. M. Masum, “A survey of methods for managing the classification and solution of data imbalance problem,” arXiv preprint, arXiv:2012.11870 (2020).
[13] M. S. Santos, P. H. Abreu, N. Japkowicz, A. Fernández, C. Soares, S. Wilk, and J. Santos, “On the joint-effect of class imbalance and overlap: A critical review,” Artif. Intell. Rev., vol. 55, no. 8, pp. 6207–6275 (2022).
[14] M. S. Santos, P. H. Abreu, N. Japkowicz, A. Fernández, and J. Santos, “A unifying view of class overlap and imbalance: Key concepts, multi-view panorama, and open avenues for research,” Inf. Fusion, vol. 89, pp. 228–253 (2023).
[15] A. Kumar, D. Singh, and R. S. Yadav, “Class overlap handling methods in imbalanced domain: A comprehensive survey,” Multimedia Tools Appl., vol. 83, no. 23, pp. 63243–63290 (2024).
[16] S. Wojciechowski and S. Wilk, “Difficulty factors and preprocessing in imbalanced data sets: An experimental study on artificial data,” Found. Comput. Decis. Sci., vol. 42, no. 2, pp. 149–176 (2017).
[17] P. Vuttipittayamongkol, E. Elyan, and A. Petrovski, “On the class overlap problem in imbalanced data classification,” Knowl.-Based Syst., vol. 212, p. 106631 (2021).
[18] G. H. Fu, Y. J. Wu, M. J. Zong, and L. Z. Yi, “Feature selection and classification by minimizing overlap degree for class-imbalanced data in metabolomics,” Chemom. Intell. Lab. Syst., vol. 196, p. 103906 (2020).
[19] A. Kumar, D. Singh, and R. S. Yadav, “Entropy and improved k-nearest neighbor search-based undersampling method to handle class overlap in imbalanced datasets,” Concurrency Comput.: Pract. Exp., vol. 36, no. 2, p. e7894 (2024).
[20] L. Gong, H. Zhang, J. Zhang, M. Wei, and Z. Huang, “A comprehensive investigation of the impact of class overlap on software defect prediction,” IEEE Trans. Softw. Eng., vol. 49, no. 4, pp. 2440–2458 (2022).
[21] C. Liu, Y. Jin, H. Chen, and X. Zhang, “Constrained oversampling: An oversampling approach to reduce noise generation in imbalanced datasets with class overlapping,” IEEE Access, vol. 10, pp. 91452–91465 (2022).
[22] H. K. Lee and S. B. Kim, “An overlap-sensitive margin classifier for imbalanced and overlapping data,” Expert Syst. Appl., vol. 98, pp. 72–83 (2018).
[23] D. Devi, S. K. Biswas, and B. Purkayastha, “Learning in presence of class imbalance and class overlapping using one-class SVM and undersampling technique,” Connection Sci., vol. 31, no. 2, pp. 105–142 (2019).
[24] F. Grina, Z. Elouedi, and E. Lefevre, “Evidential undersampling approach for imbalanced datasets with class-overlapping and noise,” in Proc. 18th Int. Conf. Modeling Decisions Artif. Intell. (MDAI), pp. 181–192 (2021).
[25] E. Alogogianni and M. Virvou, “Handling class imbalance and class overlap in machine learning applications for undeclared work prediction,” Electronics, vol. 12, no. 4, p. 913 (2023).
[26] P. Vuttipittayamongkol, E. Elyan, A. Petrovski, and C. Jayne, “Overlap-based undersampling for improving imbalanced data classification,” in Proc. IDEAL, Madrid, Spain, pp. 689–697 (2018).
[27] P. Vuttipittayamongkol and E. Elyan, “Neighbourhood-based undersampling approach for handling imbalanced and overlapped data,” Inf. Sci., vol. 509, pp. 47–70 (2020).
[28] R. Zhang, Z. Zhang, and D. Wang, “RFCL: A new under-sampling method of reducing the degree of imbalance and overlap,” Pattern Anal. Appl., vol. 24, pp. 641–654 (2021).
[29] W. Wang and D. Sun, “The improved AdaBoost algorithms for imbalanced data classification,” Inf. Sci., vol. 563, pp. 358–374 (2021).
[30] R. Patil and K. Vanjerkhede, “Random forest machine learning approach for efficient optimization of Sierpinski carpet fractal antenna,” J. Inf. Optim. Sci., vol. 45, no. 4, pp. 851–861 (2024).
[31] A. Upadhyay, M. Singh, and V. K. Yadav, “Improvised number identification using SVM and random forest classifiers,” J. Inf. Optim. Sci., vol. 41, no. 2, pp. 387–394 (2020).
[32] V. García, R. A. Mollineda, and J. S. Sánchez, “On the k-NN performance in a challenging scenario of imbalance and overlapping,” Pattern Anal. Appl., vol. 11, pp. 269–280 (2008).
[33] Z. Borsos, C. Lemnaru, and R. Potolea, “Dealing with overlap and imbalance: A new metric and approach,” Pattern Anal. Appl., vol. 21, pp. 381–395 (2018).
[34] J. A. Sáez, M. Galar, and B. Krawczyk, “Addressing the overlapping data problem in classification using the one-vs-one decomposition strategy,” IEEE Access, vol. 7, pp. 83396–83411 (2019).
[35] A. Guzmán-Ponce, R. M. Valdovinos, J. S. Sánchez, and J. R. Marcial-Romero, “A new under-sampling method to face class overlap and imbalance,” Appl. Sci., vol. 10, no. 15, p. 5164 (2020).
[36] P. Vuttipittayamongkol and E. Elyan, “Improved overlap-based undersampling for imbalanced dataset classification with application to epilepsy and Parkinson’s disease,” Int. J. Neural Syst., vol. 30, no. 8, p. 2050043 (2020).
[37] M. M. Nwe and K. T. Lynn, “KNN-based overlapping samples filter approach for classification of imbalanced data,” in Software Engineering Research, Management and Applications, pp. 55–73 (2020).
[38] B. W. Yuan, Z. L. Zhang, X. G. Luo, Y. Yu, X. H. Zou, and X. D. Zou, “OIS-RF: A novel overlap and imbalance-sensitive random forest,” Eng. Appl. Artif. Intell., vol. 104, p. 104355 (2021).
[39] Y. Zhang and F. Han, “An improved ensemble classification algorithm for imbalanced data with sample overlap,” in Proc. Int. Conf. Neural Comput. Adv. Appl., pp. 454–468 (2022).
[40] M. A. Islam, M. A. Uddin, S. Aryal, and G. Stea, “An ensemble learning approach for anomaly detection in credit card data with imbalanced and overlapped classes,” J. Inf. Secur. Appl., vol. 78, p. 103618 (2023).
[41] S. Jubair, J. Yang, and B. Ali, “Overlap to equilibrium: Oversampling imbalanced datasets using overlapping degree,” Inf. Process. Manage., vol. 62, no. 2, p. 103975 (2025).
[42] A. Natekin and A. Knoll, “Gradient boosting machines, a tutorial,” Front. Neurorobot., vol. 7, p. 21 (2013).
[43] S. J. Yen and Y. S. Lee, “Cluster-based under-sampling approaches for imbalanced data distributions,” Expert Syst. Appl., vol. 36, no. 3, pp. 5718–5727 (2009).
[44] M. A. U. H. Tahir, S. Asghar, A. Manzoor, and M. A. Noor, “A classification model for class imbalance dataset using genetic programming,” IEEE Access, vol. 7, pp. 71013–71037 (2019).
[45] G. E. Batista, R. C. Prati, and M. C. Monard, “A study of the behavior of several methods for balancing machine learning training data,” ACM SIGKDD Explor. Newsl., vol. 6, no. 1, pp. 20–29 (2004).
[46] P. Tyagi, J. Singh, and A. Gosain, “Comprehensive empirical investigation for prioritizing the pipeline of using feature selection and data resampling techniques,” J. Intell. Fuzzy Syst., vol. 46, no. 3, pp. 6019–6040 (2024).
[47] X. Tao, Q. Li, H. Zhang, and J. Li, “Real-value negative selection oversampling for imbalanced dataset learning,” Expert Syst. Appl., vol. 129, pp. 118–134 (2019).
[48] M. Kubat and S. Matwin, “Addressing the curse of imbalanced training sets: One-sided selection,” in Proc. Int. Conf. Mach. Learn. (ICML), Nashville, TN, USA, p. 179 (1997).
[49] S. Fotouhi, S. Asadi, and M. W. Kattan, “A comprehensive data-level analysis for cancer diagnosis on imbalanced data,” J. Biomed. Inform., vol. 90, p. 103089 (2019).
[50] V. H. Barella, L. P. Garcia, M. C. de Souto, A. C. Lorena, and A. C. de Carvalho, “Assessing the data complexity of imbalanced datasets,” Inf. Sci., vol. 553, pp. 83–109 (2021).
[51] J. Laurikkala, “Improving identification of difficult small classes by balancing class distribution,” in Proc. 8th Conf. Artif. Intell. Med. Eur. (AIME), Cascais, Portugal, pp. 63–66 (2001).
[52] I. Mani and I. Zhang, “kNN approach to unbalanced data distributions: A case study involving information extraction,” in Proc. Workshop Learn. Imbalanced Datasets, Washington, DC, USA, pp. 1–7 (2003).
[53] J. Demšar, “Statistical comparisons of classifiers over multiple data sets,” J. Mach. Learn. Res., vol. 7, pp. 1–30 (2006).
[54] J. Derrac, S. García, D. Molina, and F. Herrera, “A practical tutorial on the use of non-parametric statistical tests as a methodology for comparing evolutionary and swarm intelligence algorithms,” Swarm Evol. Comput., vol. 1, no. 1, pp. 3–18 (2011).

Views: 156Downloads: 96Citations: 0