TARU PUBLICATIONS
Journal of Information and Optimization Sciences cover
Open Access ·Peer-reviewed·ISSN (Online): 2169-0103·ISSN (Print): 0252-2667
Powered by:DOICrossrefiThenticate

The Journal of Information and Optimization Sciences (JIOS) is a world leading journal publishing high quality, rigorously peer-reviewed original research in all mathematically-oriented theoretical and applied topics in information sciences, optimization sciences and related areas since 1980. Subjects include but are not limited to: • Information Sciences • Optimization Sciences • Control Theory • Operational Research • Decision Sciences • Information Theory • Information Technology • Computer Networks and Communications • Mathematical Programming • Modelling and Simulation • Database Management • Applications to Engineering Sciences • Applications to Technology

Issues up to 2022 co-published with and available at:Taylor & Francis
submissions@tarupublications.com
Open Access Research Article

DOFCM-SMOTE : A technique to handle class imbalance problem

* ,

* Corresponding author · click or hover a name for details

pp. 1705–1723Vol. 46Issue 5July 2025DOI: 10.47974/JIOS-1946XML
Received:
07 Aug 2024
Published Online:
01 Jul 2025
Article type:
Research Article
Language:
EN
Article no.:
JIOS-1946
Pages:
1705–1723

Abstract

Class imbalance is a common issue in real-world datasets, where there are relatively few instances of the class of interest compared to others. The classifier’s performance is highly affected by the impurities built-in with the data like imbalanced data, noise, and class overlapping; therefore, an oversampling technique is proposed by fusing density-oriented fuzzy-c means clustering and synthetic minority oversampling technique(DOFCM-SMOTE). The density-oriented approach is intended to find outliers to address the performance issue. The proposed method works in 3 phases: 1) It detects and eliminates outliers in the dataset, 2) It Assigns membership degrees to the given samples by fuzzy c-means clustering, and 3) It then applies SMOTE to the minority cluster to balance the distribution of classes. To determine the performance, a decision tree classifier is employed as a learner. This study was conducted on ten publicly available data sets with varying imbalance ratios(IR) ranging from high to low. The findings of this study are centric on the process of clustering through which outliers are detected even before minority and majority class clusters are identified, and a general pre-processing technique, to balance the distribution of classes. The introduced approach is compared with four other oversampling techniques and the best balancing algorithm is determined. The performances are accessed through AUC and G-mean, and it was reported that the proposed technique significantly improved and outperformed the other over-sampling methods.

Keywords

Subject Classifications

68T05:Learning and adaptive systems in artificial intelligence

References

[1] J. Alcalá-Fdez, L. Sánchez, S. García, M. J. del Jesus, S. Ventura, J. M. Benítez, and F. Herrera, “KEEL: A software tool to assess evolutionary algorithms for data mining problems,” Soft Computing, vol. 13, no. 3, pp. 307–318 (2009).[2] J. Alcalá-Fdez, A. Fernández, J. Luengo, J. Derrac, S. García, L. Sánchez, and F. Herrera, “KEEL data-mining software tool: data set repository, integration of algorithms and experimental analysis framework,” Journal of Multiple-Valued Logic & Soft Computing, vol. 17, pp. 255–287 (2011).[3] R. Barandela, J. S. Sánchez, V. García, and E. Rangel, “Strategies for learning in class imbalance problems,” Pattern Recognition, vol. 36, no. 3, pp. 849–851 (2003).[4] S. Barua, M. M. Islam, K. Murase, and X. Yao, “A novel synthetic minority oversampling technique for imbalanced data set learning,” in Proc. Int. Conf. Neural Information Processing, Springer, pp. 735–744 (2011).[5] S. Barua, M. M. Islam, K. Murase, and X. Yao, “ProWSyn: Proximity weighted synthetic oversampling technique for imbalanced data set learning,” in Proc. Pacific-Asia Conf. Knowledge Discovery and Data Mining, Springer, pp. 317–328 (2013).[6] C. Bunkhumpornpat, K. Sinapiromsaran, and C. Lursinsap, “Safe-Level-SMOTE: Safe-level-synthetic minority over-sampling technique for handling the class imbalanced problem,” in Proc. Pacific-Asia Conf. Knowledge Discovery and Data Mining, Springer, pp. 475–482 (2009).[7] C. Bunkhumpornpat, K. Sinapiromsaran, and C. Lursinsap, “DBSMOTE: Density-based synthetic minority over-sampling technique,” Applied Intelligence, vol. 36, no. 3, pp. 664–684 (2012).[8] N. V. Chawla, K. W. Bowyer, L. O. Hall, and W. P. Kegelmeyer, “SMOTE: Synthetic minority over-sampling technique,” Journal of Artificial Intelligence Research, vol. 16, pp. 321–357 (2002).[9] G. Cohen, M. Hilario, H. Sax, A. Hugonnet, and V. François, “Learning from imbalanced data in surveillance of nosocomial infection,” Artificial Intelligence in Medicine, vol. 37, no. 1, pp. 7–15 (2006).[10] R. N. Dave, “Robust fuzzy clustering algorithms,” in Proc. 2nd IEEE Int. Conf. Fuzzy Systems, IEEE, pp. 1281–1286 (1993).[11] R. N. Davé and R. Krishnapuram, “Robust clustering methods: A unified view,” IEEE Transactions on Fuzzy Systems, vol. 5, no. 2, pp. 270–293 (1997).[12] C. Denoyer, “Web spam challenge 2007,” [Online]. Available: http://www.corpusetud.univ-mlv.fr/experiments/webspam/[13] G. Douzas and F. Bacao, “Self-organizing map oversampling (SOMO) for imbalanced data set learning,” Expert Systems with Applications, vol. 82, pp. 40–52 (2017).[14] M. Ester, H.-P. Kriegel, J. Sander, and X. Xu, “A density-based algorithm for discovering clusters in large spatial databases with noise,” in Proc. 2nd Int. Conf. Knowledge Discovery and Data Mining (KDD), pp. 226–231 (1996).[15] A. Fernández, S. Garcia, F. Herrera, and N. Chawla, “SMOTE for learning from imbalanced data: progress and challenges, marking the 15-year anniversary,” J. Artif. Intell. Res., vol. 61, pp. 863–905 (2018).[16] A. Ghazikhani, H. S. Yazdi, and R. Monsefi, “Class imbalance handling using wrapper-based random oversampling,” in Proc. 20th Iranian Conf. Electrical Engineering (ICEE), IEEE, pp. 611–616 (2012).[17] H. Han, W.-Y. Wang, and B.-H. Mao, “Borderline-SMOTE: A new over-sampling method in imbalanced data sets learning,” in Int. Conf. Intelligent Computing, Springer, pp. 878–887 (2005).[18] H. He, Y. Bai, E. A. Garcia, and S. Li, “ADASYN: Adaptive synthetic sampling approach for imbalanced learning,” in Proc. 2008 IEEE Int. Joint Conf. Neural Networks (IEEE World Congress on Computational Intelligence), pp. 1322–1328 (2008).[19] X. Huang, C.-Z. Zhang, and J. Yuan, “Predicting extreme financial risks on imbalanced dataset: A combined kernel FCM and kernel SMOTE based SVM classifier,” Comput. Econ., vol. 56, no. 1, pp. 187–216 (2020).[20] P. Kaur and A. Gosain, “An intelligent undersampling technique based upon intuitionistic fuzzy sets to alleviate class imbalance problem of classification with noisy environment,” Int. J. Intell. Eng. Inform., vol. 6, no. 5, pp. 417–433 (2018).[21] P. Kaur and A. Gosain, “FF-SMOTE: A metaheuristic approach to combat class imbalance in binary classification,” Appl. Artif. Intell., vol. 33, no. 5, pp. 420–439 (2019).[22] R. Krishnapuram and J. M. Keller, “A possibilistic approach to clustering,” IEEE Trans. Fuzzy Syst., vol. 1, no. 2, pp. 98–110 (1993).[23] F. Last, G. Douzas, and F. Bacao, “Oversampling for imbalanced learning based on k-means and SMOTE,” arXiv preprint arXiv:1711.00837 (2017).[24] N. Lunardon, G. Menardi, and N. Torelli, “ROSE: A package for binary imbalanced learning,” R J., vol. 6, no. 1 (2014).[25] R. Malhotra and S. Kamal, “An empirical study to investigate oversampling methods for improving software defect prediction using imbalanced data,” Neurocomputing, vol. 343, pp. 120–140 (2019).[26] I. Nekooeimehr and S. K. Lai-Yuen, “Adaptive semi-unsupervised weighted oversampling (A-SUWO) for imbalanced datasets,” Expert Syst. Appl., vol. 46, pp. 405–416 (2016).[27] N. R. Pal, K. Pal, J. M. Keller, and J. C. Bezdek, “A possibilistic fuzzy c-means clustering algorithm,” IEEE Trans. Fuzzy Syst., vol. 13, no. 4, pp. 517–530 (2005).[28] P. Kaur, I. Lamba, and G. Anjana, “DOFCM: A robust clustering technique based upon density,” Int. J. Eng. Technol., vol. 3, no. 3, p. 297 (2011).[29] S. Sharma, A. Gosain, and S. Jain, “A review of the oversampling techniques in class imbalance problem,” in Int. Conf. Innovative Computing and Communications, Springer, pp. 459–472 (2022).[30] B. Tang and H. He, “KernelADASYN: Kernel based adaptive synthetic data generation for imbalanced learning,” in 2015 IEEE Congress on Evolutionary Computation (CEC), pp. 664–671 (2015).[31] S. Vats and B. Sagar, “Performance evaluation of k-means clustering on Hadoop infrastructure,” J. Discrete Math. Sci. Cryptogr., vol. 22, no. 8, pp. 1349–1363 (2019).[32] Y. Xu, C. Wu, K. Zheng, L. Yu, and G. Zhang, “Fuzzy–Synthetic Minority Oversampling Technique: Oversampling based on fuzzy set theory for Android malware detection in imbalanced datasets,” Int. J. Distrib. Sensor Netw., vol. 13, no. 4, Article ID 1550147717703116 (2017).
Views: 81Downloads: 6Citations: 0