TARU PUBLICATIONS
Journal of Information and Optimization Sciences cover
Hybrid ·Peer-reviewed·ISSN (Online): 2169-0103·ISSN (Print): 0252-2667

WoS  JIF 2026 : 0.4 (Q4)

Powered by:Powered by

Monthly Journal: Publishes theoretical and applied research on topics in information and optimization sciences.

Issues up to 2022 co-published with and available at:Taylor & Francis
submissions@tarupublications.com
Open Access Research Article

DOFCM-SMOTE : A technique to handle class imbalance problem

* ,

* Corresponding author · click or hover a name for details

pp. 1705–1723Vol. 46Issue 5July 2025DOI: 10.47974/JIOS-1946XML
Received:
07 Aug 2024
Published Online:
15 Jul 2025
Article type:
Research Article
Language:
EN
Article no.:
JIOS-1946
Pages:
1705–1723

Abstract

Class imbalance is a common issue in real-world datasets, where there are relatively few instances of the class of interest compared to others. The classifier’s performance is highly affected by the impurities built-in with the data like imbalanced data, noise, and class overlapping; therefore, an oversampling technique is proposed by fusing density-oriented fuzzy-c means clustering and synthetic minority oversampling technique(DOFCM-SMOTE). The density-oriented approach is intended to find outliers to address the performance issue. The proposed method works in 3 phases: 1) It detects and eliminates outliers in the dataset, 2) It Assigns membership degrees to the given samples by fuzzy c-means clustering, and 3) It then applies SMOTE to the minority cluster to balance the distribution of classes. To determine the performance, a decision tree classifier is employed as a learner. This study was conducted on ten publicly available data sets with varying imbalance ratios(IR) ranging from high to low. The findings of this study are centric on the process of clustering through which outliers are detected even before minority and majority class clusters are identified, and a general pre-processing technique, to balance the distribution of classes. The introduced approach is compared with four other oversampling techniques and the best balancing algorithm is determined. The performances are accessed through AUC and G-mean, and it was reported that the proposed technique significantly improved and outperformed the other over-sampling methods.

Keywords

Subject Classifications

68T05:Learning and adaptive systems in artificial intelligence

References

[1] J. Alcalá-Fdez, L. Sánchez, S. García, M. J. del Jesus, S. Ventura, J. M. Benítez, and F. Herrera, “KEEL: A software tool to assess evolutionary algorithms for data mining problems,” Soft Computing, vol. 13, no. 3, pp. 307–318 (2009).
[2] J. Alcalá-Fdez, A. Fernández, J. Luengo, J. Derrac, S. García, L. Sánchez, and F. Herrera, “KEEL data-mining software tool: data set repository, integration of algorithms and experimental analysis framework,” Journal of Multiple-Valued Logic & Soft Computing, vol. 17, pp. 255–287 (2011).
[3] R. Barandela, J. S. Sánchez, V. García, and E. Rangel, “Strategies for learning in class imbalance problems,” Pattern Recognition, vol. 36, no. 3, pp. 849–851 (2003).
[4] S. Barua, M. M. Islam, K. Murase, and X. Yao, “A novel synthetic minority oversampling technique for imbalanced data set learning,” in Proc. Int. Conf. Neural Information Processing, Springer, pp. 735–744 (2011).
[5] S. Barua, M. M. Islam, K. Murase, and X. Yao, “ProWSyn: Proximity weighted synthetic oversampling technique for imbalanced data set learning,” in Proc. Pacific-Asia Conf. Knowledge Discovery and Data Mining, Springer, pp. 317–328 (2013).
[6] C. Bunkhumpornpat, K. Sinapiromsaran, and C. Lursinsap, “Safe-Level-SMOTE: Safe-level-synthetic minority over-sampling technique for handling the class imbalanced problem,” in Proc. Pacific-Asia Conf. Knowledge Discovery and Data Mining, Springer, pp. 475–482 (2009).
[7] C. Bunkhumpornpat, K. Sinapiromsaran, and C. Lursinsap, “DBSMOTE: Density-based synthetic minority over-sampling technique,” Applied Intelligence, vol. 36, no. 3, pp. 664–684 (2012).
[8] N. V. Chawla, K. W. Bowyer, L. O. Hall, and W. P. Kegelmeyer, “SMOTE: Synthetic minority over-sampling technique,” Journal of Artificial Intelligence Research, vol. 16, pp. 321–357 (2002).
[9] G. Cohen, M. Hilario, H. Sax, A. Hugonnet, and V. François, “Learning from imbalanced data in surveillance of nosocomial infection,” Artificial Intelligence in Medicine, vol. 37, no. 1, pp. 7–15 (2006).
[10] R. N. Dave, “Robust fuzzy clustering algorithms,” in Proc. 2nd IEEE Int. Conf. Fuzzy Systems, IEEE, pp. 1281–1286 (1993).
[11] R. N. Davé and R. Krishnapuram, “Robust clustering methods: A unified view,” IEEE Transactions on Fuzzy Systems, vol. 5, no. 2, pp. 270–293 (1997).
[12] C. Denoyer, “Web spam challenge 2007,” [Online]. Available: http://www.corpusetud.univ-mlv.fr/experiments/webspam/
[13] G. Douzas and F. Bacao, “Self-organizing map oversampling (SOMO) for imbalanced data set learning,” Expert Systems with Applications, vol. 82, pp. 40–52 (2017).
[14] M. Ester, H.-P. Kriegel, J. Sander, and X. Xu, “A density-based algorithm for discovering clusters in large spatial databases with noise,” in Proc. 2nd Int. Conf. Knowledge Discovery and Data Mining (KDD), pp. 226–231 (1996).
[15] A. Fernández, S. Garcia, F. Herrera, and N. Chawla, “SMOTE for learning from imbalanced data: progress and challenges, marking the 15-year anniversary,” J. Artif. Intell. Res., vol. 61, pp. 863–905 (2018).
[16] A. Ghazikhani, H. S. Yazdi, and R. Monsefi, “Class imbalance handling using wrapper-based random oversampling,” in Proc. 20th Iranian Conf. Electrical Engineering (ICEE), IEEE, pp. 611–616 (2012).
[17] H. Han, W.-Y. Wang, and B.-H. Mao, “Borderline-SMOTE: A new over-sampling method in imbalanced data sets learning,” in Int. Conf. Intelligent Computing, Springer, pp. 878–887 (2005).
[18] H. He, Y. Bai, E. A. Garcia, and S. Li, “ADASYN: Adaptive synthetic sampling approach for imbalanced learning,” in Proc. 2008 IEEE Int. Joint Conf. Neural Networks (IEEE World Congress on Computational Intelligence), pp. 1322–1328 (2008).
[19] X. Huang, C.-Z. Zhang, and J. Yuan, “Predicting extreme financial risks on imbalanced dataset: A combined kernel FCM and kernel SMOTE based SVM classifier,” Comput. Econ., vol. 56, no. 1, pp. 187–216 (2020).
[20] P. Kaur and A. Gosain, “An intelligent undersampling technique based upon intuitionistic fuzzy sets to alleviate class imbalance problem of classification with noisy environment,” Int. J. Intell. Eng. Inform., vol. 6, no. 5, pp. 417–433 (2018).
[21] P. Kaur and A. Gosain, “FF-SMOTE: A metaheuristic approach to combat class imbalance in binary classification,” Appl. Artif. Intell., vol. 33, no. 5, pp. 420–439 (2019).
[22] R. Krishnapuram and J. M. Keller, “A possibilistic approach to clustering,” IEEE Trans. Fuzzy Syst., vol. 1, no. 2, pp. 98–110 (1993).
[23] F. Last, G. Douzas, and F. Bacao, “Oversampling for imbalanced learning based on k-means and SMOTE,” arXiv preprint arXiv:1711.00837 (2017).
[24] N. Lunardon, G. Menardi, and N. Torelli, “ROSE: A package for binary imbalanced learning,” R J., vol. 6, no. 1 (2014).
[25] R. Malhotra and S. Kamal, “An empirical study to investigate oversampling methods for improving software defect prediction using imbalanced data,” Neurocomputing, vol. 343, pp. 120–140 (2019).
[26] I. Nekooeimehr and S. K. Lai-Yuen, “Adaptive semi-unsupervised weighted oversampling (A-SUWO) for imbalanced datasets,” Expert Syst. Appl., vol. 46, pp. 405–416 (2016).
[27] N. R. Pal, K. Pal, J. M. Keller, and J. C. Bezdek, “A possibilistic fuzzy c-means clustering algorithm,” IEEE Trans. Fuzzy Syst., vol. 13, no. 4, pp. 517–530 (2005).
[28] P. Kaur, I. Lamba, and G. Anjana, “DOFCM: A robust clustering technique based upon density,” Int. J. Eng. Technol., vol. 3, no. 3, p. 297 (2011).
[29] S. Sharma, A. Gosain, and S. Jain, “A review of the oversampling techniques in class imbalance problem,” in Int. Conf. Innovative Computing and Communications, Springer, pp. 459–472 (2022).
[30] B. Tang and H. He, “KernelADASYN: Kernel based adaptive synthetic data generation for imbalanced learning,” in 2015 IEEE Congress on Evolutionary Computation (CEC), pp. 664–671 (2015).
[31] S. Vats and B. Sagar, “Performance evaluation of k-means clustering on Hadoop infrastructure,” J. Discrete Math. Sci. Cryptogr., vol. 22, no. 8, pp. 1349–1363 (2019).
[32] Y. Xu, C. Wu, K. Zheng, L. Yu, and G. Zhang, “Fuzzy–Synthetic Minority Oversampling Technique: Oversampling based on fuzzy set theory for Android malware detection in imbalanced datasets,” Int. J. Distrib. Sensor Netw., vol. 13, no. 4, Article ID 1550147717703116 (2017).

Views: 178Downloads: 86Citations: 0