TARU PUBLICATIONS
 Journal of Statistics and Management Systems cover
Open Access ·Peer-reviewed·ISSN (Online): 2169-0014·ISSN (Print): 0972-0510
Powered by:Powered by

The Journal of Statistics and Management Systems (JSMS) is a world leading journal publishing high quality, rigorously peer-reviewed original research on theoretical and applied statistics and management systems since 1998. The scope is intentionally broad, but papers must make a novel contribution to the field to be considered for publication. Topics include, but are not limited to, the following: • Statistics • Applied Statistics • Industrial Statistics • Statistical Inference • Interdisciplinary role of Statistics • Actuarial Sciences • Decision Sciences • Managerial Aspects • Management Sciences • Management Information Systems

Issues up to 2022 co-published with and available at:Taylor & Francis Online
submissions@tarupublications.com
Open Access Research Article

Sampling strategies for handling data imbalance problem: An Extensive Review

, , *

* Corresponding author · click or hover a name for details

pp. 177–187Vol. 26Issue 1December 2022DOI: 10.47974/JSMS-957XML
Published Online:
31 Dec 2022
Article type:
Research Article
Language:
EN
Article no.:
JSMS-957
Pages:
177–187

Abstract

The imbalanced data classification is a major issue in data mining. Many researchers have proposed various solutions which addressed imbalanced data problem which is broadly categorized into data level and algorithm level. Class distributions are adjusted in data level method. Creating an algorithm or modifying the existing algorithm is an appropriate approach used in algorithm level method. Imbalanced data classification problem can be resolved by means of Sampling, Random over sampling, Random under sampling, Resampling and by SMOTE (Synthetic Minority Oversampling Techniques). Resampling includes k-means clustering, density-based clustering, neural networks and ensemble. However, no algorithm or a method has an ability to remove bias in data classification, thereby integration of kernel methods with sampling methods or integration of sampling and boosting methods or integration Kernel based with Support Vector Machines (SVM) need to be performed a great extent to get the desired accuracy and performance. The main objective of this paper is to focus on various sampling strategies that are based on sampling and resampling methods and improving the concept of learning within class imbalanced data. It also explains the objectives of the models used by several researchers and emphasized the performance along with the outcomes.

Keywords

Subject Classifications

68T07

References

[1] Hongwei Ding, Leiyang Chen, Liang Dong, Zhongwang Fu, XiaohuiCui,”Imbalanced data classification: A KNN and generative adversarial networks-based hybrid approach for intrusion detection”, Future Generation Computer Systems, 131(2022) 240-254.
[2] Wenyang Wang, Dongchu Sun, “The improved AdaBoost algorithms for imbalanced data classification”, Information Sciences, 563(2021) 358-374.
[3] Jiakun Zhao, Ju Jin, Si Chen, Ruifeng Zhang, Bilin Yu, Qingfang Liu, “A weighted hybrid ensemble method for classifying imbalanced data”, Knowledge-Based Systems, 203(2020)106087.
[4] PabasaraAthukorala, Madusha Chathurangi & Rajitha Ranasinghe, “A variant of RSA using continued fractions”, Journal of Discrete Mathematical Sciences and Cryptography, Volume 25, Issue 1 (2022)127-134.
[5] Ruchi Kaushik, Vijander Singh & Rajani Kumari, “Multi-class SVM based network intrusion detection with attribute selection using infinite feature selection technique”, Journal of Discrete Mathematical Sciences and Cryptography, Volume 24, Issue 8 (2021)2137-2153.
[6] VijayasriIyer, Bhargava Ganti, A. M. Hima Vyshnavi, P. K. Krishnan Namboori&Sriram Iyer, “Hybrid quantum computing based early detection of skin cancer”, Journal of Interdisciplinary Mathematics, Volume 23, Issue 2 (2020)347-355.
[7] Osama R. Shahin, Rasha M. Abd El-Aziz&Ahmed I. Taloba, “Detection and classification of Covid-19 in CT-lungs screening using machine learning techniques”, Journal of Interdisciplinary Mathematics, Volume 25, Issue 3 (2022)791-813.
[8] Quoc Hoan Doan, Sy-Hung Mai, Quang Thang Do, Duc-Kien Thai, “A cluster-based data splitting method for small sample and class imbalance problems in impact damage classification”, Applied Soft Computing, 120 (2022) 108628.
[9] Ming Zheng, Tong Li, Liping Sun, Taochun Wang, Biao Jie, Weiyi Yang, Mingjing Tang, Changlong Lv, “An automatic sampling ratio detection method based on genetic algorithm for imbalanced data classification”, Knowledge-Based Systems, 216(2021) 106800.
[10] Xin Gao, Bing Ren, Hao Zhang, Bohao Sun, Junliang Li, Jianhang Xu, Yang He, Kangsheng Li, “An ensemble imbalanced classification method based on model dynamic selection driven by data partition hybrid sampling”, Expert Systems with Applications, 160 (2020) 113660.
[11] Dohyun Lee, Kyoungok Kim, “An efficient method to determine sample size in oversampling based on classification complexity for imbalanced data”, Expert Systems with Applications, 184(2021) 115442.
[12] ChenxiaJin, Fachao Li, Shijie Ma, Ying Wang, “Sampling scheme-based classification rule mining method using decision tree in big data environment”, Knowledge-Based Systems, 244(2022)108522.
[13] G.Douzas, F.Bacao, F.Last, “Improving imbalanced learning through a heuristic oversampling method based on k-means and smote” Inf.Sci. 465 (2018) 1–20.
[14] Rodolfo M. Pereira, Yandre M.G. Costa, Carlos N. Silla Jr., “Toward hierarchical classification of imbalanced data using random resampling algorithms”, Information Sciences, 578(2021) 344-363.
[15] Rashmi Dubey, Jiayu Zhou, Yalin Wang, Paul M. Thompson, Jieping Ye “Analysis of sampling techniques for imbalanced data: An n = 648 ADNI study” NeuroImage 87 (2014) 220-241.
[16] Huaxiang Zhang, Mingfang Li “RWO-Sampling: A random walk over-sampling approach to imbalanced data classification” Information Fusion 20 (2014) 99-116.
[17] Myoung-Jong Kim, Dae-Ki Kang, Hong Bae Kim “Geometric mean based boosting algorithm with over-sampling to resolve data imbalance problem for bankruptcy prediction” Expert systems with Application 42 (2015) 1074-1082.
[18] Kyoham Shin, Jongmin Han, Seokho Kang “MI-MOTE: Multiple imputation-based minority oversampling technique for imbalanced and incomplete data classification” Information Science 575 (2021) 80-89.
[19] Ankit Vijayvargiya, Chandra Prakash, Rajesh Kumar, Sanjeev Bansal, Joao Manuel R. S. Tavares “Human knee abnormality detection from imbalanced sEMG data” Biomedical Signal Processing and Control 66 (2021) 102406
[20] Justin Engelmann, Stefan Lessmann “Conditional Wasserstein GAN-based oversampling of tabular data for imbalanced learning” Expert systems with Application 174 (2021) 114582.
[21] D.Devi, B.Purkayastha, “Redundancy-driven modified tomek-link based under sampling: A solution to class imbalance”PatternRecognit. Lett.93 (2017) 3–12.
[22] C.Bunkhumpornpat, K.Sinapiromsaran, “Dbmute: density-based majority under-sampling technique” Knowl. Inf. Syst. 50 (3) (2017) 827–850.
[23] Hualong Yu, Jun Ni, Jing Zhao “ACO Sampling: An ant colony optimization-based undersampling method for classifying imbalanced DNA microarray data” Neurocomputing 101 (2013) 309-318.
[24] Mikel Galar, Alberto Fernández, EdurneBarrenechea, Francisco Herrera “EUSBoost: Enhancing ensembles for highly imbalanced data-sets by evolutionary undersampling” Pattern Recognition 46 (2013) 3460-3471.
[25] Peng Cao, Jinzhu Yanga, Wei Li, Dazhe Zhao, OsmarZaianec “Ensemble-based hybrid probabilistic sampling for imbalanced data learning in lung nodule CAD” Computerized Medical Imaging and Graphics 38 (2014) 137-150.
[26] J. Hoyos-Osorio, A. Alvarez-Meza, G. Daza-Santacoloma, A. Orozco-Gutierrez, G. Castellanos-Dominguez “Relevant information undersampling to support imbalanced data classification” Neurocomputing 436 (2021) 136-146.
[27] Zhaozhao Xu, Derong Shen, TiezhengNie, Yue Kou, “A hybrid sampling algorithm combining M-SMOTE and ENN based on Random forest for medical imbalanced data”, Journal of Biomedical Informatics, 107(2020), 103465.
[28] José A. Sáez, Julián Luengo, Jerzy Stefanowski, Francisco Herrera, “SMOTE–IPF: Addressing the noisy and borderline examples problem in imbalanced classification by a re-sampling method with filtering”, Information Sciences, 291(2015) 184-203.
[29] Ming Hao, Yanli Wang∗, Stephen H. Bryant “An efficient algorithm coupled with synthetic minority over-sampling technique to classify imbalanced PubChem BioAssay data”Analytica Chimica Acta 806 (2014) 117-127.
[30] Yuqiang Fan, Xiaoyu Cui, Hua Han, Hailong Lu “Chiller fault diagnosis with field sensors using the technology of imbalanced data” Applied Thermal Engineering 159 (2019) 113933.
[31] Amir BahadorParsaa, HomaTaghipoura, Sybil Derribleb , Abolfazl (Kouros) Mohammadianc “Real-time accident detection: Coping with imbalanced data” Accident Analysis and Preventation 129 (2019) 202-210.
[32] Hong-Jie Dai, Chen-Kai Wang “Classifying adverse drug reactions from imbalanced twitter data” International Journal of Medical Informatics 129 (2019) 122-132 2021.
[33] Asniar, Nur UlfaMaulidev, KridantoSurendro “SMOTE-LOF for noise identification in imbalanced data classification” Journal of King Saud University – Computer and Information Sciences, (2021).

Views: 149Downloads: 6Citations: 1