<?xml version="1.0" encoding="UTF-8"?>
<article article-type="Research Article">
  <front>
    <journal-meta>
      <journal-id journal-id-type="publisher">journal-of-statistics-and-management-systems</journal-id>
      <journal-title-group>
        <journal-title> Journal of Statistics and Management Systems</journal-title>
      </journal-title-group>
      <issn publication-format="electronic">2169-0014</issn>
      <issn publication-format="print">0972-0510</issn>
      <publisher>
        <publisher-name>Taru Publications</publisher-name>
      </publisher>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.47974/JSMS-957</article-id>
      <title-group>
        <article-title>Sampling strategies for handling data imbalance problem: An Extensive Review</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <name>
            <surname>Veedhi</surname>
            <given-names>Bhaskar Kumar</given-names>
          </name>
          <aff>Department of Computer Science and Engineering, Siksha ‘O’ Anusandhan (Deemed to be University), Bhubaneswar, Odisha, India</aff>
        </contrib>
        <contrib contrib-type="author">
          <name>
            <surname>Mishra</surname>
            <given-names>Debahuti</given-names>
          </name>
          <aff>Department of Computer Science and Engineering, Siksha ‘O’ Anusandhan (Deemed to be University), Bhubaneswar, Odisha, India</aff>
        </contrib>
        <contrib contrib-type="author" corresp="yes">
          <name>
            <surname>Das</surname>
            <given-names>Kaberi</given-names>
          </name>
          <aff>Department of Computer Applications, Siksha ‘O’ Anusandhan (Deemed to be University), Bhubaneswar, Odisha, India</aff>
        </contrib>
      </contrib-group>
      <volume>26</volume>
      <issue>1</issue>
      <fpage>177</fpage>
      <lpage>187</lpage>
      <pub-date date-type="pub">
        <day>31</day>
        <month>12</month>
        <year>2022</year>
      </pub-date>
      <abstract>
        <p>The imbalanced data classification is a major issue in data mining. Many researchers have proposed various solutions which addressed imbalanced data problem which is broadly categorized into data level and algorithm level. Class distributions are adjusted in data level method. Creating an algorithm or modifying the existing algorithm is an appropriate approach used in algorithm level method. Imbalanced data classification problem can be resolved by means of Sampling, Random over sampling, Random under sampling, Resampling and by SMOTE (Synthetic Minority Oversampling Techniques). Resampling includes k-means clustering, density-based clustering, neural networks and ensemble. However, no algorithm or a method has an ability to remove bias in data classification, thereby integration of kernel methods with sampling methods or integration of sampling and boosting methods or integration Kernel based with Support Vector Machines (SVM) need to be performed a great extent to get the desired accuracy and performance. The main objective of this paper is to focus on various sampling strategies that are based on sampling and resampling methods and improving the concept of learning within class imbalanced data. It also explains the objectives of the models used by several researchers and emphasized the performance along with the outcomes.</p>
      </abstract>
      <kwd-group>
        <kwd>Imbalanced data</kwd>
        <kwd>Classification</kwd>
        <kwd>Skewed data</kwd>
        <kwd>Sampling methods</kwd>
        <kwd>SMOTE</kwd>
      </kwd-group>
      <custom-meta-group>
        <custom-meta>
          <meta-name>access</meta-name>
          <meta-value>open</meta-value>
        </custom-meta>
        <custom-meta>
          <meta-name>retracted</meta-name>
          <meta-value>no</meta-value>
        </custom-meta>
      </custom-meta-group>
    </article-meta>
  </front>
</article>
