<?xml version="1.0" encoding="UTF-8"?>
<article article-type="Research Article">
  <front>
    <journal-meta>
      <journal-id journal-id-type="publisher">journal-of-statistics-and-management-systems</journal-id>
      <journal-title-group>
        <journal-title> Journal of Statistics and Management Systems</journal-title>
      </journal-title-group>
      <issn publication-format="electronic">2169-0014</issn>
      <issn publication-format="print">0972-0510</issn>
      <publisher>
        <publisher-name>Taru Publications</publisher-name>
      </publisher>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.47974/JSMS-1280</article-id>
      <title-group>
        <article-title>Diabetes prediction using machine learning classifiers with oversampling and feature augmentation</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author" corresp="yes">
          <name>
            <surname>Banday</surname>
            <given-names>Mehroush</given-names>
          </name>
          <aff>Department of Computer Science, School of Engineering Sciences and Technology, Jamia Hamdard, New Delhi, 110062, India</aff>
        </contrib>
        <contrib contrib-type="author">
          <name>
            <surname>Zafar</surname>
            <given-names>Sherin</given-names>
          </name>
          <aff>Department of Computer Science, School of Engineering Sciences and Technology, Jamia Hamdard, New Delhi, 110062, India</aff>
        </contrib>
        <contrib contrib-type="author">
          <name>
            <surname>Alam</surname>
            <given-names>M. Afshar</given-names>
          </name>
          <aff>Department of Computer Science, School of Engineering Sciences and Technology, Jamia Hamdard, New Delhi, 110062, India</aff>
        </contrib>
      </contrib-group>
      <volume>27</volume>
      <issue>2</issue>
      <fpage>455</fpage>
      <lpage>464</lpage>
      <pub-date date-type="pub">
        <day>30</day>
        <month>03</month>
        <year>2024</year>
      </pub-date>
      <abstract>
        <p>Two of the leading causes of mortality in the US are diabetes and cardiovascular disease. The first step in halting the course of these disorders in patients is to recognize and anticipate them. Using survey data (as well as test findings), we assess the efficacy of machine learning algorithms for identifying patients who are at risk and pinpoint important characteristics in the data that are causing these diseases in the patients. Hunger, thirst, and frequent urination are signs of high blood sugar. Neglecting diabetes can lead to several issues. It’s critical to be able to identify diabetes early. In this study, we use supervised machine-learning techniques such as Decision Tree, KNN, (RF)Random Forest, Logistic Regression, Ada Boost, and Gradient Boosting to train on the actual data of 520 diabetics. The proposed work has achieved a more thorough comparative analysis between different datasets and their features that may be carried out to pinpoint all the essential characteristics for predicting diabetes.</p>
      </abstract>
      <kwd-group>
        <kwd>Gradient boosting</kwd>
        <kwd>KNN</kwd>
        <kwd>Ada boost</kwd>
        <kwd>Cardiovascular disease</kwd>
        <kwd>Diabetes mellitus</kwd>
      </kwd-group>
      <custom-meta-group>
        <custom-meta>
          <meta-name>access</meta-name>
          <meta-value>open</meta-value>
        </custom-meta>
        <custom-meta>
          <meta-name>retracted</meta-name>
          <meta-value>no</meta-value>
        </custom-meta>
      </custom-meta-group>
    </article-meta>
  </front>
</article>
