<?xml version="1.0" encoding="UTF-8"?>
<article article-type="Research Article">
  <front>
    <journal-meta>
      <journal-id journal-id-type="publisher">journal-of-statistics-and-management-systems</journal-id>
      <journal-title-group>
        <journal-title> Journal of Statistics and Management Systems</journal-title>
      </journal-title-group>
      <issn publication-format="electronic">2169-0014</issn>
      <issn publication-format="print">0972-0510</issn>
      <publisher>
        <publisher-name>Taru Publications</publisher-name>
      </publisher>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.47974/JSMS-971</article-id>
      <title-group>
        <article-title>Comparing the performance of eight imputation methods for propensity score matching in missing data problem</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author" corresp="yes">
          <name>
            <surname>Omurlu</surname>
            <given-names>Imran Kurt</given-names>
          </name>
          <aff>Division of Biostatistics, Faculty of Medicine, Adnan Menderes University, Aydin, Turkey</aff>
        </contrib>
        <contrib contrib-type="author">
          <name>
            <surname>Varol</surname>
            <given-names>Bugra</given-names>
          </name>
          <aff>Division of Biostatistics, Institute of Health Sciences, Adnan Menderes University, Aydin, Turkey</aff>
        </contrib>
        <contrib contrib-type="author">
          <name>
            <surname>Ture</surname>
            <given-names>Mevlut</given-names>
          </name>
          <aff>Division of Biostatistics, Faculty of Medicine, Adnan Menderes University, Aydin, Turkey</aff>
        </contrib>
      </contrib-group>
      <volume>26</volume>
      <issue>4</issue>
      <fpage>915</fpage>
      <lpage>927</lpage>
      <pub-date date-type="pub">
        <day>10</day>
        <month>08</month>
        <year>2023</year>
      </pub-date>
      <abstract>
        <p>Propensity score (PS) is a popular method to control for covariates in observational studies. A challenge in PS analyses is missing values in covariates. This study aims to investigate how different imputation methods of handling missing values of covariates in a PS analysis can affect average treatment on the treated (ATT) estimates. In this study, missing data imputation methods were evaluated using different data sets, whose covariates were low, medium, and high (r=0.10, 0.50, 0.85) correlated with each other, for n=200 units and 1000 times running simulation. Missing data structures were created according to the missing at random (MAR) mechanism and different missing rates. Different datasets were obtained after having imputed the missing values separately by eight imputation methods including mean, median, mode, hot deck, last observation carried forward (LOCF), next observation carried backward (NOCB), regression and predictive mean matching (PMM). Then the PS nearest neighbor matching was implemented and ATT scores were obtained using the imputed data sets. The predictive performance of imputation methods was compared according to ATT scores by hierarchical cluster analysis with Euclidean distance complete linkage. ATT scores of regression and PMM methods were closer to each other and these methods showed the best predictive performance. Additionally, when there were larger amounts of missing data, the PMM was the best method of choice. Ignoring missing values on covariates for PS analyses causes information loss significantly and this information loss becomes greater as the rate of missing data increases. PS analyses might be biased if missing data on covariates are also ignored. To prevent this information loss and bias, PS analyses should be performed after solving the problem of missing data with MAR mechanism on covariates by regression and PMM methods, which showed statistical superiority compared to other methods in this study.</p>
      </abstract>
      <kwd-group>
        <kwd>Missing data</kwd>
        <kwd>Imputation</kwd>
        <kwd>Simulation</kwd>
        <kwd>Propensity score</kwd>
        <kwd>Hierarchical clustering</kwd>
      </kwd-group>
      <custom-meta-group>
        <custom-meta>
          <meta-name>access</meta-name>
          <meta-value>open</meta-value>
        </custom-meta>
        <custom-meta>
          <meta-name>retracted</meta-name>
          <meta-value>no</meta-value>
        </custom-meta>
      </custom-meta-group>
    </article-meta>
  </front>
</article>
