Open Access
Research Article
Comparing the performance of eight imputation methods for propensity score matching in missing data problem
*Imran Kurt OmurluCorresponding authorikurtomurlu@gmail.comDivision of BiostatisticsFaculty of MedicineAdnan Menderes UniversityAydin, TurkeyView full profile → , Bugra VarolDivision of BiostatisticsInstitute of Health SciencesAdnan Menderes UniversityAydin, TurkeyView full profile → , Mevlut TureDivision of BiostatisticsFaculty of MedicineAdnan Menderes UniversityAydin, TurkeyView full profile →
* Corresponding author · click or hover a name for details
- Received:
- 01 Oct 2021
- Accepted:
- 01 May 2022
- Published Online:
- 10 Aug 2023
- Article type:
- Research Article
- Language:
- EN
- Article no.:
- JSMS-971
- Pages:
- 915–927
Abstract
Propensity score (PS) is a popular method to control for covariates in observational studies. A challenge in PS analyses is missing values in covariates. This study aims to investigate how different imputation methods of handling missing values of covariates in a PS analysis can affect average treatment on the treated (ATT) estimates. In this study, missing data imputation methods were evaluated using different data sets, whose covariates were low, medium, and high (r=0.10, 0.50, 0.85) correlated with each other, for n=200 units and 1000 times running simulation. Missing data structures were created according to the missing at random (MAR) mechanism and different missing rates. Different datasets were obtained after having imputed the missing values separately by eight imputation methods including mean, median, mode, hot deck, last observation carried forward (LOCF), next observation carried backward (NOCB), regression and predictive mean matching (PMM). Then the PS nearest neighbor matching was implemented and ATT scores were obtained using the imputed data sets. The predictive performance of imputation methods was compared according to ATT scores by hierarchical cluster analysis with Euclidean distance complete linkage. ATT scores of regression and PMM methods were closer to each other and these methods showed the best predictive performance. Additionally, when there were larger amounts of missing data, the PMM was the best method of choice. Ignoring missing values on covariates for PS analyses causes information loss significantly and this information loss becomes greater as the rate of missing data increases. PS analyses might be biased if missing data on covariates are also ignored. To prevent this information loss and bias, PS analyses should be performed after solving the problem of missing data with MAR mechanism on covariates by regression and PMM methods, which showed statistical superiority compared to other methods in this study.
Keywords
Subject Classifications
62B1562D1062H30
References
[1]Akmam, E.F., et al. Multiple Imputation with Predictive Mean Matching Method for Numerical Missing Data. in 2019 3rd International Conference on Informatics and Computational Sciences (ICICoS). 2019. IEEE.
[2]Austin, P.C., An introduction to propensity score methods for reducing the effects of confounding in observational studies. Multivariate behavioral research, 2011. 46(3): p. 399-424.
[3]Caliendo, M. and S. Kopeinig, Some practical guidance for the implementation of propensity score matching. Journal of Economic Surveys, 2008. 22(1): p. 31-72.
[4]Cham, H. and S.G. West, Propensity score analysis with missing data. Psychological methods, 2016. 21(3): p. 427-445.
[5]Choi, J., O.M. Dekkers, and S. le Cessie, A comparison of different methods to handle missing data in the context of propensity score analysis. European Journal of Epidemiology, 2019. 34(1): p. 23-36.
[6]Coffman, D.L., J. Zhou, and X. Cai, Comparison of methods for handling covariate missingness in propensity score estimation with a binary exposure. BMC medical research methodology, 2020. 20(1): p. 1-14.
[7]D’Agostino Jr, R.B., Propensity score methods for bias reduction in the comparison of a treatment to a non randomized control group. Statistics in medicine, 1998. 17(19): p. 2265-2281.
[8]D’Agostino Jr, R.B. and D.B. Rubin, Estimating and using propensity scores with partially missing data. Journal of the American Statistical Association, 2000. 95(451): p. 749-759.
[9]Engels, J.M. and P. Diehr, Imputation of missing longitudinal data: a comparison of methods. Journal of Clinical Epidemiology, 2003. 56(10): p. 968-976.
[10] Hosmer Jr, D.W., S. Lemeshow, and R.X. Sturdivant, Applied logistic regression. 3rd edition, Vol. 398. 2013: John Wiley & Sons , Hoboken, NJ.
[11] Jacovidis, J.N., Evaluating the performance of propensity score matching methods: A simulation study. 2017, Doctoral dissertation, James Madison University.
[12] Jadhav, A., D. Pramod, and K. Ramanathan, Comparison of performance of data imputation methods for numeric dataset. Applied Artificial Intelligence, 2019. 33(10): p. 913-933.
[13] Jochen, H., et al., Multiple imputation of missing data: a simulation study on a binary response. Open Journal of Statistics, 2013. 2013.
[14] Kalton, G. and D. Kasprzyk. Imputing for missing survey responses. in Proceedings of the section on survey research methods, American Statistical Association. 1982. American Statistical Association Cincinnati.
[15] Klamroth-Marganska, V., et al., Three-dimensional, task-specific robot therapy of the arm after stroke: a multicentre, parallel-group randomised trial. The Lancet Neurology, 2014. 13(2): p. 159-166.
[16] Little, R.J. and D.B. Rubin, Statistical analysis with missing data. 3rd edition,Vol. 793. 2019: John Wiley & Sons, Hoboken, NJ.
[17] Liu, C.-H., et al., The feature selection effect on missing value imputation of medical datasets. Applied Sciences, 2020. 10(7): p. 2344.
[18] Mattei, A., Estimating and using propensity score in presence of missing background data: an application to assess the impact of childbearing on wellbeing. Statistical Methods and Applications, 2009. 18(2): p. 257-273.
[19] Mayer, B. and B. Puschner, Propensity score adjustment of a treatment effect with missing data in psychiatric health services research. Epidemiol Biostat Public Health, 2015. 12(1): p. 10.2427.
[20] Meena, N., Singh, B., Firefly optimization based hierarchical clustering algorithm in wireless sensor network. Journal of Discrete Mathematical Sciences and Cryptography, 2021. 24(6), p. 1717-1725.
[21] Misztal, M., Imputation of missing data using R package. Acta Universitatis Lodziensis. Folia Oecologica, 2012, 269: p. 131-144.
[22] Mitra, R. and J.P. Reiter, Estimating propensity scores with missing covariate data using general location mixture models. Statistics in medicine, 2011. 30(6): p. 627-641.
[23] Mitra, R. and J.P. Reiter, A comparison of two methods of estimating propensity scores after multiple imputation. Statistical methods in medical research, 2016. 25(1): p. 188-204.
[24] Molnar, F.J., B. Hutton, and D. Fergusson, Does analysis using “last observation carried forward” introduce bias in dementia research? Cmaj, 2008. 179(8): p. 751-753.
[25] Morris, T.P., I.R. White, and P. Royston, Tuning multiple imputation by predictive mean matching and local residual draws. BMC medical research methodology, 2014. 14(1): p. 1-13.
[26] Qu, Y. and I. Lipkovich, Propensity score estimation with missing values using a multiple imputation missingness pattern (MIMP) approach. Statistics in medicine, 2009. 28(9): p. 1402-1414.
[27] Rosenbaum, P.R. and D.B. Rubin, Constructing a control group using multivariate matched sampling methods that incorporate the propensity score. The American Statistician, 1985. 39(1): p. 33-38.
[28] Rubin, D.B., Inference and missing data. Biometrika, 1976. 63(3): p. 581-592.
[29] Vink, G., et al., Predictive mean matching imputation of semicontinuous variables. Statistica Neerlandica, 2014. 68(1): p. 61-90.
[30] Zhang, J., A Comparison of Propensity Score Matching Methods in R with the MatchIt Package: A Simulation Study. 2013, University of Cincinnati, Master’s thesis. OhioLINK Electronic Theses and Dissertations Center.
Views: 272Downloads: 87Citations: 0




