{"title":"当同一数据集中存在MCAR、MAR和MNAR机制的混合时,处理缺失数据的方法的真实评估。","authors":"Brenna Gomer, Ke-Hai Yuan","doi":"10.1080/00273171.2022.2158776","DOIUrl":null,"url":null,"abstract":"<p><p>The impact of missing data on statistical inference varies depending on several factors such as the proportion of missingness, missing-data mechanism, and method employed to handle missing values. While these topics have been extensively studied, most recommendations have been made assuming that all missing values are from the same missing-data mechanism. In reality, it is very likely that a mixture of missing-data mechanisms is responsible for missing values in a dataset and even within the same pattern of missingness. Although a mixture of missing-data mechanisms and causes within a dataset is a likely scenario, the performance of popular missing-data methods under these circumstances is unknown. This study provides a realistic evaluation of methods for handling missing data in this setting using Monte Carlo simulation in the context of regression. This study also seeks to identify acceptable proportions of missing values that violate the missing-data mechanism assumed by the method used to handle missing values. Results indicate that multiple imputation (MI) performs better than other principled or ad-hoc methods. Different missing-data methods are also compared via the analysis of a real dataset in which mixtures of missingness mechanisms are created. Recommendations are provided for the use of different methods in practice.</p>","PeriodicalId":53155,"journal":{"name":"Multivariate Behavioral Research","volume":null,"pages":null},"PeriodicalIF":5.3000,"publicationDate":"2023-09-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"2","resultStr":"{\"title\":\"A Realistic Evaluation of Methods for Handling Missing Data When There is a Mixture of MCAR, MAR, and MNAR Mechanisms in the Same Dataset.\",\"authors\":\"Brenna Gomer, Ke-Hai Yuan\",\"doi\":\"10.1080/00273171.2022.2158776\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"<p><p>The impact of missing data on statistical inference varies depending on several factors such as the proportion of missingness, missing-data mechanism, and method employed to handle missing values. While these topics have been extensively studied, most recommendations have been made assuming that all missing values are from the same missing-data mechanism. In reality, it is very likely that a mixture of missing-data mechanisms is responsible for missing values in a dataset and even within the same pattern of missingness. Although a mixture of missing-data mechanisms and causes within a dataset is a likely scenario, the performance of popular missing-data methods under these circumstances is unknown. This study provides a realistic evaluation of methods for handling missing data in this setting using Monte Carlo simulation in the context of regression. This study also seeks to identify acceptable proportions of missing values that violate the missing-data mechanism assumed by the method used to handle missing values. Results indicate that multiple imputation (MI) performs better than other principled or ad-hoc methods. Different missing-data methods are also compared via the analysis of a real dataset in which mixtures of missingness mechanisms are created. Recommendations are provided for the use of different methods in practice.</p>\",\"PeriodicalId\":53155,\"journal\":{\"name\":\"Multivariate Behavioral Research\",\"volume\":null,\"pages\":null},\"PeriodicalIF\":5.3000,\"publicationDate\":\"2023-09-01\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"2\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"Multivariate Behavioral Research\",\"FirstCategoryId\":\"102\",\"ListUrlMain\":\"https://doi.org/10.1080/00273171.2022.2158776\",\"RegionNum\":3,\"RegionCategory\":\"心理学\",\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"2023/1/4 0:00:00\",\"PubModel\":\"Epub\",\"JCR\":\"Q1\",\"JCRName\":\"MATHEMATICS, INTERDISCIPLINARY APPLICATIONS\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"Multivariate Behavioral Research","FirstCategoryId":"102","ListUrlMain":"https://doi.org/10.1080/00273171.2022.2158776","RegionNum":3,"RegionCategory":"心理学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"2023/1/4 0:00:00","PubModel":"Epub","JCR":"Q1","JCRName":"MATHEMATICS, INTERDISCIPLINARY APPLICATIONS","Score":null,"Total":0}
A Realistic Evaluation of Methods for Handling Missing Data When There is a Mixture of MCAR, MAR, and MNAR Mechanisms in the Same Dataset.
The impact of missing data on statistical inference varies depending on several factors such as the proportion of missingness, missing-data mechanism, and method employed to handle missing values. While these topics have been extensively studied, most recommendations have been made assuming that all missing values are from the same missing-data mechanism. In reality, it is very likely that a mixture of missing-data mechanisms is responsible for missing values in a dataset and even within the same pattern of missingness. Although a mixture of missing-data mechanisms and causes within a dataset is a likely scenario, the performance of popular missing-data methods under these circumstances is unknown. This study provides a realistic evaluation of methods for handling missing data in this setting using Monte Carlo simulation in the context of regression. This study also seeks to identify acceptable proportions of missing values that violate the missing-data mechanism assumed by the method used to handle missing values. Results indicate that multiple imputation (MI) performs better than other principled or ad-hoc methods. Different missing-data methods are also compared via the analysis of a real dataset in which mixtures of missingness mechanisms are created. Recommendations are provided for the use of different methods in practice.
期刊介绍:
Multivariate Behavioral Research (MBR) publishes a variety of substantive, methodological, and theoretical articles in all areas of the social and behavioral sciences. Most MBR articles fall into one of two categories. Substantive articles report on applications of sophisticated multivariate research methods to study topics of substantive interest in personality, health, intelligence, industrial/organizational, and other behavioral science areas. Methodological articles present and/or evaluate new developments in multivariate methods, or address methodological issues in current research. We also encourage submission of integrative articles related to pedagogy involving multivariate research methods, and to historical treatments of interest and relevance to multivariate research methods.