Stefan Kuhle, Mary Margaret Brown, Sanja Stanojevic
{"title":"Building a better model: abandon kitchen sink regression","authors":"Stefan Kuhle, Mary Margaret Brown, Sanja Stanojevic","doi":"10.1136/archdischild-2023-326340","DOIUrl":null,"url":null,"abstract":"This paper critically examines ‘kitchen sink regression’, a practice characterised by the manual or automated selection of variables for a multivariable regression model based on p values or model-based information criteria. We highlight the pitfalls of this method, using examples from perinatal/neonatal medicine, and propose more robust alternatives. The concept of directed acyclic graphs (DAGs) is introduced as a tool for describing and analysing causal relationships. We highlight five key issues with ‘kitchen sink regression’: (1) the disregard for the directionality of variable relationships, (2) the lack of a meaningful causal interpretation of effect estimates from these models, (3) the inflated alpha error rate due to multiple testing, (4) the risk of overfitting and model instability and (5) the disregard for content expertise in model building. We advocate for the use of DAGs to guide variable selection for models that aim to examine associations between a putative risk factor and an outcome and emphasise the need for a more thoughtful and informed use of regression models in medical research.","PeriodicalId":501153,"journal":{"name":"Fetal & Neonatal","volume":"32 1","pages":""},"PeriodicalIF":0.0000,"publicationDate":"2023-12-06","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"Fetal & Neonatal","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1136/archdischild-2023-326340","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}
引用次数: 0
Abstract
This paper critically examines ‘kitchen sink regression’, a practice characterised by the manual or automated selection of variables for a multivariable regression model based on p values or model-based information criteria. We highlight the pitfalls of this method, using examples from perinatal/neonatal medicine, and propose more robust alternatives. The concept of directed acyclic graphs (DAGs) is introduced as a tool for describing and analysing causal relationships. We highlight five key issues with ‘kitchen sink regression’: (1) the disregard for the directionality of variable relationships, (2) the lack of a meaningful causal interpretation of effect estimates from these models, (3) the inflated alpha error rate due to multiple testing, (4) the risk of overfitting and model instability and (5) the disregard for content expertise in model building. We advocate for the use of DAGs to guide variable selection for models that aim to examine associations between a putative risk factor and an outcome and emphasise the need for a more thoughtful and informed use of regression models in medical research.