AI-assisted exposure-response data analysis: Quantifying heterogeneous causal effects of exposures on survival times

Louis Anthony Cox Jr. , R. Jeffrey Lewis , Saumitra V. Rege , Shubham Singh
{"title":"AI-assisted exposure-response data analysis: Quantifying heterogeneous causal effects of exposures on survival times","authors":"Louis Anthony Cox Jr. ,&nbsp;R. Jeffrey Lewis ,&nbsp;Saumitra V. Rege ,&nbsp;Shubham Singh","doi":"10.1016/j.gloepi.2024.100179","DOIUrl":null,"url":null,"abstract":"<div><div>AI-assisted data analysis can help risk analysts better understand exposure-response relationships by making it relatively easy to apply advanced statistical and machine learning methods, check their assumptions, and interpret their results. This paper demonstrates the potential of large language models (LLMs), such as ChatGPT, to facilitate statistical analyses, including survival data analyses, for health risk assessments. Through AI-guided analyses using relatively recent and advanced methods such as Individual Conditional Expectation (ICE) plots using Random Survival Forests and Heterogeneous Treatment Effects (HTEs) estimated using Causal Survival Forests, population-level exposure-response functions can be disaggregated into individual-level exposure-response functions. These reveal the extent of heterogeneity in risks across individuals for different levels of exposure, holding other variables fixed. By applying these methods to an illustrative dataset on blood lead levels (BLL) and mortality risk among never-smoker men from the NHANES III survey, we show how AI can clarify inter-individual variations in exposure-associated risks. The results add insights not easily obtained from traditional parametric or semi-parametric models such as logistic regression and Cox proportional hazards models, illustrating the advantages of non-parametric approaches for quantifying heterogeneous causal effects on survival times. This paper also suggests some practical implications of using AI in regulatory health risk assessments and public policy decisions.</div></div>","PeriodicalId":36311,"journal":{"name":"Global Epidemiology","volume":"9 ","pages":"Article 100179"},"PeriodicalIF":0.0000,"publicationDate":"2024-12-25","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC11757793/pdf/","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"Global Epidemiology","FirstCategoryId":"1085","ListUrlMain":"https://www.sciencedirect.com/science/article/pii/S2590113324000452","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}
引用次数: 0

Abstract

AI-assisted data analysis can help risk analysts better understand exposure-response relationships by making it relatively easy to apply advanced statistical and machine learning methods, check their assumptions, and interpret their results. This paper demonstrates the potential of large language models (LLMs), such as ChatGPT, to facilitate statistical analyses, including survival data analyses, for health risk assessments. Through AI-guided analyses using relatively recent and advanced methods such as Individual Conditional Expectation (ICE) plots using Random Survival Forests and Heterogeneous Treatment Effects (HTEs) estimated using Causal Survival Forests, population-level exposure-response functions can be disaggregated into individual-level exposure-response functions. These reveal the extent of heterogeneity in risks across individuals for different levels of exposure, holding other variables fixed. By applying these methods to an illustrative dataset on blood lead levels (BLL) and mortality risk among never-smoker men from the NHANES III survey, we show how AI can clarify inter-individual variations in exposure-associated risks. The results add insights not easily obtained from traditional parametric or semi-parametric models such as logistic regression and Cox proportional hazards models, illustrating the advantages of non-parametric approaches for quantifying heterogeneous causal effects on survival times. This paper also suggests some practical implications of using AI in regulatory health risk assessments and public policy decisions.
查看原文
分享 分享
微信好友 朋友圈 QQ好友 复制链接
本刊更多论文
人工智能辅助暴露-反应数据分析:量化暴露对生存时间的异质性因果效应。
人工智能辅助数据分析可以帮助风险分析师更好地理解暴露-反应关系,使其相对容易地应用先进的统计和机器学习方法,检查他们的假设,并解释他们的结果。本文展示了大型语言模型(llm)的潜力,例如ChatGPT,以促进统计分析,包括生存数据分析,用于健康风险评估。通过人工智能引导的分析,使用相对最新和先进的方法,如使用随机生存森林的个体条件期望(ICE)图和使用因果生存森林估计的异质处理效应(HTEs),可以将种群水平的暴露-反应函数分解为个体水平的暴露-反应函数。这些揭示了不同暴露水平的个体之间风险的异质性程度,保持其他变量不变。通过将这些方法应用于NHANES III调查中从不吸烟男性的血铅水平(BLL)和死亡风险的说明性数据集,我们展示了人工智能如何阐明暴露相关风险的个体间差异。结果增加了传统参数或半参数模型(如逻辑回归和Cox比例风险模型)不易获得的见解,说明了非参数方法在量化异质性因果效应对生存时间的优势。本文还提出了在监管卫生风险评估和公共政策决策中使用人工智能的一些实际意义。
本文章由计算机程序翻译,如有差异,请以英文原文为准。
求助全文
约1分钟内获得全文 去求助
来源期刊
Global Epidemiology
Global Epidemiology Medicine-Infectious Diseases
CiteScore
5.00
自引率
0.00%
发文量
22
审稿时长
39 days
期刊最新文献
Interpretation of the IARC quantitative bias analysis of talc and ovarian cancer Comment on “do certain blood groups increase COVID-19 severity and mortality?” Statistically significant results from low-power analyses: A comedy of errors Barriers and facilitators to healthcare access for migrants in Morocco: A narrative review A temporal network analysis of drug co-prescription during antidepressants and anxiolytics dispensing in the Netherlands from 2018 to 2022
×
引用
GB/T 7714-2015
复制
MLA
复制
APA
复制
导出至
BibTeX EndNote RefMan NoteFirst NoteExpress
×
×
提示
您的信息不完整,为了账户安全,请先补充。
现在去补充
×
提示
您因"违规操作"
具体请查看互助需知
我知道了
×
提示
现在去查看 取消
×
提示
确定
0
微信
客服QQ
Book学术公众号 扫码关注我们
反馈
×
意见反馈
请填写您的意见或建议
请填写您的手机或邮箱
已复制链接
已复制链接
快去分享给好友吧!
我知道了
×
扫码分享
扫码分享
Book学术官方微信
Book学术文献互助
Book学术文献互助群
群 号:604180095
Book学术
文献互助 智能选刊 最新文献 互助须知 联系我们:info@booksci.cn
Book学术提供免费学术资源搜索服务,方便国内外学者检索中英文文献。致力于提供最便捷和优质的服务体验。
Copyright © 2023 Book学术 All rights reserved.
ghs 京公网安备 11010802042870号 京ICP备2023020795号-1