Using Transfer Learning for Improved Mortality Prediction in a Data-Scarce Hospital Setting.

Biomedical informatics insights Pub Date : 2017-06-12 eCollection Date: 2017-01-01 DOI:10.1177/1178222617712994
Thomas Desautels, Jacob Calvert, Jana Hoffman, Qingqing Mao, Melissa Jay, Grant Fletcher, Chris Barton, Uli Chettipally, Yaniv Kerem, Ritankar Das
{"title":"Using Transfer Learning for Improved Mortality Prediction in a Data-Scarce Hospital Setting.","authors":"Thomas Desautels,&nbsp;Jacob Calvert,&nbsp;Jana Hoffman,&nbsp;Qingqing Mao,&nbsp;Melissa Jay,&nbsp;Grant Fletcher,&nbsp;Chris Barton,&nbsp;Uli Chettipally,&nbsp;Yaniv Kerem,&nbsp;Ritankar Das","doi":"10.1177/1178222617712994","DOIUrl":null,"url":null,"abstract":"<p><p>Algorithm-based clinical decision support (CDS) systems associate patient-derived health data with outcomes of interest, such as in-hospital mortality. However, the quality of such associations often depends on the availability of site-specific training data. Without sufficient quantities of data, the underlying statistical apparatus cannot differentiate useful patterns from noise and, as a result, may underperform. This initial training data burden limits the widespread, out-of-the-box, use of machine learning-based risk scoring systems. In this study, we implement a statistical transfer learning technique, which uses a large \"source\" data set to drastically reduce the amount of data needed to perform well on a \"target\" site for which training data are scarce. We test this transfer technique with <i>AutoTriage</i>, a mortality prediction algorithm, on patient charts from the Beth Israel Deaconess Medical Center (the source) and a population of 48 249 adult inpatients from University of California San Francisco Medical Center (the target institution). We find that the amount of training data required to surpass 0.80 area under the receiver operating characteristic (AUROC) on the target set decreases from more than 4000 patients to fewer than 220. This performance is superior to the Modified Early Warning Score (AUROC: 0.76) and corresponds to a decrease in clinical data collection time from approximately 6 months to less than 10 days. Our results highlight the usefulness of transfer learning in the specialization of CDS systems to new hospital sites, without requiring expensive and time-consuming data collection efforts.</p>","PeriodicalId":88397,"journal":{"name":"Biomedical informatics insights","volume":"9 ","pages":"1178222617712994"},"PeriodicalIF":0.0000,"publicationDate":"2017-06-12","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://sci-hub-pdf.com/10.1177/1178222617712994","citationCount":"42","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"Biomedical informatics insights","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1177/1178222617712994","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"2017/1/1 0:00:00","PubModel":"eCollection","JCR":"","JCRName":"","Score":null,"Total":0}
引用次数: 42

Abstract

Algorithm-based clinical decision support (CDS) systems associate patient-derived health data with outcomes of interest, such as in-hospital mortality. However, the quality of such associations often depends on the availability of site-specific training data. Without sufficient quantities of data, the underlying statistical apparatus cannot differentiate useful patterns from noise and, as a result, may underperform. This initial training data burden limits the widespread, out-of-the-box, use of machine learning-based risk scoring systems. In this study, we implement a statistical transfer learning technique, which uses a large "source" data set to drastically reduce the amount of data needed to perform well on a "target" site for which training data are scarce. We test this transfer technique with AutoTriage, a mortality prediction algorithm, on patient charts from the Beth Israel Deaconess Medical Center (the source) and a population of 48 249 adult inpatients from University of California San Francisco Medical Center (the target institution). We find that the amount of training data required to surpass 0.80 area under the receiver operating characteristic (AUROC) on the target set decreases from more than 4000 patients to fewer than 220. This performance is superior to the Modified Early Warning Score (AUROC: 0.76) and corresponds to a decrease in clinical data collection time from approximately 6 months to less than 10 days. Our results highlight the usefulness of transfer learning in the specialization of CDS systems to new hospital sites, without requiring expensive and time-consuming data collection efforts.

Abstract Image

Abstract Image

Abstract Image

查看原文
分享 分享
微信好友 朋友圈 QQ好友 复制链接
本刊更多论文
在数据稀缺的医院环境中使用迁移学习改进死亡率预测。
基于算法的临床决策支持(CDS)系统将患者衍生的健康数据与感兴趣的结果(如住院死亡率)关联起来。然而,这种联系的质量往往取决于具体地点培训数据的可得性。如果没有足够数量的数据,底层的统计仪器就无法区分有用的模式和噪声,结果可能表现不佳。这种初始的训练数据负担限制了基于机器学习的风险评分系统的广泛使用。在本研究中,我们实现了一种统计迁移学习技术,该技术使用大型“源”数据集来大幅减少在训练数据稀缺的“目标”站点上表现良好所需的数据量。我们使用AutoTriage(一种死亡率预测算法)对来自Beth Israel Deaconess医疗中心(来源)的患者图表和来自加州大学旧金山医疗中心(目标机构)的48249名成年住院患者进行了测试。我们发现,在目标集的接收者操作特征(AUROC)下,超过0.80面积所需的训练数据量从4000多名患者减少到不到220名患者。这一性能优于改良早期预警评分(AUROC: 0.76),并对应于临床数据收集时间从大约6个月减少到不到10天。我们的研究结果强调了转移学习在CDS系统专业化到新医院的有用性,而不需要昂贵和耗时的数据收集工作。
本文章由计算机程序翻译,如有差异,请以英文原文为准。
求助全文
约1分钟内获得全文 去求助
来源期刊
自引率
0.00%
发文量
0
期刊最新文献
A Data-Driven Approach to Predicting Septic Shock in the Intensive Care Unit A Genome Model to Explain Major Features of Neurodevelopmental Disorders in Newborns. Mathematical Model for Computer-Assisted Modification of Medication Dosing Rules. Applying Supervised Machine Learning to Identify Which Patient Characteristics Identify the Highest Rates of Mortality Post-Interhospital Transfer. Coalitional Game Theory Facilitates Identification of Non-Coding Variants Associated With Autism.
×
引用
GB/T 7714-2015
复制
MLA
复制
APA
复制
导出至
BibTeX EndNote RefMan NoteFirst NoteExpress
×
×
提示
您的信息不完整,为了账户安全,请先补充。
现在去补充
×
提示
您因"违规操作"
具体请查看互助需知
我知道了
×
提示
现在去查看 取消
×
提示
确定
0
微信
客服QQ
Book学术公众号 扫码关注我们
反馈
×
意见反馈
请填写您的意见或建议
请填写您的手机或邮箱
已复制链接
已复制链接
快去分享给好友吧!
我知道了
×
扫码分享
扫码分享
Book学术官方微信
Book学术文献互助
Book学术文献互助群
群 号:481959085
Book学术
文献互助 智能选刊 最新文献 互助须知 联系我们:info@booksci.cn
Book学术提供免费学术资源搜索服务,方便国内外学者检索中英文文献。致力于提供最便捷和优质的服务体验。
Copyright © 2023 Book学术 All rights reserved.
ghs 京公网安备 11010802042870号 京ICP备2023020795号-1