Enhancing the interoperability and transparency of real-world data extraction in clinical research: evaluating the feasibility and impact of a ChatGLM implementation in Chinese hospital settings.
Bin Wang, Junkai Lai, Han Cao, Feifei Jin, Qiang Li, Mingkun Tang, Chen Yao, Ping Zhang
{"title":"Enhancing the interoperability and transparency of real-world data extraction in clinical research: evaluating the feasibility and impact of a ChatGLM implementation in Chinese hospital settings.","authors":"Bin Wang, Junkai Lai, Han Cao, Feifei Jin, Qiang Li, Mingkun Tang, Chen Yao, Ping Zhang","doi":"10.1093/ehjdh/ztae066","DOIUrl":null,"url":null,"abstract":"<p><strong>Aims: </strong>This study aims to assess the feasibility and impact of the implementation of the ChatGLM for real-world data (RWD) extraction in hospital settings. The primary focus of this research is on the effectiveness of ChatGLM-driven data extraction compared with that of manual processes associated with the electronic source data repository (ESDR) system.</p><p><strong>Methods and results: </strong>The researchers developed the ESDR system, which integrates ChatGLM, electronic case report forms (eCRFs), and electronic health records. The LLaMA (Large Language Model Meta AI) model was also deployed to compare the extraction accuracy of ChatGLM in free-text forms. A single-centre retrospective cohort study served as a pilot case. Five eCRF forms of 63 subjects, including free-text forms and discharge medication, were evaluated. Data collection involved electronic medical and prescription records collected from 13 departments. The ChatGLM-assisted process was associated with an estimated efficiency improvement of 80.7% in the eCRF data transcription time. The initial manual input accuracy for free-text forms was 99.59%, the ChatGLM data extraction accuracy was 77.13%, and the LLaMA data extraction accuracy was 43.86%. The challenges associated with the use of ChatGLM focus on prompt design, prompt output consistency, prompt output verification, and integration with hospital information systems.</p><p><strong>Conclusion: </strong>The main contribution of this study is to validate the use of ESDR tools to address the interoperability and transparency challenges of using ChatGLM for RWD extraction in Chinese hospital settings.</p>","PeriodicalId":72965,"journal":{"name":"European heart journal. Digital health","volume":"5 6","pages":"712-724"},"PeriodicalIF":3.9000,"publicationDate":"2024-09-12","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC11570364/pdf/","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"European heart journal. Digital health","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1093/ehjdh/ztae066","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"2024/11/1 0:00:00","PubModel":"eCollection","JCR":"Q1","JCRName":"CARDIAC & CARDIOVASCULAR SYSTEMS","Score":null,"Total":0}
引用次数: 0
Abstract
Aims: This study aims to assess the feasibility and impact of the implementation of the ChatGLM for real-world data (RWD) extraction in hospital settings. The primary focus of this research is on the effectiveness of ChatGLM-driven data extraction compared with that of manual processes associated with the electronic source data repository (ESDR) system.
Methods and results: The researchers developed the ESDR system, which integrates ChatGLM, electronic case report forms (eCRFs), and electronic health records. The LLaMA (Large Language Model Meta AI) model was also deployed to compare the extraction accuracy of ChatGLM in free-text forms. A single-centre retrospective cohort study served as a pilot case. Five eCRF forms of 63 subjects, including free-text forms and discharge medication, were evaluated. Data collection involved electronic medical and prescription records collected from 13 departments. The ChatGLM-assisted process was associated with an estimated efficiency improvement of 80.7% in the eCRF data transcription time. The initial manual input accuracy for free-text forms was 99.59%, the ChatGLM data extraction accuracy was 77.13%, and the LLaMA data extraction accuracy was 43.86%. The challenges associated with the use of ChatGLM focus on prompt design, prompt output consistency, prompt output verification, and integration with hospital information systems.
Conclusion: The main contribution of this study is to validate the use of ESDR tools to address the interoperability and transparency challenges of using ChatGLM for RWD extraction in Chinese hospital settings.