Revealing Patient Dissatisfaction With Health Care Resource Allocation in Multiple Dimensions Using Large Language Models and the International Classification of Diseases 11th Revision: Aspect-Based Sentiment Analysis.

IF 8.2 2区 医学 Q1 HEALTH CARE SCIENCES & SERVICES Journal of Medical Internet Research Pub Date : 2025-03-17 DOI:10.2196/66344
Jiaxuan Li, Yunchu Yang, Chao Mao, Patrick Cheong-Iao Pang, Quanjing Zhu, Dejian Xu, Yapeng Wang
{"title":"Revealing Patient Dissatisfaction With Health Care Resource Allocation in Multiple Dimensions Using Large Language Models and the International Classification of Diseases 11th Revision: Aspect-Based Sentiment Analysis.","authors":"Jiaxuan Li, Yunchu Yang, Chao Mao, Patrick Cheong-Iao Pang, Quanjing Zhu, Dejian Xu, Yapeng Wang","doi":"10.2196/66344","DOIUrl":null,"url":null,"abstract":"<p><strong>Background: </strong>Accurately measuring the health care needs of patients with different diseases remains a public health challenge for health care management worldwide. There is a need for new computational methods to be able to assess the health care resources required by patients with different diseases to avoid wasting resources.</p><p><strong>Objective: </strong>This study aimed to assessing dissatisfaction with allocation of health care resources from the perspective of patients with different diseases that can help optimize resource allocation and better achieve several of the Sustainable Development Goals (SDGs), such as SDG 3 (\"Good Health and Well-being\"). Our goal was to show the effectiveness and practicality of large language models (LLMs) in assessing the distribution of health care resources.</p><p><strong>Methods: </strong>We used aspect-based sentiment analysis (ABSA), which can divide textual data into several aspects for sentiment analysis. In this study, we used Chat Generative Pretrained Transformer (ChatGPT) to perform ABSA of patient reviews based on 3 aspects (patient experience, physician skills and efficiency, and infrastructure and administration)00 in which we embedded chain-of-thought (CoT) prompting and compared the performance of Chinese and English LLMs on a Chinese dataset. Additionally, we used the International Classification of Diseases 11th Revision (ICD-11) application programming interface (API) to classify the sentiment analysis results into different disease categories.</p><p><strong>Results: </strong>We evaluated the performance of the models by comparing predicted sentiments (either positive or negative) with the labels judged by human evaluators in terms of the aforementioned 3 aspects. The results showed that ChatGPT 3.5 is superior in a combination of stability, expense, and runtime considerations compared to ChatGPT-4o and Qwen-7b. The weighted total precision of our method based on the ABSA of patient reviews was 0.907, while the average accuracy of all 3 sampling methods was 0.893. Both values suggested that the model was able to achieve our objective. Using our approach, we identified that dissatisfaction is highest for sex-related diseases and lowest for circulatory diseases and that the need for better infrastructure and administration is much higher for blood-related diseases than for other diseases in China.</p><p><strong>Conclusions: </strong>The results prove that our method with LLMs can use patient reviews and the ICD-11 classification to assess the health care needs of patients with different diseases, which can assist with resource allocation rationally.</p>","PeriodicalId":16337,"journal":{"name":"Journal of Medical Internet Research","volume":"27 ","pages":"e66344"},"PeriodicalIF":8.2000,"publicationDate":"2025-03-17","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC11959199/pdf/","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"Journal of Medical Internet Research","FirstCategoryId":"3","ListUrlMain":"https://doi.org/10.2196/66344","RegionNum":2,"RegionCategory":"医学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"Q1","JCRName":"HEALTH CARE SCIENCES & SERVICES","Score":null,"Total":0}
引用次数: 0

Abstract

Background: Accurately measuring the health care needs of patients with different diseases remains a public health challenge for health care management worldwide. There is a need for new computational methods to be able to assess the health care resources required by patients with different diseases to avoid wasting resources.

Objective: This study aimed to assessing dissatisfaction with allocation of health care resources from the perspective of patients with different diseases that can help optimize resource allocation and better achieve several of the Sustainable Development Goals (SDGs), such as SDG 3 ("Good Health and Well-being"). Our goal was to show the effectiveness and practicality of large language models (LLMs) in assessing the distribution of health care resources.

Methods: We used aspect-based sentiment analysis (ABSA), which can divide textual data into several aspects for sentiment analysis. In this study, we used Chat Generative Pretrained Transformer (ChatGPT) to perform ABSA of patient reviews based on 3 aspects (patient experience, physician skills and efficiency, and infrastructure and administration)00 in which we embedded chain-of-thought (CoT) prompting and compared the performance of Chinese and English LLMs on a Chinese dataset. Additionally, we used the International Classification of Diseases 11th Revision (ICD-11) application programming interface (API) to classify the sentiment analysis results into different disease categories.

Results: We evaluated the performance of the models by comparing predicted sentiments (either positive or negative) with the labels judged by human evaluators in terms of the aforementioned 3 aspects. The results showed that ChatGPT 3.5 is superior in a combination of stability, expense, and runtime considerations compared to ChatGPT-4o and Qwen-7b. The weighted total precision of our method based on the ABSA of patient reviews was 0.907, while the average accuracy of all 3 sampling methods was 0.893. Both values suggested that the model was able to achieve our objective. Using our approach, we identified that dissatisfaction is highest for sex-related diseases and lowest for circulatory diseases and that the need for better infrastructure and administration is much higher for blood-related diseases than for other diseases in China.

Conclusions: The results prove that our method with LLMs can use patient reviews and the ICD-11 classification to assess the health care needs of patients with different diseases, which can assist with resource allocation rationally.

Abstract Image

Abstract Image

Abstract Image

查看原文
分享 分享
微信好友 朋友圈 QQ好友 复制链接
本刊更多论文
基于大语言模型和《国际疾病分类》第11版:基于方面的情感分析在多维度上揭示患者对医疗资源分配的不满。
背景:准确测量不同疾病患者的卫生保健需求仍然是全球卫生保健管理面临的公共卫生挑战。需要新的计算方法来评估不同疾病患者所需的卫生保健资源,以避免资源浪费。目的:本研究旨在从不同疾病患者的角度评估对卫生保健资源分配的不满,以帮助优化资源分配,更好地实现可持续发展目标(SDG),如SDG 3(“良好健康和福祉”)。我们的目标是展示大型语言模型(llm)在评估医疗资源分布方面的有效性和实用性。方法:采用基于方面的情感分析(ABSA)方法,将文本数据分成几个方面进行情感分析。在这项研究中,我们使用聊天生成预训练转换器(ChatGPT)基于3个方面(患者经验、医生技能和效率、基础设施和管理)00对患者评论进行ABSA,其中我们嵌入了思维链(CoT)提示,并比较了中文和英文法学硕士在中文数据集上的表现。此外,我们使用国际疾病分类第11版(ICD-11)应用程序编程接口(API)将情感分析结果划分为不同的疾病类别。结果:我们通过将预测的情绪(积极或消极)与人类评估者根据上述三个方面判断的标签进行比较来评估模型的性能。结果表明,与ChatGPT- 40和Qwen-7b相比,ChatGPT 3.5在稳定性、费用和运行时间方面都优于ChatGPT- 40。基于患者评价ABSA的方法加权总精密度为0.907,3种抽样方法的平均精密度为0.893。这两个值都表明该模型能够实现我们的目标。通过我们的方法,我们发现,在中国,与性有关的疾病的满意度最高,而与循环系统疾病的满意度最低,并且与其他疾病相比,血液相关疾病对更好的基础设施和管理的需求要高得多。结论:结果证明我们的llm方法可以利用患者回顾和ICD-11分类来评估不同疾病患者的医疗保健需求,有助于资源的合理配置。
本文章由计算机程序翻译,如有差异,请以英文原文为准。
求助全文
约1分钟内获得全文 去求助
来源期刊
CiteScore
14.40
自引率
5.40%
发文量
654
审稿时长
1 months
期刊介绍: The Journal of Medical Internet Research (JMIR) is a highly respected publication in the field of health informatics and health services. With a founding date in 1999, JMIR has been a pioneer in the field for over two decades. As a leader in the industry, the journal focuses on digital health, data science, health informatics, and emerging technologies for health, medicine, and biomedical research. It is recognized as a top publication in these disciplines, ranking in the first quartile (Q1) by Impact Factor. Notably, JMIR holds the prestigious position of being ranked #1 on Google Scholar within the "Medical Informatics" discipline.
期刊最新文献
Implementing Supported Digital Enhanced Cognitive Behavior Therapy for Binge Eating Disorder in Routine Care: Mixed Methods Service Evaluation. How Pandemics Have Reshaped the Respiratory Virus Data Landscape in Europe: Scoping Review. AI Models Could Improve Diagnosis and Care for Rare Diseases. Describing a National Chatbot Deployed by the Ministry of Health in Malawi During the COVID-19 Pandemic: Retrospective Data Analysis. Development and Clinical Evaluation of a Large Language Model-Based System for Generating Patient-Friendly Echocardiography Reports: Two-Stage Retrospective Validation and Prospective Survey Study.
×
引用
GB/T 7714-2015
复制
MLA
复制
APA
复制
导出至
BibTeX EndNote RefMan NoteFirst NoteExpress
×
×
提示
您的信息不完整,为了账户安全,请先补充。
现在去补充
×
提示
您因"违规操作"
具体请查看互助需知
我知道了
×
提示
现在去查看 取消
×
提示
确定
0
微信
客服QQ
Book学术公众号 扫码关注我们
反馈
×
意见反馈
请填写您的意见或建议
请填写您的手机或邮箱
已复制链接
已复制链接
快去分享给好友吧!
我知道了
×
扫码分享
扫码分享
Book学术官方微信
Book学术官方微信
Book学术文献互助
Book学术文献互助群
群 号:604180095
Book学术
文献互助 智能选刊 最新文献 互助须知 联系我们:info@booksci.cn
Book学术提供免费学术资源搜索服务,方便国内外学者检索中英文文献。致力于提供最便捷和优质的服务体验。
Copyright © 2023 Book学术 All rights reserved.
ghs 京公网安备 11010802042870号 京ICP备2023020795号-1