通过在相似性度量中结合结构化变量标签关系来改善患者聚类。

IF 3.4 3区 医学 Q1 HEALTH CARE SCIENCES & SERVICES BMC Medical Research Methodology Pub Date : 2025-03-15 DOI:10.1186/s12874-025-02459-8
Judith Lambert, Anne-Louise Leutenegger, Anaïs Baudot, Anne-Sophie Jannot
{"title":"通过在相似性度量中结合结构化变量标签关系来改善患者聚类。","authors":"Judith Lambert, Anne-Louise Leutenegger, Anaïs Baudot, Anne-Sophie Jannot","doi":"10.1186/s12874-025-02459-8","DOIUrl":null,"url":null,"abstract":"<p><strong>Background: </strong>Patient stratification is the cornerstone of numerous health investigations, serving to enhance the estimation of treatment efficacy and facilitating patient matching. To stratify patients, similarity measures between patients can be computed from clinical variables contained in medical health records. These variables have both values and labels structured in ontologies or other classification systems. The relevance of considering variable label relationships in the computation of patient similarity measures has been poorly studied.</p><p><strong>Objective: </strong>We adapt and evaluate several weighted versions of the Cosine similarity in order to consider structured label relationships to compute patient similarities from a medico-administrative database.</p><p><strong>Materials and methods: </strong>As a use case, we clustered patients aged 60 years from their annual medicine reimbursements contained in the Échantillon Généraliste des Bénéficiaires, a random sample of a French medico-administrative database. We used four patient similarity measures: the standard Cosine similarity, a weighted Cosine similarity measure that includes variable frequencies and two weighted Cosine similarity measures that consider variable label relationships. We construct patient networks from each similarity measure and identify clusters of patients using the Markov Cluster algorithm. We evaluate the performance of the different similarity measures with enrichment tests based on patient diagnoses.</p><p><strong>Results: </strong>The weighted similarity measures that include structured variable label relationships perform better to identify similar patients. Indeed, using these weighted measures, we identify more clusters associated with different diagnose enrichment. Importantly, the enrichment tests provide clinically interpretable insights into these patient clusters.</p><p><strong>Conclusion: </strong>Considering label relationships when computing patient similarities improves stratification of patients regarding their health status.</p>","PeriodicalId":9114,"journal":{"name":"BMC Medical Research Methodology","volume":"25 1","pages":"72"},"PeriodicalIF":3.4000,"publicationDate":"2025-03-15","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC11910865/pdf/","citationCount":"0","resultStr":"{\"title\":\"Improving patient clustering by incorporating structured variable label relationships in similarity measures.\",\"authors\":\"Judith Lambert, Anne-Louise Leutenegger, Anaïs Baudot, Anne-Sophie Jannot\",\"doi\":\"10.1186/s12874-025-02459-8\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"<p><strong>Background: </strong>Patient stratification is the cornerstone of numerous health investigations, serving to enhance the estimation of treatment efficacy and facilitating patient matching. To stratify patients, similarity measures between patients can be computed from clinical variables contained in medical health records. These variables have both values and labels structured in ontologies or other classification systems. The relevance of considering variable label relationships in the computation of patient similarity measures has been poorly studied.</p><p><strong>Objective: </strong>We adapt and evaluate several weighted versions of the Cosine similarity in order to consider structured label relationships to compute patient similarities from a medico-administrative database.</p><p><strong>Materials and methods: </strong>As a use case, we clustered patients aged 60 years from their annual medicine reimbursements contained in the Échantillon Généraliste des Bénéficiaires, a random sample of a French medico-administrative database. We used four patient similarity measures: the standard Cosine similarity, a weighted Cosine similarity measure that includes variable frequencies and two weighted Cosine similarity measures that consider variable label relationships. We construct patient networks from each similarity measure and identify clusters of patients using the Markov Cluster algorithm. We evaluate the performance of the different similarity measures with enrichment tests based on patient diagnoses.</p><p><strong>Results: </strong>The weighted similarity measures that include structured variable label relationships perform better to identify similar patients. Indeed, using these weighted measures, we identify more clusters associated with different diagnose enrichment. Importantly, the enrichment tests provide clinically interpretable insights into these patient clusters.</p><p><strong>Conclusion: </strong>Considering label relationships when computing patient similarities improves stratification of patients regarding their health status.</p>\",\"PeriodicalId\":9114,\"journal\":{\"name\":\"BMC Medical Research Methodology\",\"volume\":\"25 1\",\"pages\":\"72\"},\"PeriodicalIF\":3.4000,\"publicationDate\":\"2025-03-15\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC11910865/pdf/\",\"citationCount\":\"0\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"BMC Medical Research Methodology\",\"FirstCategoryId\":\"3\",\"ListUrlMain\":\"https://doi.org/10.1186/s12874-025-02459-8\",\"RegionNum\":3,\"RegionCategory\":\"医学\",\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"Q1\",\"JCRName\":\"HEALTH CARE SCIENCES & SERVICES\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"BMC Medical Research Methodology","FirstCategoryId":"3","ListUrlMain":"https://doi.org/10.1186/s12874-025-02459-8","RegionNum":3,"RegionCategory":"医学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"Q1","JCRName":"HEALTH CARE SCIENCES & SERVICES","Score":null,"Total":0}
引用次数: 0

摘要

背景:患者分层是许多健康调查的基础,有助于提高治疗效果的估计和促进患者匹配。为了对患者进行分层,可以从医疗健康记录中包含的临床变量计算患者之间的相似性度量。这些变量具有在本体或其他分类系统中结构化的值和标签。在计算患者相似度时考虑变量标签关系的相关性研究很少。目的:我们调整和评估余弦相似度的几个加权版本,以便考虑结构化标签关系来计算来自医疗管理数据库的患者相似度。材料和方法:作为一个用例,我们从法国医疗管理数据库的随机样本Échantillon gsamnastaliste des bsamnsamiciciaires中收集了60岁的年度医疗报销患者。我们使用了四种患者相似度度量:标准余弦相似度,包括可变频率的加权余弦相似度和考虑可变标签关系的两个加权余弦相似度度量。我们从每个相似度量构建患者网络,并使用马尔可夫聚类算法识别患者簇。我们评估了不同的相似性措施的性能与富集试验基于患者的诊断。结果:包含结构化变量标签关系的加权相似度度量在识别相似患者方面表现更好。事实上,使用这些加权措施,我们确定了更多与不同诊断富集相关的集群。重要的是,富集试验为这些患者群提供了临床可解释的见解。结论:在计算患者相似度时考虑标签关系可以改善患者健康状况的分层。
本文章由计算机程序翻译,如有差异,请以英文原文为准。

摘要图片

摘要图片

摘要图片

查看原文
分享 分享
微信好友 朋友圈 QQ好友 复制链接
本刊更多论文
Improving patient clustering by incorporating structured variable label relationships in similarity measures.

Background: Patient stratification is the cornerstone of numerous health investigations, serving to enhance the estimation of treatment efficacy and facilitating patient matching. To stratify patients, similarity measures between patients can be computed from clinical variables contained in medical health records. These variables have both values and labels structured in ontologies or other classification systems. The relevance of considering variable label relationships in the computation of patient similarity measures has been poorly studied.

Objective: We adapt and evaluate several weighted versions of the Cosine similarity in order to consider structured label relationships to compute patient similarities from a medico-administrative database.

Materials and methods: As a use case, we clustered patients aged 60 years from their annual medicine reimbursements contained in the Échantillon Généraliste des Bénéficiaires, a random sample of a French medico-administrative database. We used four patient similarity measures: the standard Cosine similarity, a weighted Cosine similarity measure that includes variable frequencies and two weighted Cosine similarity measures that consider variable label relationships. We construct patient networks from each similarity measure and identify clusters of patients using the Markov Cluster algorithm. We evaluate the performance of the different similarity measures with enrichment tests based on patient diagnoses.

Results: The weighted similarity measures that include structured variable label relationships perform better to identify similar patients. Indeed, using these weighted measures, we identify more clusters associated with different diagnose enrichment. Importantly, the enrichment tests provide clinically interpretable insights into these patient clusters.

Conclusion: Considering label relationships when computing patient similarities improves stratification of patients regarding their health status.

求助全文
通过发布文献求助,成功后即可免费获取论文全文。 去求助
来源期刊
BMC Medical Research Methodology
BMC Medical Research Methodology 医学-卫生保健
CiteScore
6.50
自引率
2.50%
发文量
298
审稿时长
3-8 weeks
期刊介绍: BMC Medical Research Methodology is an open access journal publishing original peer-reviewed research articles in methodological approaches to healthcare research. Articles on the methodology of epidemiological research, clinical trials and meta-analysis/systematic review are particularly encouraged, as are empirical studies of the associations between choice of methodology and study outcomes. BMC Medical Research Methodology does not aim to publish articles describing scientific methods or techniques: these should be directed to the BMC journal covering the relevant biomedical subject area.
期刊最新文献
Respondent fatigue in clinical practice: exploring the challenges of increasing PROMS and PREMS utilization in a tertiary hospital. From annotation to adaptation: extracting temporal relations in French clinical narratives. Modelling longitudinal and time-to-event data: a phase IV simulation study comparing R package implementations of joint models with time-varying Cox proportional-hazards regression, and the two-stage approach. Reducing selection bias risk to enhance RCT validity: sandwich mixed randomization outperforms permuted block design. Guidance for protocol content and reporting of dog-assisted interventions in randomised controlled trials: explanation and elaboration of the SPIRIT 2025 and CONSORT 2025 extensions.
×
引用
GB/T 7714-2015
复制
MLA
复制
APA
复制
导出至
BibTeX EndNote RefMan NoteFirst NoteExpress
×
×
提示
您的信息不完整,为了账户安全,请先补充。
现在去补充
×
提示
您因"违规操作"
具体请查看互助需知
我知道了
×
提示
现在去查看 取消
×
提示
确定
0
微信
客服QQ
Book学术公众号 扫码关注我们
反馈
×
意见反馈
请填写您的意见或建议
请填写您的手机或邮箱
已复制链接
已复制链接
快去分享给好友吧!
我知道了
×
扫码分享
扫码分享
Book学术官方微信
Book学术官方微信
Book学术文献互助
Book学术文献互助群
群 号:604180095
Book学术
文献互助 智能选刊 最新文献 互助须知 联系我们:info@booksci.cn
Book学术提供免费学术资源搜索服务,方便国内外学者检索中英文文献。致力于提供最便捷和优质的服务体验。
Copyright © 2023 Book学术 All rights reserved.
ghs 京公网安备 11010802042870号 京ICP备2023020795号-1