metLinkR: Facilitating Metaanalysis of Human Metabolomics Data through Automated Linking of Metabolite Identifiers.

IF 3.8 2区 生物学 Q1 BIOCHEMICAL RESEARCH METHODS Journal of Proteome Research Pub Date : 2025-05-02 Epub Date: 2025-04-04 DOI:10.1021/acs.jproteome.4c01051
Andrew Patt, Iris Pang, Fred Lee, Chiraag Gohel, Eoin Fahy, Vicki Stevens, David Ruggieri, Steven C Moore, Ewy A Mathé
{"title":"metLinkR: Facilitating Metaanalysis of Human Metabolomics Data through Automated Linking of Metabolite Identifiers.","authors":"Andrew Patt, Iris Pang, Fred Lee, Chiraag Gohel, Eoin Fahy, Vicki Stevens, David Ruggieri, Steven C Moore, Ewy A Mathé","doi":"10.1021/acs.jproteome.4c01051","DOIUrl":null,"url":null,"abstract":"<p><p>Metabolites are referenced in spectral, structural and pathway databases with a diverse array of schemas, including various internal database identifiers and large tables of common name synonyms. Cross-linking metabolite identifiers is a required step for meta-analysis of metabolomic results across studies but made difficult due to the lack of a consensus identifier system. We have implemented metLinkR, an R package that leverages RefMet and RaMP-DB to automate and simplify cross-linking metabolite identifiers across studies and generating common names. MetLinkR accepts as input metabolite common names and identifiers from five different databases (HMDB, KEGG, ChEBI, LIPIDMAPS and PubChem) to exhaustively search for possible overlap in supplied metabolites from input data sets. In an example of 13 metabolomic data sets totaling 10,400 metabolites, metLinkR identified and provided common names for 1377 metabolites in common between at least 2 data sets in less than 18 min and produced standardized names for 74.4% of the input metabolites. In another example comprising five data sets with 3512 metabolites, metLinkR identified 715 metabolites in common between at least two data sets in under 12 min and produced standardized names for 82.3% of the input metabolites. Outputs of MetLInR include output tables and metrics allowing users to readily double check the mappings and to get an overview of chemical classes represented. Overall, MetLinkR provides a streamlined solution for a common task in metabolomic epidemiology and other fields that meta-analyze metabolomic data. The R package, vignette and source code are freely downloadable at https://github.com/ncats/metLinkR.</p>","PeriodicalId":48,"journal":{"name":"Journal of Proteome Research","volume":" ","pages":"2403-2407"},"PeriodicalIF":3.8000,"publicationDate":"2025-05-02","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC12053952/pdf/","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"Journal of Proteome Research","FirstCategoryId":"99","ListUrlMain":"https://doi.org/10.1021/acs.jproteome.4c01051","RegionNum":2,"RegionCategory":"生物学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"2025/4/4 0:00:00","PubModel":"Epub","JCR":"Q1","JCRName":"BIOCHEMICAL RESEARCH METHODS","Score":null,"Total":0}
引用次数: 0

Abstract

Metabolites are referenced in spectral, structural and pathway databases with a diverse array of schemas, including various internal database identifiers and large tables of common name synonyms. Cross-linking metabolite identifiers is a required step for meta-analysis of metabolomic results across studies but made difficult due to the lack of a consensus identifier system. We have implemented metLinkR, an R package that leverages RefMet and RaMP-DB to automate and simplify cross-linking metabolite identifiers across studies and generating common names. MetLinkR accepts as input metabolite common names and identifiers from five different databases (HMDB, KEGG, ChEBI, LIPIDMAPS and PubChem) to exhaustively search for possible overlap in supplied metabolites from input data sets. In an example of 13 metabolomic data sets totaling 10,400 metabolites, metLinkR identified and provided common names for 1377 metabolites in common between at least 2 data sets in less than 18 min and produced standardized names for 74.4% of the input metabolites. In another example comprising five data sets with 3512 metabolites, metLinkR identified 715 metabolites in common between at least two data sets in under 12 min and produced standardized names for 82.3% of the input metabolites. Outputs of MetLInR include output tables and metrics allowing users to readily double check the mappings and to get an overview of chemical classes represented. Overall, MetLinkR provides a streamlined solution for a common task in metabolomic epidemiology and other fields that meta-analyze metabolomic data. The R package, vignette and source code are freely downloadable at https://github.com/ncats/metLinkR.

Abstract Image

Abstract Image

查看原文
分享 分享
微信好友 朋友圈 QQ好友 复制链接
本刊更多论文
metLinkR:通过代谢物标识符的自动链接促进人类代谢组学数据的元分析。
代谢物在光谱、结构和通路数据库中以多种模式被引用,包括各种内部数据库标识符和常用名称同义词的大型表。交联代谢物标识符是跨研究代谢组学结果荟萃分析的必要步骤,但由于缺乏共识的标识符系统而变得困难。我们已经实现了metLinkR,这是一个R包,利用RefMet和RaMP-DB来自动化和简化跨研究的交联代谢物标识符,并生成通用名称。MetLinkR接受来自五个不同数据库(HMDB, KEGG, ChEBI, LIPIDMAPS和PubChem)的代谢物通用名称和标识符作为输入,以从输入数据集中穷尽地搜索所提供代谢物的可能重叠。在13个代谢组学数据集共计10,400个代谢物的示例中,metLinkR在不到18分钟的时间内识别并提供了至少2个数据集之间共有的1377个代谢物的通用名称,并为74.4%的输入代谢物产生了标准化名称。在另一个包含包含3512种代谢物的5个数据集的例子中,metLinkR在12分钟内识别出至少两个数据集之间共有的715种代谢物,并为82.3%的输入代谢物产生了标准化名称。MetLInR的输出包括输出表和度量,允许用户轻松地再次检查映射并获得所表示的化学类的概述。总的来说,MetLinkR为代谢组学流行病学和其他领域的代谢组学数据荟萃分析提供了一个简化的解决方案。R包、小插图和源代码可以在https://github.com/ncats/metLinkR上免费下载。
本文章由计算机程序翻译,如有差异,请以英文原文为准。
求助全文
约1分钟内获得全文 去求助
来源期刊
Journal of Proteome Research
Journal of Proteome Research 生物-生化研究方法
CiteScore
9.00
自引率
4.50%
发文量
251
审稿时长
3 months
期刊介绍: Journal of Proteome Research publishes content encompassing all aspects of global protein analysis and function, including the dynamic aspects of genomics, spatio-temporal proteomics, metabonomics and metabolomics, clinical and agricultural proteomics, as well as advances in methodology including bioinformatics. The theme and emphasis is on a multidisciplinary approach to the life sciences through the synergy between the different types of "omics".
期刊最新文献
Proteomic Adaptations in the Spinal Cord of a Breast Cancer Model of Paclitaxel-Induced Peripheral Neuropathy. SoftHybrid: A Hybrid Imputation Algorithm Optimized for Single-Cell Proteomics Data. MSstatsQC-ML: A Supervised Machine Learning Approach to Monitor System Suitability and Quality Control in Mass Spectrometry-Based Proteomics. Collection on Forensic Proteomics. Improved Protein Identification in Shotgun Proteomics with a Group-Level Extension of the LPGF Model.
×
引用
GB/T 7714-2015
复制
MLA
复制
APA
复制
导出至
BibTeX EndNote RefMan NoteFirst NoteExpress
×
×
提示
您的信息不完整,为了账户安全,请先补充。
现在去补充
×
提示
您因"违规操作"
具体请查看互助需知
我知道了
×
提示
现在去查看 取消
×
提示
确定
0
微信
客服QQ
Book学术公众号 扫码关注我们
反馈
×
意见反馈
请填写您的意见或建议
请填写您的手机或邮箱
已复制链接
已复制链接
快去分享给好友吧!
我知道了
×
扫码分享
扫码分享
Book学术官方微信
Book学术官方微信
Book学术文献互助
Book学术文献互助群
群 号:604180095
Book学术
文献互助 智能选刊 最新文献 互助须知 联系我们:info@booksci.cn
Book学术提供免费学术资源搜索服务,方便国内外学者检索中英文文献。致力于提供最便捷和优质的服务体验。
Copyright © 2023 Book学术 All rights reserved.
ghs 京公网安备 11010802042870号 京ICP备2023020795号-1