Generating barcodes for nanopore sequencing data with PRO

IF 6.2 3区 综合性期刊 Q1 Multidisciplinary Fundamental Research Pub Date : 2024-07-01 DOI:10.1016/j.fmre.2024.04.014
{"title":"Generating barcodes for nanopore sequencing data with PRO","authors":"","doi":"10.1016/j.fmre.2024.04.014","DOIUrl":null,"url":null,"abstract":"<div><p>DNA barcodes, short and unique DNA sequences, play a crucial role in sample identification when processing many samples simultaneously, which helps reduce experimental costs. Nevertheless, the low quality of long-read sequencing makes it difficult to identify barcodes accurately, which poses significant challenges for the design of barcodes for large numbers of samples in a single sequencing run. Here, we present a comprehensive study of the generation of barcodes and develop a tool, PRO, that can be used for selecting optimal barcode sets and demultiplexing. We formulate the barcode design problem as a combinatorial problem and prove that finding the optimal largest barcode set in a given DNA sequence space in which all sequences have the same length is theoretically NP-complete. For practical applications, we developed the novel method PRO by introducing the probability divergence between two DNA sequences to expand the capacity of barcode kits while ensuring demultiplexing accuracy. Specifically, the maximum size of the barcode kits designed by PRO is 2,292, which keeps the length of barcodes the same as that of the official ones used by Oxford Nanopore Technologies (ONT). We validated the performance of PRO on a simulated nanopore dataset with high error rates. The demultiplexing accuracy of PRO reached 98.29% for a barcode kit of size 2,922, 4.31% higher than that of Guppy, the official demultiplexing tool. When the size of the barcode kit generated by PRO is the same as the official size provided by ONT, both tools show superior and comparable demultiplexing accuracy.</p></div>","PeriodicalId":34602,"journal":{"name":"Fundamental Research","volume":null,"pages":null},"PeriodicalIF":6.2000,"publicationDate":"2024-07-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.sciencedirect.com/science/article/pii/S2667325824001924/pdfft?md5=ff23f335fd4d810b3e30973aca3b4f7d&pid=1-s2.0-S2667325824001924-main.pdf","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"Fundamental Research","FirstCategoryId":"1085","ListUrlMain":"https://www.sciencedirect.com/science/article/pii/S2667325824001924","RegionNum":3,"RegionCategory":"综合性期刊","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"Q1","JCRName":"Multidisciplinary","Score":null,"Total":0}
引用次数: 0

Abstract

DNA barcodes, short and unique DNA sequences, play a crucial role in sample identification when processing many samples simultaneously, which helps reduce experimental costs. Nevertheless, the low quality of long-read sequencing makes it difficult to identify barcodes accurately, which poses significant challenges for the design of barcodes for large numbers of samples in a single sequencing run. Here, we present a comprehensive study of the generation of barcodes and develop a tool, PRO, that can be used for selecting optimal barcode sets and demultiplexing. We formulate the barcode design problem as a combinatorial problem and prove that finding the optimal largest barcode set in a given DNA sequence space in which all sequences have the same length is theoretically NP-complete. For practical applications, we developed the novel method PRO by introducing the probability divergence between two DNA sequences to expand the capacity of barcode kits while ensuring demultiplexing accuracy. Specifically, the maximum size of the barcode kits designed by PRO is 2,292, which keeps the length of barcodes the same as that of the official ones used by Oxford Nanopore Technologies (ONT). We validated the performance of PRO on a simulated nanopore dataset with high error rates. The demultiplexing accuracy of PRO reached 98.29% for a barcode kit of size 2,922, 4.31% higher than that of Guppy, the official demultiplexing tool. When the size of the barcode kit generated by PRO is the same as the official size provided by ONT, both tools show superior and comparable demultiplexing accuracy.

查看原文
分享 分享
微信好友 朋友圈 QQ好友 复制链接
本刊更多论文
用 PRO 生成纳米孔测序数据的条形码
DNA 条形码是一种简短而独特的 DNA 序列,在同时处理许多样本时对样本识别起着至关重要的作用,有助于降低实验成本。然而,由于长读程测序的质量较低,很难准确识别条形码,这给在一次测序中为大量样本设计条形码带来了巨大挑战。在此,我们对条形码的生成进行了全面研究,并开发了一种可用于选择最佳条形码集和解复用的工具 PRO。我们将条形码设计问题表述为一个组合问题,并证明在所有序列长度相同的给定 DNA 序列空间中寻找最优的最大条形码集在理论上是 NP-完全的。在实际应用中,我们开发了新方法 PRO,通过引入两个 DNA 序列之间的概率发散,在确保解复用精度的同时扩大了条形码套件的容量。具体来说,PRO 所设计的条形码试剂盒的最大尺寸为 2292,与牛津纳米孔技术公司(ONT)正式使用的条形码长度相同。我们在误差率较高的模拟纳米孔数据集上验证了 PRO 的性能。在条形码大小为 2,922 的情况下,PRO 的解复用准确率达到 98.29%,比官方解复用工具 Guppy 高出 4.31%。当 PRO 生成的条形码工具包大小与 ONT 提供的官方大小相同时,两种工具的解复用准确率都很高,不相上下。
本文章由计算机程序翻译,如有差异,请以英文原文为准。
求助全文
约1分钟内获得全文 去求助
来源期刊
Fundamental Research
Fundamental Research Multidisciplinary-Multidisciplinary
CiteScore
4.00
自引率
1.60%
发文量
294
审稿时长
79 days
期刊介绍:
期刊最新文献
Structural rejuvenation of a well-aged metallic glass Reaction mechanism of toluene decomposition in non-thermal plasma: How does it compare with benzene? Event-triggered neuroadaptive predefined practical finite-time control for dynamic positioning vessels: A time-based generator approach Unsupervised learning of interacting topological phases from experimental observables Reactions with Criegee intermediates are the dominant gas-phase sink for formyl fluoride in the atmosphere
×
引用
GB/T 7714-2015
复制
MLA
复制
APA
复制
导出至
BibTeX EndNote RefMan NoteFirst NoteExpress
×
×
提示
您的信息不完整,为了账户安全,请先补充。
现在去补充
×
提示
您因"违规操作"
具体请查看互助需知
我知道了
×
提示
现在去查看 取消
×
提示
确定
0
微信
客服QQ
Book学术公众号 扫码关注我们
反馈
×
意见反馈
请填写您的意见或建议
请填写您的手机或邮箱
已复制链接
已复制链接
快去分享给好友吧!
我知道了
×
扫码分享
扫码分享
Book学术官方微信
Book学术文献互助
Book学术文献互助群
群 号:481959085
Book学术
文献互助 智能选刊 最新文献 互助须知 联系我们:info@booksci.cn
Book学术提供免费学术资源搜索服务,方便国内外学者检索中英文文献。致力于提供最便捷和优质的服务体验。
Copyright © 2023 Book学术 All rights reserved.
ghs 京公网安备 11010802042870号 京ICP备2023020795号-1