GigaByte (Hong Kong, China)最新文献

英文中文

Chromosomal-level genome assembly of the long-spined sea urchin Diadema setosum (Leske, 1778). 长刺海胆 Diadema setosum (Leske, 1778) 染色体级基因组组装。

GigaByte (Hong Kong, China)

Pub Date : 2024-04-25 eCollection Date: 2024-01-01 DOI: 10.46471/gigabyte.121

The long-spined sea urchin Diadema setosum is an algal and coral feeder widely distributed in the Indo-Pacific that can cause severe bioerosion on the reef community. However, the lack of genomic information has hindered the study of its ecology and evolution. Here, we report the chromosomal-level genome (885.8 Mb) of the long-spined sea urchin D. setosum using a combination of PacBio long-read sequencing and Omni-C scaffolding technology. The assembled genome contains a scaffold N50 length of 38.3 Mb, 98.1% of complete BUSCO (Geno, metazoa_odb10) genes (the single copy score is 97.8% and the duplication score is 0.3%), and 98.6% of the sequences are anchored to 22 pseudo-molecules/chromosomes. A total of 27,478 gene models have were annotated, reaching a total of 28,414 transcripts, including 5,384 tRNA and 23,030 protein-coding genes. The high-quality genome of D. setosum presented here is a valuable resource for the ecological and evolutionary studies of this coral reef-associated sea urchin.

长刺海胆（Diadema setosum）是一种广泛分布于印度洋-太平洋地区的藻类和珊瑚喂食者，可对珊瑚礁群落造成严重的生物侵蚀。然而，基因组信息的缺乏阻碍了对其生态学和进化的研究。在这里，我们利用 PacBio 长线程测序技术和 Omni-C 支架技术，报告了长棘海胆 D. setosum 的染色体级基因组（885.8 Mb）。组装的基因组包含 38.3 Mb 的支架 N50 长度，98.1% 的完整 BUSCO（Geno，metazoa_odb10）基因（单拷贝得分 97.8%，重复得分 0.3%），98.6% 的序列锚定在 22 个伪分子/染色体上。共注释了 27,478 个基因模型，共有 28,414 个转录本，包括 5,384 个 tRNA 和 23,030 个编码蛋白质的基因。这里展示的高质量 D. setosum 基因组是研究这种与珊瑚礁相关的海胆的生态和进化的宝贵资源。

引用次数: 0

Chromosome-level genome assembly of the common chiton, Liolophura japonica (Lischke, 1873). 普通甲壳动物 Liolophura japonica (Lischke, 1873) 染色体级基因组组装。

IF 1.2

GigaByte (Hong Kong, China)

Pub Date : 2024-04-25 eCollection Date: 2024-01-01 DOI: 10.46471/gigabyte.123

Chitons (Polyplacophora) are marine molluscs that can be found worldwide from cold waters to the tropics, and play important ecological roles in the environment. However, only two chiton genomes have been sequenced to date. The chiton Liolophura japonica (Lischke, 1873) is one of the most abundant polyplacophorans found throughout East Asia. Our PacBio HiFi reads and Omni-C sequencing data resulted in a high-quality near chromosome-level genome assembly of ∼609 Mb with a scaffold N50 length of 37.34 Mb (96.1% BUSCO). A total of 28,233 genes were predicted, including 28,010 protein-coding ones. The repeat content (27.89%) was similar to that of other Chitonidae species and approximately three times lower than that of the Hanleyidae chiton genome. The genomic resources provided by this work will help to expand our understanding of the evolution of molluscs and the ecological adaptation of chitons.

甲壳纲（Polyplacophora）是一种海洋软体动物，从寒冷的水域到热带地区都有分布，在环境中扮演着重要的生态角色。然而，迄今为止只有两个甲壳动物基因组被测序。Liolophura japonica（Lischke，1873 年）是在整个东亚发现的最丰富的多孔软体动物之一。我们的 PacBio HiFi 读数和 Omni-C 测序数据产生了一个 ∼609 Mb 的高质量近染色体级基因组，支架 N50 长度为 37.34 Mb（96.1% BUSCO）。共预测出 28,233 个基因，包括 28,010 个编码蛋白质的基因。重复含量（27.89%）与壳斗科其他物种相似，比汉雷科壳斗鱼基因组低约三倍。这项工作提供的基因组资源将有助于拓展我们对软体动物进化和甲壳动物生态适应性的认识。

引用次数: 0

Genome assembly of the edible jelly fungus Dacryopinax spathularia (Dacrymycetaceae). 食用果冻真菌 Dacryopinax spathularia（Dacrymycetaceae）的基因组组装。

GigaByte (Hong Kong, China)

Pub Date : 2024-04-25 eCollection Date: 2024-01-01 DOI: 10.46471/gigabyte.120

The edible jelly fungus Dacryopinax spathularia (Dacrymycetaceae) is wood-decaying and can be commonly found worldwide. It has found application in food additives, given its ability to synthesize long-chain glycolipids, among other uses. In this study, we present the genome assembly of D. spathularia using a combination of PacBio HiFi reads and Omni-C data. The genome size is 29.2 Mb. It has high sequence contiguity and completeness, with a scaffold N50 of 1.925 Mb and a 92.0% BUSCO score. A total of 11,510 protein-coding genes and 474.7 kb repeats (accounting for 1.62% of the genome) were predicted. The D. spathularia genome assembly generated in this study provides a valuable resource for understanding their ecology, such as their wood-decaying capability, their evolutionary relationships with other fungi, and their unique biology and applications in the food industry.

可食用的果冻真菌 Dacryopinax spathularia（Dacrymycetaceae）是一种木材腐生菌，在世界各地都能常见到。由于它具有合成长链糖脂的能力，因此在食品添加剂等方面也有应用。在这项研究中，我们结合使用 PacBio HiFi 读数和 Omni-C 数据，完成了 D. spathularia 的基因组组装。基因组大小为 29.2 Mb。它具有较高的序列连续性和完整性，支架 N50 为 1.925 Mb，BUSCO 得分为 92.0%。共预测出 11,510 个编码蛋白质的基因和 474.7 kb 的重复序列（占基因组的 1.62%）。本研究中生成的 D. spathularia 基因组组装为了解其生态学（如木材腐烂能力）、与其他真菌的进化关系以及其独特的生物学特性和在食品工业中的应用提供了宝贵的资源。

引用次数: 0

Genome assembly of the milky mangrove Excoecaria agallocha. 牛奶红树林 Excoecaria agallocha 的基因组组装。

GigaByte (Hong Kong, China)

Pub Date : 2024-04-25 eCollection Date: 2024-01-01 DOI: 10.46471/gigabyte.119

The milky mangrove Excoecaria agallocha is a latex-secreting mangrove that are distributed in tropical and subtropical regions. While its poisonous latex is regarded as a potential source of phytochemicals for biomedical applications, the genomic resources of E. agallocha remains limited. Here, we present a chromosomal level genome of E. agallocha, assembled from the combination of PacBio long-read sequencing and Omni-C data. The resulting assembly size is 1,332.45 Mb and has high contiguity and completeness with a scaffold N50 of 58.9 Mb and a BUSCO score of 98.4%, with 86.08% of sequences anchored to 18 pseudomolecules. 73,740 protein-coding genes were also predicted. The milky mangrove genome provides a useful resource for further understanding the biosynthesis of phytochemical compounds in E. agallocha.

乳汁红树林（Excoecaria agallocha）是一种分泌乳汁的红树林，分布于热带和亚热带地区。虽然其有毒的乳汁被认为是生物医学应用中植物化学物质的潜在来源，但 E. agallocha 的基因组资源仍然有限。在这里，我们展示了结合 PacBio 长线程测序和 Omni-C 数据组装的 E. agallocha 染色体级基因组。组装结果大小为 1,332.45 Mb，具有很高的连续性和完整性，支架 N50 为 58.9 Mb，BUSCO 得分为 98.4%，其中 86.08% 的序列锚定在 18 个假分子上。此外，还预测了 73,740 个编码蛋白质的基因。乳汁红树林基因组为进一步了解 E. agallocha 植物化学物质的生物合成提供了有用的资源。

引用次数: 0

Bridging Biodiversity and Health: The Global Biodiversity Information Facility's initiative on open data on vectors of human diseases. 连接生物多样性与健康：全球生物多样性信息基金关于人类疾病媒介开放数据的倡议。

GigaByte (Hong Kong, China)

Pub Date : 2024-04-11 eCollection Date: 2024-01-01 DOI: 10.46471/gigabyte.117

Paloma Shimabukuro, Quentin Groom, Florence Fouque, Lindsay Campbell, Theeraphap Chareonviriyaphap, Josiane Etang, Sylvie Manguin, Marianne Sinka, Dmitry Schigel, Kate Ingenloff

There is an increased awareness of the importance of data publication, data sharing, and open science to support research, monitoring and control of vector-borne disease (VBD). Here we describe the efforts of the Global Biodiversity Information Facility (GBIF) as well as the World Health Special Programme on Research and Training in Diseases of Poverty (TDR) to promote publication of data related to vectors of diseases. In 2020, a GBIF task group of experts was formed to provide advice and support efforts aimed at enhancing the coverage and accessibility of data on vectors of human diseases within GBIF. Various strategies, such as organizing training courses and publishing data papers, were used to increase this content. This editorial introduces the outcome of a second call for data papers partnered by the TDR, GBIF and GigaScience Press in the journal GigaByte. Biodiversity and infectious diseases are linked in complex ways. These links can involve changes from the microorganism level to that of the habitat, and there are many ways in which these factors interact to affect human health. One way to tackle disease control and possibly elimination, is to provide stakeholders with access to a wide range of data shared under the FAIR principles, so it is possible to support early detection, analyses and evaluation, and to promote policy improvements and/or development.

人们越来越意识到数据发布、数据共享和开放科学对于支持病媒生物疾病（VBD）的研究、监测和控制的重要性。在此，我们将介绍全球生物多样性信息基金（GBIF）以及世界贫困疾病研究和培训特别计划（TDR）为促进病媒相关数据的发布所做的努力。2020 年，成立了 GBIF 专家工作组，以提供建议和支持旨在加强 GBIF 内人类疾病病媒数据的覆盖面和可获取性的工作。为增加这方面的内容，采取了各种策略，如组织培训课程、发表数据论文等。这篇社论介绍了由TDR、GBIF和GigaScience出版社合作在《GigaByte》杂志上第二次征集数据论文的结果。生物多样性与传染性疾病之间有着复杂的联系。这些联系可能涉及从微生物层面到栖息地层面的变化，这些因素通过多种方式相互作用，影响人类健康。解决疾病控制和可能的消除问题的方法之一，是让利益相关者能够访问在 FAIR 原则下共享的各种数据，从而支持早期检测、分析和评估，并促进政策改进和/或发展。

{"title":"Bridging Biodiversity and Health: The Global Biodiversity Information Facility's initiative on open data on vectors of human diseases.","authors":"Paloma Shimabukuro, Quentin Groom, Florence Fouque, Lindsay Campbell, Theeraphap Chareonviriyaphap, Josiane Etang, Sylvie Manguin, Marianne Sinka, Dmitry Schigel, Kate Ingenloff","doi":"10.46471/gigabyte.117","DOIUrl":"10.46471/gigabyte.117","url":null,"abstract":"There is an increased awareness of the importance of data publication, data sharing, and open science to support research, monitoring and control of vector-borne disease (VBD). Here we describe the efforts of the Global Biodiversity Information Facility (GBIF) as well as the World Health Special Programme on Research and Training in Diseases of Poverty (TDR) to promote publication of data related to vectors of diseases. In 2020, a GBIF task group of experts was formed to provide advice and support efforts aimed at enhancing the coverage and accessibility of data on vectors of human diseases within GBIF. Various strategies, such as organizing training courses and publishing data papers, were used to increase this content. This editorial introduces the outcome of a second call for data papers partnered by the TDR, GBIF and GigaScience Press in the journal GigaByte. Biodiversity and infectious diseases are linked in complex ways. These links can involve changes from the microorganism level to that of the habitat, and there are many ways in which these factors interact to affect human health. One way to tackle disease control and possibly elimination, is to provide stakeholders with access to a wide range of data shared under the FAIR principles, so it is possible to support early detection, analyses and evaluation, and to promote policy improvements and/or development.","PeriodicalId":73157,"journal":{"name":"GigaByte (Hong Kong, China)","volume":"2024 ","pages":"gigabyte117"},"PeriodicalIF":0.0,"publicationDate":"2024-04-11","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC11027195/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"140860840","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":0,"RegionCategory":"","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}

引用次数: 0

Snakemake workflows for long-read bacterial genome assembly and evaluation. 用于长读数细菌基因组组装和评估的 Snakemake 工作流程。

GigaByte (Hong Kong, China)

Pub Date : 2024-04-01 eCollection Date: 2024-01-01 DOI: 10.46471/gigabyte.116

Peter Menzel

With the advancement of long-read sequencing technologies and their increasing use for bacterial genomics, several methods for generating genome assemblies from error-prone long reads have been developed. These are complemented by various tools for assembly polishing using either long reads, short reads, or reference genomes. End users are therefore left with a plethora of possible combinations of programs for obtaining a final trusted assembly. Hence, there is also a need to measure the completeness and accuracy of such assemblies, for which, again, several evaluation methods implemented in various programs are available. In order to automatically run multiple genome assembly and evaluation programs at once, I developed two workflows for the workflow management system Snakemake, which provide end users with an easy-to-run solution for testing various genome assemblies from their sequencing data. Both workflows use the conda packaging system, so there is no need for manual installation of each program.

Availability & implementation: The workflows are available as open source software under the MIT license at github.com/pmenzel/ont-assembly-snake and github.com/pmenzel/score-assemblies.

随着长读数测序技术的发展及其在细菌基因组学中的应用日益广泛，已经开发出了几种从容易出错的长读数中生成基因组装配的方法。此外，还有各种利用长读数、短读数或参考基因组进行组装抛光的工具。因此，最终用户只能通过大量可能的程序组合来获得最终可信的组装结果。因此，还需要对这些组装的完整性和准确性进行测量，为此，在各种程序中也提供了多种评估方法。为了一次自动运行多个基因组组装和评估程序，我为工作流管理系统 Snakemake 开发了两个工作流，为终端用户提供了一个易于运行的解决方案，以测试其测序数据中的各种基因组组装。这两个工作流程都使用 conda 打包系统，因此无需手动安装每个程序：这两个工作流均为 MIT 许可下的开源软件，分别位于 github.com/pmenzel/ont-assembly-snake 和 github.com/pmenzel/score-assemblies。

{"title":"Snakemake workflows for long-read bacterial genome assembly and evaluation.","authors":"Peter Menzel","doi":"10.46471/gigabyte.116","DOIUrl":"10.46471/gigabyte.116","url":null,"abstract":"With the advancement of long-read sequencing technologies and their increasing use for bacterial genomics, several methods for generating genome assemblies from error-prone long reads have been developed. These are complemented by various tools for assembly polishing using either long reads, short reads, or reference genomes. End users are therefore left with a plethora of possible combinations of programs for obtaining a final trusted assembly. Hence, there is also a need to measure the completeness and accuracy of such assemblies, for which, again, several evaluation methods implemented in various programs are available. In order to automatically run multiple genome assembly and evaluation programs at once, I developed two workflows for the workflow management system Snakemake, which provide end users with an easy-to-run solution for testing various genome assemblies from their sequencing data. Both workflows use the conda packaging system, so there is no need for manual installation of each program.Availability & implementation: The workflows are available as open source software under the MIT license at github.com/pmenzel/ont-assembly-snake and github.com/pmenzel/score-assemblies.","PeriodicalId":73157,"journal":{"name":"GigaByte (Hong Kong, China)","volume":"2024 ","pages":"gigabyte116"},"PeriodicalIF":0.0,"publicationDate":"2024-04-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC11000499/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"140874304","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":0,"RegionCategory":"","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}

引用次数: 0

Whole genome assembly and annotation of the King Angelfish (Holacanthus passer) gives insight into the evolution of marine fishes of the Tropical Eastern Pacific. 帝王姬鱼（Holacanthus passer）的全基因组组装和注释有助于深入了解热带东太平洋海洋鱼类的进化。

GigaByte (Hong Kong, China)

Pub Date : 2024-03-21 eCollection Date: 2024-01-01 DOI: 10.46471/gigabyte.115

Remy Gatins, Carlos F Arias, Carlos Sánchez, Giacomo Bernardi, Luis F De León

Holacanthus angelfishes are some of the most iconic marine fishes of the Tropical Eastern Pacific (TEP). However, very limited genomic resources currently exist for the genus. In this study we: (i) assembled and annotated the nuclear genome of the King Angelfish (Holacanthus passer), and (ii) examined the demographic history of H. passer in the TEP. We generated 43.8 Gb of ONT and 97.3 Gb Illumina reads representing 75× and 167× coverage, respectively. The final genome assembly size was 583 Mb with a contig N50 of 5.7 Mb, which captured 97.5% of the complete Actinoterygii Benchmarking Universal Single-Copy Orthologs (BUSCOs). Repetitive elements accounted for 5.09% of the genome, and 33,889 protein-coding genes were predicted, of which 22,984 were functionally annotated. Our demographic analysis suggests that population expansions of H. passer occurred prior to the last glacial maximum (LGM) and were more likely shaped by events associated with the closure of the Isthmus of Panama. This result is surprising, given that most rapid population expansions in both freshwater and marine organisms have been reported to occur globally after the LGM. Overall, this annotated genome assembly provides a novel molecular resource to study the evolution of Holacanthus angelfishes, while facilitating research into local adaptation, speciation, and introgression in marine fishes.

天使鱼（Holacanthus angelfishes）是东太平洋热带地区（TEP）一些最具代表性的海洋鱼类。然而，目前该属的基因组资源非常有限。在这项研究中，我们(i) 组装并注释了帝王吴郭鱼（Holacanthus passer）的核基因组，(ii) 研究了帝王吴郭鱼在热带东太平洋的种群历史。我们生成了 43.8 Gb ONT 和 97.3 Gb Illumina 读数，覆盖率分别为 75 倍和 167 倍。最终的基因组组装大小为 583 Mb，等位基因 N50 为 5.7 Mb，捕获了 97.5% 的完整的放线虫基准通用单拷贝同源物（BUSCOs）。重复元件占基因组的 5.09%，预测了 33,889 个编码蛋白质的基因，其中 22,984 个已进行了功能注释。我们的人口学分析表明，H. passer的种群扩张发生在上一个冰川极盛时期（LGM）之前，更有可能是由与巴拿马地峡关闭相关的事件形成的。这一结果令人惊讶，因为据报道，淡水和海洋生物的大多数快速种群扩张都发生在全球大冰川时期之后。总之，该注释基因组的组装为研究 Holacanthus Angelf 鱼的进化提供了新的分子资源，同时也促进了对海洋鱼类的局部适应、物种分化和引种的研究。

{"title":"Whole genome assembly and annotation of the King Angelfish (Holacanthus passer) gives insight into the evolution of marine fishes of the Tropical Eastern Pacific.","authors":"Remy Gatins, Carlos F Arias, Carlos Sánchez, Giacomo Bernardi, Luis F De León","doi":"10.46471/gigabyte.115","DOIUrl":"10.46471/gigabyte.115","url":null,"abstract":"Holacanthus angelfishes are some of the most iconic marine fishes of the Tropical Eastern Pacific (TEP). However, very limited genomic resources currently exist for the genus. In this study we: (i) assembled and annotated the nuclear genome of the King Angelfish (Holacanthus passer), and (ii) examined the demographic history of H. passer in the TEP. We generated 43.8 Gb of ONT and 97.3 Gb Illumina reads representing 75× and 167× coverage, respectively. The final genome assembly size was 583 Mb with a contig N50 of 5.7 Mb, which captured 97.5% of the complete Actinoterygii Benchmarking Universal Single-Copy Orthologs (BUSCOs). Repetitive elements accounted for 5.09% of the genome, and 33,889 protein-coding genes were predicted, of which 22,984 were functionally annotated. Our demographic analysis suggests that population expansions of H. passer occurred prior to the last glacial maximum (LGM) and were more likely shaped by events associated with the closure of the Isthmus of Panama. This result is surprising, given that most rapid population expansions in both freshwater and marine organisms have been reported to occur globally after the LGM. Overall, this annotated genome assembly provides a novel molecular resource to study the evolution of Holacanthus angelfishes, while facilitating research into local adaptation, speciation, and introgression in marine fishes.","PeriodicalId":73157,"journal":{"name":"GigaByte (Hong Kong, China)","volume":"2024 ","pages":"gigabyte115"},"PeriodicalIF":0.0,"publicationDate":"2024-03-21","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC10973836/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"140320042","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":0,"RegionCategory":"","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}

引用次数: 0

Molecular Property Diagnostic Suite for COVID-19 (MPDS^COVID-19): an open-source disease-specific drug discovery portal. 用于 COVID-19 的分子特性诊断套件（MPDSCOVID-19）：一个开源的疾病特异性药物发现门户网站。

GigaByte (Hong Kong, China)

Pub Date : 2024-03-14 eCollection Date: 2024-01-01 DOI: 10.46471/gigabyte.114

Lipsa Priyadarsinee, Esther Jamir, Selvaraman Nagamani, Hridoy Jyoti Mahanta, Nandan Kumar, Lijo John, Himakshi Sarma, Asheesh Kumar, Anamika Singh Gaur, Rosaleen Sahoo, S Vaikundamani, N Arul Murugan, U Deva Priyakumar, G P S Raghava, Prasad V Bharatam, Ramakrishnan Parthasarathi, V Subramanian, G Madhavi Sastry, G Narahari Sastry

Molecular Property Diagnostic Suite (MPDS) was conceived and developed as an open-source disease-specific web portal based on Galaxy. MPDS^COVID-19 was developed for COVID-19 as a one-stop solution for drug discovery research. Galaxy platforms enable the creation of customized workflows connecting various modules in the web server. The architecture of MPDS^COVID-19 effectively employs Galaxy v22.04 features, which are ported on CentOS 7.8 and Python 3.7. MPDS^COVID-19 provides significant updates and the addition of several new tools updated after six years. Tools developed by our group in Perl/Python and open-source tools are collated and integrated into MPDS^COVID-19 using XML scripts. Our MPDS suite aims to facilitate transparent and open innovation. This approach significantly helps bring inclusiveness in the community while promoting free access and participation in software development.

Availability & implementation: The MPDS^COVID-19 portal can be accessed at https://mpds.neist.res.in:8085/.

Molecular Property Diagnostic Suite (MPDS) 是基于 Galaxy 构想和开发的一个开源疾病特定门户网站。MPDSCOVID-19 是为 COVID-19 开发的药物发现研究一站式解决方案。通过 Galaxy 平台，可以创建连接网络服务器中各种模块的定制工作流程。MPDSCOVID-19 的架构有效利用了 Galaxy v22.04 的功能，并在 CentOS 7.8 和 Python 3.7 上进行了移植。MPDSCOVID-19 在六年后进行了重大更新，并增加了几个新工具。我们小组用 Perl/Python 开发的工具和开源工具经过整理，使用 XML 脚本集成到了 MPDSCOVID-19 中。我们的 MPDS 套件旨在促进透明、开放的创新。这种方法大大有助于提高社区的包容性，同时促进自由访问和参与软件开发：MPDSCOVID-19 门户网站的访问网址为 https://mpds.neist.res.in:8085/。

{"title":"Molecular Property Diagnostic Suite for COVID-19 (MPDSCOVID-19): an open-source disease-specific drug discovery portal.","authors":"Lipsa Priyadarsinee, Esther Jamir, Selvaraman Nagamani, Hridoy Jyoti Mahanta, Nandan Kumar, Lijo John, Himakshi Sarma, Asheesh Kumar, Anamika Singh Gaur, Rosaleen Sahoo, S Vaikundamani, N Arul Murugan, U Deva Priyakumar, G P S Raghava, Prasad V Bharatam, Ramakrishnan Parthasarathi, V Subramanian, G Madhavi Sastry, G Narahari Sastry","doi":"10.46471/gigabyte.114","DOIUrl":"10.46471/gigabyte.114","url":null,"abstract":"Molecular Property Diagnostic Suite (MPDS) was conceived and developed as an open-source disease-specific web portal based on Galaxy. MPDSCOVID-19 was developed for COVID-19 as a one-stop solution for drug discovery research. Galaxy platforms enable the creation of customized workflows connecting various modules in the web server. The architecture of MPDSCOVID-19 effectively employs Galaxy v22.04 features, which are ported on CentOS 7.8 and Python 3.7. MPDSCOVID-19 provides significant updates and the addition of several new tools updated after six years. Tools developed by our group in Perl/Python and open-source tools are collated and integrated into MPDSCOVID-19 using XML scripts. Our MPDS suite aims to facilitate transparent and open innovation. This approach significantly helps bring inclusiveness in the community while promoting free access and participation in software development.Availability & implementation: The MPDSCOVID-19 portal can be accessed at https://mpds.neist.res.in:8085/.","PeriodicalId":73157,"journal":{"name":"GigaByte (Hong Kong, China)","volume":"2024 ","pages":"gigabyte114"},"PeriodicalIF":0.0,"publicationDate":"2024-03-14","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC10958779/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"140208383","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":0,"RegionCategory":"","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}

引用次数: 0

Julearn: an easy-to-use library for leakage-free evaluation and inspection of ML models. Julearn：一个易于使用的库，用于对 ML 模型进行无泄漏评估和检查。

GigaByte (Hong Kong, China)

Pub Date : 2024-03-07 eCollection Date: 2024-01-01 DOI: 10.46471/gigabyte.113

Sami Hamdan, Shammi More, Leonard Sasse, Vera Komeyer, Kaustubh R Patil, Federico Raimondo

The fast-paced development of machine learning (ML) and its increasing adoption in research challenge researchers without extensive training in ML. In neuroscience, ML can help understand brain-behavior relationships, diagnose diseases and develop biomarkers using data from sources like magnetic resonance imaging and electroencephalography. Primarily, ML builds models to make accurate predictions on unseen data. Researchers evaluate models' performance and generalizability using techniques such as cross-validation (CV). However, choosing a CV scheme and evaluating an ML pipeline is challenging and, if done improperly, can lead to overestimated results and incorrect interpretations. Here, we created julearn, an open-source Python library allowing researchers to design and evaluate complex ML pipelines without encountering common pitfalls. We present the rationale behind julearn's design, its core features, and showcase three examples of previously-published research projects. Julearn simplifies the access to ML providing an easy-to-use environment. With its design, unique features, simple interface, and practical documentation, it poses as a useful Python-based library for research projects.

机器学习（ML）的发展日新月异，在研究领域的应用也日益广泛，这对没有接受过广泛 ML 培训的研究人员提出了挑战。在神经科学领域，ML 可以帮助理解大脑与行为之间的关系，利用磁共振成像和脑电图等数据源诊断疾病和开发生物标记物。ML 主要是建立模型，对未见数据进行准确预测。研究人员使用交叉验证（CV）等技术评估模型的性能和可推广性。然而，选择交叉验证方案和评估 ML 管道具有挑战性，如果操作不当，可能会导致结果被高估和解释错误。在这里，我们创建了 julearn，这是一个开源 Python 库，允许研究人员设计和评估复杂的 ML 管道，而不会遇到常见的陷阱。我们介绍了 julearn 的设计原理、核心功能，并展示了之前发表的三个研究项目实例。Julearn 提供了一个易于使用的环境，简化了对 ML 的访问。凭借其设计、独特的功能、简单的界面和实用的文档，它成为研究项目中一个有用的基于 Python 的库。

{"title":"Julearn: an easy-to-use library for leakage-free evaluation and inspection of ML models.","authors":"Sami Hamdan, Shammi More, Leonard Sasse, Vera Komeyer, Kaustubh R Patil, Federico Raimondo","doi":"10.46471/gigabyte.113","DOIUrl":"10.46471/gigabyte.113","url":null,"abstract":"The fast-paced development of machine learning (ML) and its increasing adoption in research challenge researchers without extensive training in ML. In neuroscience, ML can help understand brain-behavior relationships, diagnose diseases and develop biomarkers using data from sources like magnetic resonance imaging and electroencephalography. Primarily, ML builds models to make accurate predictions on unseen data. Researchers evaluate models' performance and generalizability using techniques such as cross-validation (CV). However, choosing a CV scheme and evaluating an ML pipeline is challenging and, if done improperly, can lead to overestimated results and incorrect interpretations. Here, we created julearn, an open-source Python library allowing researchers to design and evaluate complex ML pipelines without encountering common pitfalls. We present the rationale behind julearn's design, its core features, and showcase three examples of previously-published research projects. Julearn simplifies the access to ML providing an easy-to-use environment. With its design, unique features, simple interface, and practical documentation, it poses as a useful Python-based library for research projects.","PeriodicalId":73157,"journal":{"name":"GigaByte (Hong Kong, China)","volume":"2024 ","pages":"gigabyte113"},"PeriodicalIF":0.0,"publicationDate":"2024-03-07","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC10940896/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"140144689","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":0,"RegionCategory":"","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}

引用次数: 0

An improved chromosome-level genome assembly of perennial ryegrass (Lolium perenne L.). 改进的多年生黑麦草（Lolium perenne L.）染色体组水平基因组组装。

GigaByte (Hong Kong, China)

Pub Date : 2024-03-06 eCollection Date: 2024-01-01 DOI: 10.46471/gigabyte.112

Yutang Chen, Roland Kölliker, Martin Mascher, Dario Copetti, Axel Himmelbach, Nils Stein, Bruno Studer

This work is an update and extension of the previously published article "Ultralong Oxford Nanopore Reads Enable the Development of a Reference-Grade Perennial Ryegrass Genome Assembly" by Frei et al. The published genome assembly of the doubled haploid perennial ryegrass (Lolium perenne L.) genotype Kyuss (Kyuss v1.0) marked a milestone for forage grass research and breeding. However, order and orientation errors may exist in the pseudo-chromosomes of Kyuss, since barley (Hordeum vulgare L.), which diverged 30 million years ago from perennial ryegrass, was used as the reference to scaffold Kyuss. To correct for structural errors possibly present in the published Kyuss assembly, we de novo assembled the genome again and generated 50-fold coverage high-throughput chromosome conformation capture (Hi-C) data to assist pseudo-chromosome construction. The resulting new chromosome-level assembly Kyuss v2.0 showed improved quality with high contiguity (contig N50 = 120 Mb), high completeness (total BUSCO score = 99%), high base-level accuracy (QV = 50), and correct pseudo-chromosome structure (validated by Hi-C contact map). This new assembly will serve as a better reference genome for Lolium spp. and greatly benefit the forage and turf grass research community.

这项工作是对 Frei 等人以前发表的文章《超长牛津纳米孔读数促成了参考级多年生黑麦草基因组组装的开发》的更新和扩展。双倍单倍体多年生黑麦草（Lolium perenne L.）基因型 Kyuss（Kyuss v1.0）基因组组装的发表标志着牧草研究和育种的一个里程碑。然而，由于大麦（Hordeum vulgare L.）与多年生黑麦草在 3000 万年前就已分化，因此 Kyuss 的假染色体可能存在顺序和方向错误。为了纠正已发表的 Kyuss 组装中可能存在的结构错误，我们重新组装了基因组，并生成了 50 倍覆盖率的高通量染色体构象捕获（Hi-C）数据，以帮助构建假染色体。由此产生的新的染色体级组装Kyuss v2.0显示出更高的质量，具有高毗连性（毗连N50 = 120 Mb）、高完整性（BUSCO总分 = 99%）、高碱基水平准确性（QV = 50）和正确的假染色体结构（由Hi-C接触图验证）。这一新的基因组将成为更好的洛仑草属（Lolium spp.）参考基因组，对牧草和草坪草研究界大有裨益。

{"title":"An improved chromosome-level genome assembly of perennial ryegrass (Lolium perenne L.).","authors":"Yutang Chen, Roland Kölliker, Martin Mascher, Dario Copetti, Axel Himmelbach, Nils Stein, Bruno Studer","doi":"10.46471/gigabyte.112","DOIUrl":"10.46471/gigabyte.112","url":null,"abstract":"This work is an update and extension of the previously published article \"Ultralong Oxford Nanopore Reads Enable the Development of a Reference-Grade Perennial Ryegrass Genome Assembly\" by Frei et al. The published genome assembly of the doubled haploid perennial ryegrass (Lolium perenne L.) genotype Kyuss (Kyuss v1.0) marked a milestone for forage grass research and breeding. However, order and orientation errors may exist in the pseudo-chromosomes of Kyuss, since barley (Hordeum vulgare L.), which diverged 30 million years ago from perennial ryegrass, was used as the reference to scaffold Kyuss. To correct for structural errors possibly present in the published Kyuss assembly, we de novo assembled the genome again and generated 50-fold coverage high-throughput chromosome conformation capture (Hi-C) data to assist pseudo-chromosome construction. The resulting new chromosome-level assembly Kyuss v2.0 showed improved quality with high contiguity (contig N50 = 120 Mb), high completeness (total BUSCO score = 99%), high base-level accuracy (QV = 50), and correct pseudo-chromosome structure (validated by Hi-C contact map). This new assembly will serve as a better reference genome for Lolium spp. and greatly benefit the forage and turf grass research community.","PeriodicalId":73157,"journal":{"name":"GigaByte (Hong Kong, China)","volume":"2024 ","pages":"gigabyte112"},"PeriodicalIF":0.0,"publicationDate":"2024-03-06","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://www.ncbi.nlm.nih.gov/pmc/articles/PMC10940895/pdf/","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"140144688","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":0,"RegionCategory":"","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}

引用次数: 0

首页上一页

下一页尾页

类型

全部化学•材料生命科学医学物理工程技术环境•农林材料科学地球科学法学管理学化学环境科学与生态学计算机科学教育学经济学农林科学人文科学生物学数学物理与天体物理心理学综合性期刊其他工业工程理学历史学农学文学信息工程

数据库

全部 ACS Publications Elsevier ieeexplore Springer The Royal Society of Chemistry Wiley

期刊

GigaByte (Hong Kong, China)

全部 Acc. Chem. Res. ACS Applied Bio Materials ACS Appl. Electron. Mater. ACS Appl. Energy Mater. ACS Appl. Mater. Interfaces ACS Appl. Nano Mater. ACS Appl. Polym. Mater. ACS BIOMATER-SCI ENG ACS Catal. ACS Cent. Sci. ACS Chem. Biol. ACS Chemical Health & Safety ACS Chem. Neurosci. ACS Comb. Sci. ACS Earth Space Chem. ACS Energy Lett. ACS Infect. Dis. ACS Macro Lett. ACS Mater. Lett. ACS Med. Chem. Lett. ACS Nano ACS Omega ACS Photonics ACS Sens. ACS Sustainable Chem. Eng. ACS Synth. Biol. Anal. Chem. BIOCHEMISTRY-US Bioconjugate Chem. BIOMACROMOLECULES Chem. Res. Toxicol. Chem. Rev. Chem. Mater. CRYST GROWTH DES ENERG FUEL Environ. Sci. Technol. Environ. Sci. Technol. Lett. Eur. J. Inorg. Chem. IND ENG CHEM RES Inorg. Chem. J. Agric. Food. Chem. J. Chem. Eng. Data J. Chem. Educ. J. Chem. Inf. Model. J. Chem. Theory Comput. J. Med. Chem. J. Nat. Prod. J PROTEOME RES J. Am. Chem. Soc. LANGMUIR MACROMOLECULES Mol. Pharmaceutics Nano Lett. Org. Lett. ORG PROCESS RES DEV ORGANOMETALLICS J. Org. Chem. J. Phys. Chem. J. Phys. Chem. A J. Phys. Chem. B J. Phys. Chem. C J. Phys. Chem. Lett. Analyst Anal. Methods Biomater. Sci. Catal. Sci. Technol. Chem. Commun. Chem. Soc. Rev. CHEM EDUC RES PRACT CRYSTENGCOMM Dalton Trans. Energy Environ. Sci. ENVIRON SCI-NANO ENVIRON SCI-PROC IMP ENVIRON SCI-WAT RES Faraday Discuss. Food Funct. Green Chem. Inorg. Chem. Front. Integr. Biol. J. Anal. At. Spectrom. J. Mater. Chem. A J. Mater. Chem. B J. Mater. Chem. C Lab Chip Mater. Chem. Front. Mater. Horiz. MEDCHEMCOMM Metallomics Mol. Biosyst. Mol. Syst. Des. Eng. Nanoscale Nanoscale Horiz. Nat. Prod. Rep. New J. Chem. Org. Biomol. Chem. Org. Chem. Front. PHOTOCH PHOTOBIO SCI PCCP Polym. Chem.

﹀