20D-Dynamic Representation of Protein Sequences Combined with K-means Clustering.

IF 1.6 4区 医学 Q4 BIOCHEMICAL RESEARCH METHODS Combinatorial chemistry & high throughput screening Pub Date : 2025-02-26 DOI:10.2174/0113862073359729250220131623
Dorota Bielińska-Wąż, Piotr Wąż, Agata Błaczkowska
{"title":"20D-Dynamic Representation of Protein Sequences Combined with K-means Clustering.","authors":"Dorota Bielińska-Wąż, Piotr Wąż, Agata Błaczkowska","doi":"10.2174/0113862073359729250220131623","DOIUrl":null,"url":null,"abstract":"<p><strong>Objective: </strong>The objective of this research is to demonstrate that alignment-free bioinformatics approaches are effective tools for analyzing the similarity and dissimilarity of protein sequences. All numerical parameters representing sequences are expressed analytically, ensuring precision, clarity, and efficient processing, even for large datasets and long sequences. Additionally, a novel approach for identifying previously unknown virus strains is introduced.</p><p><strong>Methods: </strong>A novel approach is proposed, integrating the unique features of our newly developed method, the 20D-Dynamic Representation of Protein Sequences, with the K-means clustering algorithm. The sequences are represented as clouds of material points in a 20-dimensional space (20D-dynamic graphs), with their spatial distribution being unique to each protein sequence. The numerical parameters, referred to as descriptors in molecular similarity theory, represent quantities characteristic of dynamic systems and serve as input data for the K-means clustering algorithm.</p><p><strong>Results: </strong>Examples of the application of the approach are presented, including projections of the 20D-dynamic graphs onto 3D spaces, which serve as a visual tool for comparing sequences. Additionally, cluster plots for the analyzed sequences are provided using the proposed method.</p><p><strong>Conclusion: </strong>It has been demonstrated that the 20D-Dynamic Representation of Protein Sequences, combined with the K-means clustering algorithm, successfully classifies subtypes of influenza A virus strains.</p>","PeriodicalId":10491,"journal":{"name":"Combinatorial chemistry & high throughput screening","volume":" ","pages":""},"PeriodicalIF":1.6000,"publicationDate":"2025-02-26","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"Combinatorial chemistry & high throughput screening","FirstCategoryId":"3","ListUrlMain":"https://doi.org/10.2174/0113862073359729250220131623","RegionNum":4,"RegionCategory":"医学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"Q4","JCRName":"BIOCHEMICAL RESEARCH METHODS","Score":null,"Total":0}
引用次数: 0

Abstract

Objective: The objective of this research is to demonstrate that alignment-free bioinformatics approaches are effective tools for analyzing the similarity and dissimilarity of protein sequences. All numerical parameters representing sequences are expressed analytically, ensuring precision, clarity, and efficient processing, even for large datasets and long sequences. Additionally, a novel approach for identifying previously unknown virus strains is introduced.

Methods: A novel approach is proposed, integrating the unique features of our newly developed method, the 20D-Dynamic Representation of Protein Sequences, with the K-means clustering algorithm. The sequences are represented as clouds of material points in a 20-dimensional space (20D-dynamic graphs), with their spatial distribution being unique to each protein sequence. The numerical parameters, referred to as descriptors in molecular similarity theory, represent quantities characteristic of dynamic systems and serve as input data for the K-means clustering algorithm.

Results: Examples of the application of the approach are presented, including projections of the 20D-dynamic graphs onto 3D spaces, which serve as a visual tool for comparing sequences. Additionally, cluster plots for the analyzed sequences are provided using the proposed method.

Conclusion: It has been demonstrated that the 20D-Dynamic Representation of Protein Sequences, combined with the K-means clustering algorithm, successfully classifies subtypes of influenza A virus strains.

查看原文
分享 分享
微信好友 朋友圈 QQ好友 复制链接
本刊更多论文
求助全文
约1分钟内获得全文 去求助
来源期刊
CiteScore
3.10
自引率
5.60%
发文量
327
审稿时长
7.5 months
期刊介绍: Combinatorial Chemistry & High Throughput Screening (CCHTS) publishes full length original research articles and reviews/mini-reviews dealing with various topics related to chemical biology (High Throughput Screening, Combinatorial Chemistry, Chemoinformatics, Laboratory Automation and Compound management) in advancing drug discovery research. Original research articles and reviews in the following areas are of special interest to the readers of this journal: Target identification and validation Assay design, development, miniaturization and comparison High throughput/high content/in silico screening and associated technologies Label-free detection technologies and applications Stem cell technologies Biomarkers ADMET/PK/PD methodologies and screening Probe discovery and development, hit to lead optimization Combinatorial chemistry (e.g. small molecules, peptide, nucleic acid or phage display libraries) Chemical library design and chemical diversity Chemo/bio-informatics, data mining Compound management Pharmacognosy Natural Products Research (Chemistry, Biology and Pharmacology of Natural Products) Natural Product Analytical Studies Bipharmaceutical studies of Natural products Drug repurposing Data management and statistical analysis Laboratory automation, robotics, microfluidics, signal detection technologies Current & Future Institutional Research Profile Technology transfer, legal and licensing issues Patents.
期刊最新文献
Computational Study for Preparation of Benzoimidazo[1,2-a]pyrimidines from Reaction of Benzaldehyde, Indanedione and 1H-benzo[d]imidazol-2- amine. Agaricus blazei Murill Extract (FA-2-b-β) Induces Ferroptosis in Diffuse Large B-Cell Lymphoma via the Nrf2/Ho-1 Pathway. Dry Powder Inhaler of Sustained-Release Microspheres Containing Glycyrrhizin: Factorial Design and Optimization. Revealing the Mechanism of Buzhong Yiqi Tang in Ameliorating Autoimmune Thyroiditis via the Toll-like Receptor Pathway. 20D-Dynamic Representation of Protein Sequences Combined with K-means Clustering.
×
引用
GB/T 7714-2015
复制
MLA
复制
APA
复制
导出至
BibTeX EndNote RefMan NoteFirst NoteExpress
×
×
提示
您的信息不完整,为了账户安全,请先补充。
现在去补充
×
提示
您因"违规操作"
具体请查看互助需知
我知道了
×
提示
现在去查看 取消
×
提示
确定
0
微信
客服QQ
Book学术公众号 扫码关注我们
反馈
×
意见反馈
请填写您的意见或建议
请填写您的手机或邮箱
已复制链接
已复制链接
快去分享给好友吧!
我知道了
×
扫码分享
扫码分享
Book学术官方微信
Book学术文献互助
Book学术文献互助群
群 号:481959085
Book学术
文献互助 智能选刊 最新文献 互助须知 联系我们:info@booksci.cn
Book学术提供免费学术资源搜索服务,方便国内外学者检索中英文文献。致力于提供最便捷和优质的服务体验。
Copyright © 2023 Book学术 All rights reserved.
ghs 京公网安备 11010802042870号 京ICP备2023020795号-1