{"title":"HPClas:基于 catBoost 的数据驱动型嗜卤蛋白质识别方法","authors":"Shantong Hu, Xiao-Yong Wang, Zhikang Wang, Menghan Jiang, Shihui Wang, Wenya Wang, Jiangning Song, Guimin Zhang","doi":"10.1002/mlf2.12125","DOIUrl":null,"url":null,"abstract":"Halophilic proteins possess unique structural properties and show high stability under extreme conditions. This distinct characteristic makes them invaluable for application in various aspects such as bioenergy, pharmaceuticals, environmental clean‐up, and energy production. Generally, halophilic proteins are discovered and characterized through labor‐intensive and time‐consuming wet lab experiments. In this study, we introduce the Halophilic Protein Classifier (HPClas), a machine learning‐based classifier developed using the catBoost ensemble learning technique to identify halophilic proteins. Extensive in silico calculations were conducted on a large public dataset of 12,574 samples and HPClas achieved an area under the receiver operating characteristic curve (AUROC) of 0.844 on an independent test set of 200 samples. The source code and curated dataset of HPClas are publicly available at https://github.com/Showmake2/HPClas. In conclusion, HPClas can be explored as a promising tool to aid in the identification of halophilic proteins and accelerate their application in different fields.","PeriodicalId":94145,"journal":{"name":"mLife","volume":null,"pages":null},"PeriodicalIF":4.5000,"publicationDate":"2024-07-20","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":"{\"title\":\"HPClas: A data‐driven approach for identifying halophilic proteins based on catBoost\",\"authors\":\"Shantong Hu, Xiao-Yong Wang, Zhikang Wang, Menghan Jiang, Shihui Wang, Wenya Wang, Jiangning Song, Guimin Zhang\",\"doi\":\"10.1002/mlf2.12125\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"Halophilic proteins possess unique structural properties and show high stability under extreme conditions. This distinct characteristic makes them invaluable for application in various aspects such as bioenergy, pharmaceuticals, environmental clean‐up, and energy production. Generally, halophilic proteins are discovered and characterized through labor‐intensive and time‐consuming wet lab experiments. In this study, we introduce the Halophilic Protein Classifier (HPClas), a machine learning‐based classifier developed using the catBoost ensemble learning technique to identify halophilic proteins. Extensive in silico calculations were conducted on a large public dataset of 12,574 samples and HPClas achieved an area under the receiver operating characteristic curve (AUROC) of 0.844 on an independent test set of 200 samples. The source code and curated dataset of HPClas are publicly available at https://github.com/Showmake2/HPClas. In conclusion, HPClas can be explored as a promising tool to aid in the identification of halophilic proteins and accelerate their application in different fields.\",\"PeriodicalId\":94145,\"journal\":{\"name\":\"mLife\",\"volume\":null,\"pages\":null},\"PeriodicalIF\":4.5000,\"publicationDate\":\"2024-07-20\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"0\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"mLife\",\"FirstCategoryId\":\"0\",\"ListUrlMain\":\"https://doi.org/10.1002/mlf2.12125\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"Q1\",\"JCRName\":\"MICROBIOLOGY\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"mLife","FirstCategoryId":"0","ListUrlMain":"https://doi.org/10.1002/mlf2.12125","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"Q1","JCRName":"MICROBIOLOGY","Score":null,"Total":0}
HPClas: A data‐driven approach for identifying halophilic proteins based on catBoost
Halophilic proteins possess unique structural properties and show high stability under extreme conditions. This distinct characteristic makes them invaluable for application in various aspects such as bioenergy, pharmaceuticals, environmental clean‐up, and energy production. Generally, halophilic proteins are discovered and characterized through labor‐intensive and time‐consuming wet lab experiments. In this study, we introduce the Halophilic Protein Classifier (HPClas), a machine learning‐based classifier developed using the catBoost ensemble learning technique to identify halophilic proteins. Extensive in silico calculations were conducted on a large public dataset of 12,574 samples and HPClas achieved an area under the receiver operating characteristic curve (AUROC) of 0.844 on an independent test set of 200 samples. The source code and curated dataset of HPClas are publicly available at https://github.com/Showmake2/HPClas. In conclusion, HPClas can be explored as a promising tool to aid in the identification of halophilic proteins and accelerate their application in different fields.