Ruilin Su, Jingbo Lv, Yahui Xue, Sheng Jiang, Lei Zhou, Li Jiang, Junyan Tan, Zhencai Shen, Ping Zhong, Jianfeng Liu
{"title":"Genomic selection in pig breeding: comparative analysis of machine learning algorithms","authors":"Ruilin Su, Jingbo Lv, Yahui Xue, Sheng Jiang, Lei Zhou, Li Jiang, Junyan Tan, Zhencai Shen, Ping Zhong, Jianfeng Liu","doi":"10.1186/s12711-025-00957-3","DOIUrl":null,"url":null,"abstract":"The effectiveness of genomic prediction (GP) significantly influences breeding progress, and employing SNP markers to predict phenotypic values is a pivotal aspect of pig breeding. Machine learning (ML) methods are usually used to predict phenotypic values since their advantages in processing high dimensional data. While, the existing researches have not indicated which ML methods are suitable for most pig genomic prediction. Therefore, it is necessary to select appropriate methods from a large number of ML methods as long as genomic prediction is performed. This paper compared the performance of popular ML methods in predicting pig phenotypes and then found out suitable methods for most traits. In this paper, five commonly used datasets from other literatures were utilized to compare the performance of different ML methods. The experimental results demonstrate that Stacking performs best on the PIC dataset where the trait information is hidden, and the performs of kernel ridge regression with rbf kernel (KRR-rbf) closely follows. Support vector regression (SVR) performs best in predicting reproductive traits, followed by genomic best linear unbiased prediction (GBLUP). GBLUP achieves the best performance on growth traits, with SVR as the second best. GBLUP achieves good performance for GP problems. Similarly, the Stacking, SVR, and KRR-RBF methods also achieve high prediction accuracy. Moreover, LR statistical analysis shows that Stacking, SVR and KRR are stable. When applying ML methods for phenotypic values prediction in pigs, we recommend these three approaches.","PeriodicalId":55120,"journal":{"name":"Genetics Selection Evolution","volume":"38 1","pages":""},"PeriodicalIF":3.6000,"publicationDate":"2025-03-10","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"Genetics Selection Evolution","FirstCategoryId":"99","ListUrlMain":"https://doi.org/10.1186/s12711-025-00957-3","RegionNum":1,"RegionCategory":"农林科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"Q1","JCRName":"AGRICULTURE, DAIRY & ANIMAL SCIENCE","Score":null,"Total":0}
引用次数: 0
Abstract
The effectiveness of genomic prediction (GP) significantly influences breeding progress, and employing SNP markers to predict phenotypic values is a pivotal aspect of pig breeding. Machine learning (ML) methods are usually used to predict phenotypic values since their advantages in processing high dimensional data. While, the existing researches have not indicated which ML methods are suitable for most pig genomic prediction. Therefore, it is necessary to select appropriate methods from a large number of ML methods as long as genomic prediction is performed. This paper compared the performance of popular ML methods in predicting pig phenotypes and then found out suitable methods for most traits. In this paper, five commonly used datasets from other literatures were utilized to compare the performance of different ML methods. The experimental results demonstrate that Stacking performs best on the PIC dataset where the trait information is hidden, and the performs of kernel ridge regression with rbf kernel (KRR-rbf) closely follows. Support vector regression (SVR) performs best in predicting reproductive traits, followed by genomic best linear unbiased prediction (GBLUP). GBLUP achieves the best performance on growth traits, with SVR as the second best. GBLUP achieves good performance for GP problems. Similarly, the Stacking, SVR, and KRR-RBF methods also achieve high prediction accuracy. Moreover, LR statistical analysis shows that Stacking, SVR and KRR are stable. When applying ML methods for phenotypic values prediction in pigs, we recommend these three approaches.
期刊介绍:
Genetics Selection Evolution invites basic, applied and methodological content that will aid the current understanding and the utilization of genetic variability in domestic animal species. Although the focus is on domestic animal species, research on other species is invited if it contributes to the understanding of the use of genetic variability in domestic animals. Genetics Selection Evolution publishes results from all levels of study, from the gene to the quantitative trait, from the individual to the population, the breed or the species. Contributions concerning both the biological approach, from molecular genetics to quantitative genetics, as well as the mathematical approach, from population genetics to statistics, are welcome. Specific areas of interest include but are not limited to: gene and QTL identification, mapping and characterization, analysis of new phenotypes, high-throughput SNP data analysis, functional genomics, cytogenetics, genetic diversity of populations and breeds, genetic evaluation, applied and experimental selection, genomic selection, selection efficiency, and statistical methodology for the genetic analysis of phenotypes with quantitative and mixed inheritance.