Choosing SNPs using feature selection.

Proceedings. IEEE Computational Systems Bioinformatics Conference Pub Date : 2005-01-01 DOI:10.1109/csb.2005.22

Tu Minh Phuong, Zhen Lin, Russ B Altman

引用次数: 89

Abstract

A major challenge for genomewide disease association studies is the high cost of genotyping large number of single nucleotide polymorphisms (SNP). The correlations between SNPs, however, make it possible to select a parsimonious set of informative SNPs, known as "tagging" SNPs, able to capture most variation in a population. Considerable research interest has recently focused on the development of methods for finding such SNPs. In this paper, we present an efficient method for finding tagging SNPs. The method does not involve computation-intensive search for SNP subsets but discards redundant SNPs using a feature selection algorithm. In contrast to most existing methods, the method presented here does not limit itself to using only correlations between SNPs in local groups. By using correlations that occur across different chromosomal regions, the method can reduce the number of globally redundant SNPs. Experimental results show that the number of tagging SNPs selected by our method is smaller than by using block-based methods.

查看原文

微信好友朋友圈 QQ好友复制链接

本刊更多论文

使用特征选择选择snp。

全基因组疾病关联研究的一个主要挑战是对大量单核苷酸多态性(SNP)进行基因分型的高成本。然而，snp之间的相关性使得选择一组信息丰富的snp成为可能，这些snp被称为“标记”snp，能够捕获种群中的大多数变异。相当大的研究兴趣最近集中在寻找这种snp的方法的发展上。在本文中，我们提出了一种有效的方法来寻找标记snp。该方法不涉及对SNP子集的计算密集型搜索，而是使用特征选择算法丢弃冗余SNP。与大多数现有方法相比，本文提出的方法并不局限于仅使用本地群体中snp之间的相关性。通过使用发生在不同染色体区域的相关性，该方法可以减少全局冗余snp的数量。实验结果表明，与基于块的方法相比，该方法选择的标记snp数量更少。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文去求助

来源期刊

Proceedings. IEEE Computational Systems Bioinformatics Conference

自引率

0.00%

发文量