单独抽样下混合概率对交叉验证偏差的影响

2013 IEEE International Workshop on Genomic Signal Processing and Statistics Pub Date : 2013-11-01 DOI:10.1109/GENSIPS.2013.6735947

A. Zollanvari, U. Braga-Neto, E. Dougherty

{"title":"单独抽样下混合概率对交叉验证偏差的影响","authors":"A. Zollanvari, U. Braga-Neto, E. Dougherty","doi":"10.1109/GENSIPS.2013.6735947","DOIUrl":null,"url":null,"abstract":"Cross-validation is commonly used to estimate the overall error rate of a designed classifier in a small-sample expression study. The true error of the classifier is a function of the prior probabilities of the classes. With random sampling these can be estimated consistently in terms of the class sample sizes, but when sampling is separate, meaning these sample sizes are determined prior to sampling, there are no reasonable estimates from the data and the prior probabilities must be “estimated” outside the experiment. We have conducted a set of simulations to study the bias of cross-validation as a function of these “estimates”. The results show that a poor choice for estimating these probabilities can significantly increase the bias of cross-validation as an estimator of the true error.","PeriodicalId":336511,"journal":{"name":"2013 IEEE International Workshop on Genomic Signal Processing and Statistics","volume":"34 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2013-11-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":"{\"title\":\"Effect of mixing probabilities on the bias of cross-validation under separate sampling\",\"authors\":\"A. Zollanvari, U. Braga-Neto, E. Dougherty\",\"doi\":\"10.1109/GENSIPS.2013.6735947\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"Cross-validation is commonly used to estimate the overall error rate of a designed classifier in a small-sample expression study. The true error of the classifier is a function of the prior probabilities of the classes. With random sampling these can be estimated consistently in terms of the class sample sizes, but when sampling is separate, meaning these sample sizes are determined prior to sampling, there are no reasonable estimates from the data and the prior probabilities must be “estimated” outside the experiment. We have conducted a set of simulations to study the bias of cross-validation as a function of these “estimates”. The results show that a poor choice for estimating these probabilities can significantly increase the bias of cross-validation as an estimator of the true error.\",\"PeriodicalId\":336511,\"journal\":{\"name\":\"2013 IEEE International Workshop on Genomic Signal Processing and Statistics\",\"volume\":\"34 1\",\"pages\":\"0\"},\"PeriodicalIF\":0.0000,\"publicationDate\":\"2013-11-01\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"0\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"2013 IEEE International Workshop on Genomic Signal Processing and Statistics\",\"FirstCategoryId\":\"1085\",\"ListUrlMain\":\"https://doi.org/10.1109/GENSIPS.2013.6735947\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"\",\"JCRName\":\"\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"2013 IEEE International Workshop on Genomic Signal Processing and Statistics","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1109/GENSIPS.2013.6735947","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 0

摘要

交叉验证通常用于估计小样本表达研究中设计分类器的总体错误率。分类器的真实误差是类的先验概率的函数。通过随机抽样，可以根据类样本大小一致地估计这些，但当抽样是分开的，意味着这些样本大小是在抽样之前确定的，从数据中没有合理的估计，先验概率必须在实验之外“估计”。我们进行了一组模拟来研究交叉验证的偏差作为这些“估计”的函数。结果表明，对于估计这些概率的不良选择会显著增加交叉验证作为真实误差估计的偏差。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

查看原文

微信好友朋友圈 QQ好友复制链接

本刊更多论文

Effect of mixing probabilities on the bias of cross-validation under separate sampling

Cross-validation is commonly used to estimate the overall error rate of a designed classifier in a small-sample expression study. The true error of the classifier is a function of the prior probabilities of the classes. With random sampling these can be estimated consistently in terms of the class sample sizes, but when sampling is separate, meaning these sample sizes are determined prior to sampling, there are no reasonable estimates from the data and the prior probabilities must be “estimated” outside the experiment. We have conducted a set of simulations to study the bias of cross-validation as a function of these “estimates”. The results show that a poor choice for estimating these probabilities can significantly increase the bias of cross-validation as an estimator of the true error.

求助全文

通过发布文献求助，成功后即可免费获取论文全文。去求助

来源期刊

2013 IEEE International Workshop on Genomic Signal Processing and Statistics

自引率

0.00%

发文量