分布式顺序估计程序

IF 0.8 4区数学 Q3 STATISTICS & PROBABILITY Canadian Journal of Statistics-Revue Canadienne De Statistique Pub Date : 2023-02-11 DOI:10.1002/cjs.11762

Zhuojian Chen, Zhanfeng Wang, Yuan-chin Ivan Chang

{"title":"分布式顺序估计程序","authors":"Zhuojian Chen, Zhanfeng Wang, Yuan-chin Ivan Chang","doi":"10.1002/cjs.11762","DOIUrl":null,"url":null,"abstract":"<p>Data collected from distributed sources or sites commonly have different distributions or contaminated observations. Active learning procedures allow us to assess data when recruiting new data into model building. Thus, combining several active learning procedures together is a promising idea, even when the collected data set is contaminated. Here, we study how to conduct and integrate several adaptive sequential procedures at a time to produce a valid result via several machines or a parallel-computing framework. To avoid distraction by complicated modelling processes, we use confidence set estimation for linear models to illustrate the proposed method and discuss the approach's statistical properties. We then evaluate its performance using both synthetic and real data. We have implemented our method using Python and made it available through Github at https://github.com/zhuojianc/dsep.</p>","PeriodicalId":55281,"journal":{"name":"Canadian Journal of Statistics-Revue Canadienne De Statistique","volume":"52 1","pages":"271-290"},"PeriodicalIF":0.8000,"publicationDate":"2023-02-11","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":"{\"title\":\"Distributed sequential estimation procedures\",\"authors\":\"Zhuojian Chen, Zhanfeng Wang, Yuan-chin Ivan Chang\",\"doi\":\"10.1002/cjs.11762\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"<p>Data collected from distributed sources or sites commonly have different distributions or contaminated observations. Active learning procedures allow us to assess data when recruiting new data into model building. Thus, combining several active learning procedures together is a promising idea, even when the collected data set is contaminated. Here, we study how to conduct and integrate several adaptive sequential procedures at a time to produce a valid result via several machines or a parallel-computing framework. To avoid distraction by complicated modelling processes, we use confidence set estimation for linear models to illustrate the proposed method and discuss the approach's statistical properties. We then evaluate its performance using both synthetic and real data. We have implemented our method using Python and made it available through Github at https://github.com/zhuojianc/dsep.</p>\",\"PeriodicalId\":55281,\"journal\":{\"name\":\"Canadian Journal of Statistics-Revue Canadienne De Statistique\",\"volume\":\"52 1\",\"pages\":\"271-290\"},\"PeriodicalIF\":0.8000,\"publicationDate\":\"2023-02-11\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"0\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"Canadian Journal of Statistics-Revue Canadienne De Statistique\",\"FirstCategoryId\":\"100\",\"ListUrlMain\":\"https://onlinelibrary.wiley.com/doi/10.1002/cjs.11762\",\"RegionNum\":4,\"RegionCategory\":\"数学\",\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"Q3\",\"JCRName\":\"STATISTICS & PROBABILITY\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"Canadian Journal of Statistics-Revue Canadienne De Statistique","FirstCategoryId":"100","ListUrlMain":"https://onlinelibrary.wiley.com/doi/10.1002/cjs.11762","RegionNum":4,"RegionCategory":"数学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"Q3","JCRName":"STATISTICS & PROBABILITY","Score":null,"Total":0}

引用次数: 0

摘要

从分散的来源或地点收集到的数据通常具有不同的分布或受污染的观测结果。主动学习程序允许我们在收集新数据建立模型时对数据进行评估。因此，将多个主动学习程序结合在一起是一个很有前景的想法，即使收集到的数据集受到了污染。在这里，我们研究了如何通过多台机器或并行计算框架，同时进行并整合多个自适应序列程序，以产生有效的结果。为了避免复杂建模过程的干扰，我们使用线性模型的置信集估计来说明所提出的方法，并讨论该方法的统计特性。然后，我们使用合成数据和真实数据对其性能进行评估。我们使用 Python 实现了我们的方法，并通过 Github 发布在 https://github.com/zhuojianc/dsep 上。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

查看原文

微信好友朋友圈 QQ好友复制链接

本刊更多论文

Distributed sequential estimation procedures

Data collected from distributed sources or sites commonly have different distributions or contaminated observations. Active learning procedures allow us to assess data when recruiting new data into model building. Thus, combining several active learning procedures together is a promising idea, even when the collected data set is contaminated. Here, we study how to conduct and integrate several adaptive sequential procedures at a time to produce a valid result via several machines or a parallel-computing framework. To avoid distraction by complicated modelling processes, we use confidence set estimation for linear models to illustrate the proposed method and discuss the approach's statistical properties. We then evaluate its performance using both synthetic and real data. We have implemented our method using Python and made it available through Github at https://github.com/zhuojianc/dsep.

求助全文

通过发布文献求助，成功后即可免费获取论文全文。去求助

来源期刊

Canadian Journal of Statistics-Revue Canadienne De Statistique 数学-统计学与概率论

CiteScore

1.40

自引率

0.00%

发文量

审稿时长

>12 weeks

期刊介绍： The Canadian Journal of Statistics is the official journal of the Statistical Society of Canada. It has a reputation internationally as an excellent journal. The editorial board is comprised of statistical scientists with applied, computational, methodological, theoretical and probabilistic interests. Their role is to ensure that the journal continues to provide an international forum for the discipline of Statistics. The journal seeks papers making broad points of interest to many readers, whereas papers making important points of more specific interest are better placed in more specialized journals. The levels of innovation and impact are key in the evaluation of submitted manuscripts.

期刊最新文献

Issue Information True and false discoveries with independent and sequential e-values Issue Information Multiple change-point detection for regression curves Robust estimation of loss-based measures of model performance under covariate shift