基于模型的重尾分布新简约混合聚类

IF 1.4 4区数学 Q2 STATISTICS & PROBABILITY Asta-Advances in Statistical Analysis Pub Date : 2022-01-14 DOI:10.1007/s10182-021-00430-8

Salvatore D. Tomarchio, Luca Bagnato, Antonio Punzo

{"title":"基于模型的重尾分布新简约混合聚类","authors":"Salvatore D. Tomarchio, Luca Bagnato, Antonio Punzo","doi":"10.1007/s10182-021-00430-8","DOIUrl":null,"url":null,"abstract":"<div><p>Two families of parsimonious mixture models are introduced for model-based clustering. They are based on two multivariate distributions-the shifted exponential normal and the tail-inflated normal-recently introduced in the literature as heavy-tailed generalizations of the multivariate normal. Parsimony is attained by the eigen-decomposition of the component scale matrices, as well as by the imposition of a constraint on the tailedness parameters. Identifiability conditions are also provided. Two variants of the expectation-maximization algorithm are presented for maximum likelihood parameter estimation. Parameter recovery and clustering performance are investigated via a simulation study. Comparisons with the unconstrained mixture models are obtained as by-product. A further simulated analysis is conducted to assess how sensitive our and some well-established parsimonious competitors are to their own generative scheme. Lastly, our and the competing models are evaluated in terms of fitting and clustering on three real datasets.</p></div>","PeriodicalId":55446,"journal":{"name":"Asta-Advances in Statistical Analysis","volume":"106 2","pages":"315 - 347"},"PeriodicalIF":1.4000,"publicationDate":"2022-01-14","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"6","resultStr":"{\"title\":\"Model-based clustering via new parsimonious mixtures of heavy-tailed distributions\",\"authors\":\"Salvatore D. Tomarchio, Luca Bagnato, Antonio Punzo\",\"doi\":\"10.1007/s10182-021-00430-8\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"<div><p>Two families of parsimonious mixture models are introduced for model-based clustering. They are based on two multivariate distributions-the shifted exponential normal and the tail-inflated normal-recently introduced in the literature as heavy-tailed generalizations of the multivariate normal. Parsimony is attained by the eigen-decomposition of the component scale matrices, as well as by the imposition of a constraint on the tailedness parameters. Identifiability conditions are also provided. Two variants of the expectation-maximization algorithm are presented for maximum likelihood parameter estimation. Parameter recovery and clustering performance are investigated via a simulation study. Comparisons with the unconstrained mixture models are obtained as by-product. A further simulated analysis is conducted to assess how sensitive our and some well-established parsimonious competitors are to their own generative scheme. Lastly, our and the competing models are evaluated in terms of fitting and clustering on three real datasets.</p></div>\",\"PeriodicalId\":55446,\"journal\":{\"name\":\"Asta-Advances in Statistical Analysis\",\"volume\":\"106 2\",\"pages\":\"315 - 347\"},\"PeriodicalIF\":1.4000,\"publicationDate\":\"2022-01-14\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"6\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"Asta-Advances in Statistical Analysis\",\"FirstCategoryId\":\"100\",\"ListUrlMain\":\"https://link.springer.com/article/10.1007/s10182-021-00430-8\",\"RegionNum\":4,\"RegionCategory\":\"数学\",\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"Q2\",\"JCRName\":\"STATISTICS & PROBABILITY\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"Asta-Advances in Statistical Analysis","FirstCategoryId":"100","ListUrlMain":"https://link.springer.com/article/10.1007/s10182-021-00430-8","RegionNum":4,"RegionCategory":"数学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"Q2","JCRName":"STATISTICS & PROBABILITY","Score":null,"Total":0}

引用次数: 6

摘要

引入了两类简约混合模型用于基于模型的聚类。它们基于两个多变量分布，即最近在文献中引入的移位指数正态和尾部膨胀正态，作为多变量正态的重尾推广。通过分量尺度矩阵的本征分解以及对尾性参数施加约束来获得简洁性。还提供了可识别性条件。针对最大似然参数估计，提出了期望最大化算法的两种变体。通过仿真研究研究了参数恢复和聚类性能。作为副产品，获得了与无约束混合物模型的比较。进行了进一步的模拟分析，以评估我们和一些公认的吝啬竞争对手对他们自己的生成方案的敏感程度。最后，我们和竞争模型在三个真实数据集上进行了拟合和聚类评估。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

查看原文

微信好友朋友圈 QQ好友复制链接

本刊更多论文

Model-based clustering via new parsimonious mixtures of heavy-tailed distributions

Two families of parsimonious mixture models are introduced for model-based clustering. They are based on two multivariate distributions-the shifted exponential normal and the tail-inflated normal-recently introduced in the literature as heavy-tailed generalizations of the multivariate normal. Parsimony is attained by the eigen-decomposition of the component scale matrices, as well as by the imposition of a constraint on the tailedness parameters. Identifiability conditions are also provided. Two variants of the expectation-maximization algorithm are presented for maximum likelihood parameter estimation. Parameter recovery and clustering performance are investigated via a simulation study. Comparisons with the unconstrained mixture models are obtained as by-product. A further simulated analysis is conducted to assess how sensitive our and some well-established parsimonious competitors are to their own generative scheme. Lastly, our and the competing models are evaluated in terms of fitting and clustering on three real datasets.

求助全文

通过发布文献求助，成功后即可免费获取论文全文。去求助

来源期刊

Asta-Advances in Statistical Analysis 数学-统计学与概率论

CiteScore

2.20

自引率

14.30%

发文量

审稿时长

>12 weeks

期刊介绍： AStA - Advances in Statistical Analysis, a journal of the German Statistical Society, is published quarterly and presents original contributions on statistical methods and applications and review articles.

期刊最新文献

A note on dynamic spatiotemporal ARCH models: small- and large-sample results Advances in spatial econometrics and geostatistics: methods, theory, and applications Fuzzy C-modes clustering with spatial regularization and noise cluster Margin-closed regime-switching multivariate time series models Editorial