K. Amaral, Zihan Li, W. Ding, S. Crouter, Ping Chen
{"title":"SummerTime: Variable-length Time Series Summarization with Application to Physical Activity Analysis","authors":"K. Amaral, Zihan Li, W. Ding, S. Crouter, Ping Chen","doi":"10.1145/3532628","DOIUrl":null,"url":null,"abstract":"SummerTime seeks to summarize global time-series signals and provides a fixed-length, robust representation of the variable-length time series. Many machine learning methods depend on data instances with a fixed number of features. As a result, those methods cannot be directly applied to variable-length time series data. Existing methods such as sliding windows can lose minority local information. Summarization conducted by the SummerTime method will be a fixed-length feature vector which can be used in place of the time series dataset for use with classical machine learning methods. We use Gaussian Mixture models (GMM) over small same-length disjoint windows in the time series to group local data into clusters. The time series’ rate of membership for each cluster will be a feature in the summarization. By making use of variational methods, GMM converges to a more robust mixture, meaning the clusters are more resistant to noise and overfitting. Further, the model is naturally capable of converging to an appropriate cluster count. We validate our method on a challenging real-world dataset, an imbalanced physical activity dataset with a variable-length time series structure. We compare our results to state-of-the-art studies and show high-quality improvement by classifying with only the summarization.","PeriodicalId":72043,"journal":{"name":"ACM transactions on computing for healthcare","volume":"3 1","pages":"1 - 15"},"PeriodicalIF":0.0000,"publicationDate":"2022-09-14","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"ACM transactions on computing for healthcare","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1145/3532628","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}
引用次数: 0
Abstract
SummerTime seeks to summarize global time-series signals and provides a fixed-length, robust representation of the variable-length time series. Many machine learning methods depend on data instances with a fixed number of features. As a result, those methods cannot be directly applied to variable-length time series data. Existing methods such as sliding windows can lose minority local information. Summarization conducted by the SummerTime method will be a fixed-length feature vector which can be used in place of the time series dataset for use with classical machine learning methods. We use Gaussian Mixture models (GMM) over small same-length disjoint windows in the time series to group local data into clusters. The time series’ rate of membership for each cluster will be a feature in the summarization. By making use of variational methods, GMM converges to a more robust mixture, meaning the clusters are more resistant to noise and overfitting. Further, the model is naturally capable of converging to an appropriate cluster count. We validate our method on a challenging real-world dataset, an imbalanced physical activity dataset with a variable-length time series structure. We compare our results to state-of-the-art studies and show high-quality improvement by classifying with only the summarization.