利用Hadoop在某电力公司进行时间序列搜索

BigMine '13 Pub Date : 2013-08-11 DOI:10.1145/2501221.2501224

Alice Berard, G. Hébrail

{"title":"利用Hadoop在某电力公司进行时间序列搜索","authors":"Alice Berard, G. Hébrail","doi":"10.1145/2501221.2501224","DOIUrl":null,"url":null,"abstract":"In this paper, we investigate the possibilities offered by the Hadoop eco-system for searching time series in an electric power company (Top-K or range-queries based on a similarity measure). There has been much work done on speeding up the search of time series in a large dataset, mainly by designing efficient indexing techniques preceded by reduction techniques. In this paper, we do not follow these approaches but focus on using the brutal force of distributed computations in the Hadoop environment. We propose an implementation of time series search functions in Hadoop and describe experiments on a large database of electric power consumption curves (35M customers observed during 1 month at a 30' sampling rate). We also show that this architecture supports easily the computation of several distances for the same query with a small response time overhead: this is very useful in practice when the end-user does not know very well which distance to use.","PeriodicalId":441216,"journal":{"name":"BigMine '13","volume":"17 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2013-08-11","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"13","resultStr":"{\"title\":\"Searching time series with Hadoop in an electric power company\",\"authors\":\"Alice Berard, G. Hébrail\",\"doi\":\"10.1145/2501221.2501224\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"In this paper, we investigate the possibilities offered by the Hadoop eco-system for searching time series in an electric power company (Top-K or range-queries based on a similarity measure). There has been much work done on speeding up the search of time series in a large dataset, mainly by designing efficient indexing techniques preceded by reduction techniques. In this paper, we do not follow these approaches but focus on using the brutal force of distributed computations in the Hadoop environment. We propose an implementation of time series search functions in Hadoop and describe experiments on a large database of electric power consumption curves (35M customers observed during 1 month at a 30' sampling rate). We also show that this architecture supports easily the computation of several distances for the same query with a small response time overhead: this is very useful in practice when the end-user does not know very well which distance to use.\",\"PeriodicalId\":441216,\"journal\":{\"name\":\"BigMine '13\",\"volume\":\"17 1\",\"pages\":\"0\"},\"PeriodicalIF\":0.0000,\"publicationDate\":\"2013-08-11\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"13\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"BigMine '13\",\"FirstCategoryId\":\"1085\",\"ListUrlMain\":\"https://doi.org/10.1145/2501221.2501224\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"\",\"JCRName\":\"\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"BigMine '13","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1145/2501221.2501224","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 13

摘要

在本文中，我们研究了Hadoop生态系统为在电力公司中搜索时间序列(Top-K或基于相似性度量的范围查询)提供的可能性。在加速大型数据集中时间序列的搜索方面已经做了很多工作，主要是通过在约简技术之前设计高效的索引技术。在本文中，我们不遵循这些方法，而是专注于在Hadoop环境中使用分布式计算的残酷力量。我们提出了在Hadoop中实现时间序列搜索功能，并描述了在一个大型电力消耗曲线数据库上的实验(以30'的采样率在1个月内观察到35M客户)。我们还展示了该体系结构支持以很小的响应时间开销轻松地计算相同查询的多个距离:当最终用户不太清楚该使用哪个距离时，这在实践中非常有用。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

查看原文

微信好友朋友圈 QQ好友复制链接

本刊更多论文

Searching time series with Hadoop in an electric power company

In this paper, we investigate the possibilities offered by the Hadoop eco-system for searching time series in an electric power company (Top-K or range-queries based on a similarity measure). There has been much work done on speeding up the search of time series in a large dataset, mainly by designing efficient indexing techniques preceded by reduction techniques. In this paper, we do not follow these approaches but focus on using the brutal force of distributed computations in the Hadoop environment. We propose an implementation of time series search functions in Hadoop and describe experiments on a large database of electric power consumption curves (35M customers observed during 1 month at a 30' sampling rate). We also show that this architecture supports easily the computation of several distances for the same query with a small response time overhead: this is very useful in practice when the end-user does not know very well which distance to use.

求助全文

通过发布文献求助，成功后即可免费获取论文全文。去求助

来源期刊

BigMine '13

自引率

0.00%

发文量

期刊最新文献

Forecasting building occupancy using sensor network data Maintaining connected components for infinite graph streams Soft-CsGDT: soft cost-sensitive Gaussian decision tree for cost-sensitive classification of data streams Data-driven study of urban infrastructure to enable city-wide ubiquitous computing Big & personal: data and models behind netflix recommendations