{"title":"TT-HEALpix:用于大规模天文目录高效交叉匹配的新数据索引策略","authors":"Qing Zhao, Chengkui Zhang, Hao Li, Tingting Zhao, Chenzhou Cui, Dongwei Fan","doi":"10.1088/1538-3873/ad2721","DOIUrl":null,"url":null,"abstract":"\n Cross-matching is an indispensable operation in the data preparation, analysis, and research processes of multi-band astronomy and time-domain astronomy. Multi-catalog time-series data reconstruction is an important part of time-domain astronomy. In the large-scale distributed reconstruction process, boundary problems have always affected the accuracy of time-series data. To optimize these boundary problems and improve data precision, this paper proposes a new hybrid astronomical data indexing method called Translated Transformation based HEALPix Dual Index (TT-HEALPix). Under the reasonable Healpix division level, by translation transformation, the two indexes before and after the transformation form a unique pseudo-hybrid index strategy, which not only retains the advantages of the hybrid index scheme suitable for large-scale parallel computing, but also compensates for its shortage of high omission at the block boundary position. Based on TT-HEALPix, this paper completes the multi-catalog time-series reconstruction process on the Spark platform and compares it with the HEALPix+HTM hybrid indexing strategy. The experiments demonstrate that TT-HEALPix has significant advantages over the traditional HEALPix+HTM hybrid indexing method in terms of data accuracy and cross-matching efficiency. At level 9 of the Healpix index, TT-HEALPix achieves a 6%–19% improvement in cross-matching efficiency in a distributed environment compared to HEALPix+HTM. In terms of data accuracy, for the AST3-II dataset at level 9, TT-HEALPix has 62.2% accuracy improvement over HEALPix and 45.5% improvement over HEALPix+HTM. In conclusion, the proposed novel indexing strategy, TT-HEALPix, is better suited to the efficiency and accuracy requirements of cross-match.","PeriodicalId":20820,"journal":{"name":"Publications of the Astronomical Society of the Pacific","volume":null,"pages":null},"PeriodicalIF":3.3000,"publicationDate":"2024-03-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":"{\"title\":\"TT-HEALpix: A New Data Indexing Strategy for Efficient Cross-match of Large-scale Astronomical Catalogs\",\"authors\":\"Qing Zhao, Chengkui Zhang, Hao Li, Tingting Zhao, Chenzhou Cui, Dongwei Fan\",\"doi\":\"10.1088/1538-3873/ad2721\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"\\n Cross-matching is an indispensable operation in the data preparation, analysis, and research processes of multi-band astronomy and time-domain astronomy. Multi-catalog time-series data reconstruction is an important part of time-domain astronomy. In the large-scale distributed reconstruction process, boundary problems have always affected the accuracy of time-series data. To optimize these boundary problems and improve data precision, this paper proposes a new hybrid astronomical data indexing method called Translated Transformation based HEALPix Dual Index (TT-HEALPix). Under the reasonable Healpix division level, by translation transformation, the two indexes before and after the transformation form a unique pseudo-hybrid index strategy, which not only retains the advantages of the hybrid index scheme suitable for large-scale parallel computing, but also compensates for its shortage of high omission at the block boundary position. Based on TT-HEALPix, this paper completes the multi-catalog time-series reconstruction process on the Spark platform and compares it with the HEALPix+HTM hybrid indexing strategy. The experiments demonstrate that TT-HEALPix has significant advantages over the traditional HEALPix+HTM hybrid indexing method in terms of data accuracy and cross-matching efficiency. At level 9 of the Healpix index, TT-HEALPix achieves a 6%–19% improvement in cross-matching efficiency in a distributed environment compared to HEALPix+HTM. In terms of data accuracy, for the AST3-II dataset at level 9, TT-HEALPix has 62.2% accuracy improvement over HEALPix and 45.5% improvement over HEALPix+HTM. In conclusion, the proposed novel indexing strategy, TT-HEALPix, is better suited to the efficiency and accuracy requirements of cross-match.\",\"PeriodicalId\":20820,\"journal\":{\"name\":\"Publications of the Astronomical Society of the Pacific\",\"volume\":null,\"pages\":null},\"PeriodicalIF\":3.3000,\"publicationDate\":\"2024-03-01\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"0\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"Publications of the Astronomical Society of the Pacific\",\"FirstCategoryId\":\"101\",\"ListUrlMain\":\"https://doi.org/10.1088/1538-3873/ad2721\",\"RegionNum\":3,\"RegionCategory\":\"物理与天体物理\",\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"Q2\",\"JCRName\":\"ASTRONOMY & ASTROPHYSICS\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"Publications of the Astronomical Society of the Pacific","FirstCategoryId":"101","ListUrlMain":"https://doi.org/10.1088/1538-3873/ad2721","RegionNum":3,"RegionCategory":"物理与天体物理","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"Q2","JCRName":"ASTRONOMY & ASTROPHYSICS","Score":null,"Total":0}
TT-HEALpix: A New Data Indexing Strategy for Efficient Cross-match of Large-scale Astronomical Catalogs
Cross-matching is an indispensable operation in the data preparation, analysis, and research processes of multi-band astronomy and time-domain astronomy. Multi-catalog time-series data reconstruction is an important part of time-domain astronomy. In the large-scale distributed reconstruction process, boundary problems have always affected the accuracy of time-series data. To optimize these boundary problems and improve data precision, this paper proposes a new hybrid astronomical data indexing method called Translated Transformation based HEALPix Dual Index (TT-HEALPix). Under the reasonable Healpix division level, by translation transformation, the two indexes before and after the transformation form a unique pseudo-hybrid index strategy, which not only retains the advantages of the hybrid index scheme suitable for large-scale parallel computing, but also compensates for its shortage of high omission at the block boundary position. Based on TT-HEALPix, this paper completes the multi-catalog time-series reconstruction process on the Spark platform and compares it with the HEALPix+HTM hybrid indexing strategy. The experiments demonstrate that TT-HEALPix has significant advantages over the traditional HEALPix+HTM hybrid indexing method in terms of data accuracy and cross-matching efficiency. At level 9 of the Healpix index, TT-HEALPix achieves a 6%–19% improvement in cross-matching efficiency in a distributed environment compared to HEALPix+HTM. In terms of data accuracy, for the AST3-II dataset at level 9, TT-HEALPix has 62.2% accuracy improvement over HEALPix and 45.5% improvement over HEALPix+HTM. In conclusion, the proposed novel indexing strategy, TT-HEALPix, is better suited to the efficiency and accuracy requirements of cross-match.
期刊介绍:
The Publications of the Astronomical Society of the Pacific (PASP), the technical journal of the Astronomical Society of the Pacific (ASP), has been published regularly since 1889, and is an integral part of the ASP''s mission to advance the science of astronomy and disseminate astronomical information. The journal provides an outlet for astronomical results of a scientific nature and serves to keep readers in touch with current astronomical research. It contains refereed research and instrumentation articles, invited and contributed reviews, tutorials, and dissertation summaries.