Accelerating Number Theoretic Transformations for Bootstrappable Homomorphic Encryption on GPUs

2020 IEEE International Symposium on Workload Characterization (IISWC) Pub Date : 2020-10-01 DOI:10.1109/IISWC50251.2020.00033

Sangpyo Kim, Wonkyung Jung, J. Park, Jung Ho Ahn

{"title":"Accelerating Number Theoretic Transformations for Bootstrappable Homomorphic Encryption on GPUs","authors":"Sangpyo Kim, Wonkyung Jung, J. Park, Jung Ho Ahn","doi":"10.1109/IISWC50251.2020.00033","DOIUrl":null,"url":null,"abstract":"Homomorphic encryption (HE) draws huge attention as it provides a way of privacy-preserving computations on encrypted messages. Number Theoretic Transform (NTT), a specialized form of Discrete Fourier Transform (DFT) in the finite field of integers, is the key algorithm that enables fast computation on encrypted ciphertexts in HE. Prior works have accelerated NTT and its inverse transformation on a popular parallel processing platform, GPU, by leveraging DFT optimization techniques. However, these GPU-based studies lack a comprehensive analysis of the primary differences between NTT and DFT or only consider small HE parameters that have tight constraints in the number of arithmetic operations that can be performed without decryption. In this paper, we analyze the algorithmic characteristics of NTT and DFT and assess the performance of NTT when we apply the optimizations that are commonly applicable to both DFT and NTT on modern GPUs. From the analysis, we identify that NTT suffers from severe main-memory bandwidth bottleneck on large HE parameter sets. To tackle the main-memory bandwidth issue, we propose a novel NTT-specific on-the-fly root generation scheme dubbed on-the-fly twiddling (OT). Compared to the baseline radix-2 NTT implementation, after applying all the optimizations, including OT, we achieve 4.2⨯ speedup on a modern GPU.","PeriodicalId":365983,"journal":{"name":"2020 IEEE International Symposium on Workload Characterization (IISWC)","volume":"12 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2020-10-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"40","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"2020 IEEE International Symposium on Workload Characterization (IISWC)","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1109/IISWC50251.2020.00033","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 40

Abstract

Homomorphic encryption (HE) draws huge attention as it provides a way of privacy-preserving computations on encrypted messages. Number Theoretic Transform (NTT), a specialized form of Discrete Fourier Transform (DFT) in the finite field of integers, is the key algorithm that enables fast computation on encrypted ciphertexts in HE. Prior works have accelerated NTT and its inverse transformation on a popular parallel processing platform, GPU, by leveraging DFT optimization techniques. However, these GPU-based studies lack a comprehensive analysis of the primary differences between NTT and DFT or only consider small HE parameters that have tight constraints in the number of arithmetic operations that can be performed without decryption. In this paper, we analyze the algorithmic characteristics of NTT and DFT and assess the performance of NTT when we apply the optimizations that are commonly applicable to both DFT and NTT on modern GPUs. From the analysis, we identify that NTT suffers from severe main-memory bandwidth bottleneck on large HE parameter sets. To tackle the main-memory bandwidth issue, we propose a novel NTT-specific on-the-fly root generation scheme dubbed on-the-fly twiddling (OT). Compared to the baseline radix-2 NTT implementation, after applying all the optimizations, including OT, we achieve 4.2⨯ speedup on a modern GPU.

查看原文

微信好友朋友圈 QQ好友复制链接

本刊更多论文

gpu上可自引导同态加密的加速数论变换

同态加密(HE)为加密消息提供了一种保护隐私的计算方法，引起了人们的广泛关注。数论变换(NTT)是离散傅立叶变换(DFT)在有限整数域中的一种特殊形式，是实现HE中加密密文快速计算的关键算法。先前的工作通过利用DFT优化技术，加速了NTT及其在流行的并行处理平台GPU上的逆变换。然而，这些基于gpu的研究缺乏对NTT和DFT之间主要差异的全面分析，或者只考虑在没有解密的情况下可以执行的算术运算数量有严格限制的小HE参数。在本文中，我们分析了NTT和DFT的算法特征，并在现代gpu上应用通常适用于DFT和NTT的优化时评估了NTT的性能。从分析中，我们发现NTT在大HE参数集上存在严重的主存带宽瓶颈。为了解决主存带宽问题，我们提出了一种新的ntt特定的动态根生成方案，称为动态捻动(OT)。与基线基数-2 NTT实现相比，在应用所有优化(包括OT)后，我们在现代GPU上实现了4.2加速。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文去求助

来源期刊

2020 IEEE International Symposium on Workload Characterization (IISWC)

自引率

0.00%

发文量

期刊最新文献

Organizing Committee : IISWC 2020 Characterizing the impact of last-level cache replacement policies on big-data workloads AI on the Edge: Characterizing AI-based IoT Applications Using Specialized Edge Architectures Empirical Analysis and Modeling of Compute Times of CNN Operations on AWS Cloud Reliability Modeling of NISQ- Era Quantum Computers