IEEE Journal on Emerging and Selected Topics in Circuits and Systems最新文献

英文中文

IEEE Circuits and Systems Society Information 电气和电子工程师学会电路与系统协会信息

IF 3.7 2区工程技术 Q2 ENGINEERING, ELECTRICAL & ELECTRONIC

IEEE Journal on Emerging and Selected Topics in Circuits and Systems

Pub Date : 2024-09-16 DOI: 10.1109/JETCAS.2024.3450049

引用次数: 0

IEEE Journal on Emerging and Selected Topics in Circuits and Systems Publication Information 电气和电子工程师学会电路与系统新专题与选题期刊》出版信息

IF 3.7 2区工程技术 Q2 ENGINEERING, ELECTRICAL & ELECTRONIC

IEEE Journal on Emerging and Selected Topics in Circuits and Systems

Pub Date : 2024-09-16 DOI: 10.1109/JETCAS.2024.3450055

引用次数: 0

Guest Editorial Chip and Package-Scale Communication-Aware Architectures for General-Purpose, Domain-Specific, and Quantum Computing Systems 特邀编辑通用、特定领域和量子计算系统的芯片和封装级通信感知架构

IF 3.7 2区工程技术 Q2 ENGINEERING, ELECTRICAL & ELECTRONIC

IEEE Journal on Emerging and Selected Topics in Circuits and Systems

Pub Date : 2024-09-16 DOI: 10.1109/JETCAS.2024.3445208

Abhijit Das;Maurizio Palesi;John Kim;Partha Pratim Pande

This Special Issue of IEEE Journal on Emerging and Selected Topics in Circuits and Systems (JETCAS) is devoted to advancing the field of chip and package-scale communications across diverse computing domains, bridging academic research and industrial innovation. As we enter a new golden age of computer architecture, marked by both challenges and opportunities, the anticipated end of Moore’s law necessitates reimagining the future of computing systems as we approach the physical limits of transistors. Three leading approaches to address these challenges include the chiplet paradigm, domain-specific customization, and quantum computing. However, these architectural and technological innovations have shifted the primary bottleneck from computation to communication. Consequently, on-chip and on-package communication now play a critical role in determining the performance, efficiency, and scalability of general-purpose, domain-specific, and quantum computing systems. Their ever-growing importance has garnered significant attention from both academia and industry.

本期《IEEE 电路与系统新兴选题期刊》（JETCAS）特刊致力于推动不同计算领域的芯片和封装级通信领域的发展，为学术研究和产业创新搭建桥梁。随着我们进入计算机体系结构的新黄金时代，挑战与机遇并存，摩尔定律的预期终结要求我们在接近晶体管物理极限时重新构想计算系统的未来。应对这些挑战的三种主要方法包括芯片范式、特定领域定制和量子计算。然而，这些架构和技术创新将主要瓶颈从计算转移到了通信。因此，片上和封装上通信现在在决定通用、特定领域和量子计算系统的性能、效率和可扩展性方面发挥着至关重要的作用。它们的重要性与日俱增，引起了学术界和产业界的极大关注。

{"title":"Guest Editorial Chip and Package-Scale Communication-Aware Architectures for General-Purpose, Domain-Specific, and Quantum Computing Systems","authors":"Abhijit Das;Maurizio Palesi;John Kim;Partha Pratim Pande","doi":"10.1109/JETCAS.2024.3445208","DOIUrl":"https://doi.org/10.1109/JETCAS.2024.3445208","url":null,"abstract":"This Special Issue of IEEE Journal on Emerging and Selected Topics in Circuits and Systems (JETCAS) is devoted to advancing the field of chip and package-scale communications across diverse computing domains, bridging academic research and industrial innovation. As we enter a new golden age of computer architecture, marked by both challenges and opportunities, the anticipated end of Moore’s law necessitates reimagining the future of computing systems as we approach the physical limits of transistors. Three leading approaches to address these challenges include the chiplet paradigm, domain-specific customization, and quantum computing. However, these architectural and technological innovations have shifted the primary bottleneck from computation to communication. Consequently, on-chip and on-package communication now play a critical role in determining the performance, efficiency, and scalability of general-purpose, domain-specific, and quantum computing systems. Their ever-growing importance has garnered significant attention from both academia and industry.","PeriodicalId":48827,"journal":{"name":"IEEE Journal on Emerging and Selected Topics in Circuits and Systems","volume":"14 3","pages":"349-353"},"PeriodicalIF":3.7,"publicationDate":"2024-09-16","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://ieeexplore.ieee.org/stamp/stamp.jsp?tp=&arnumber=10680692","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"142235954","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"工程技术","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}

引用次数: 0

IEEE Journal on Emerging and Selected Topics in Circuits and Systems Information for Authors IEEE 《电路与系统新兴选题》期刊作者须知

IF 3.7 2区工程技术 Q2 ENGINEERING, ELECTRICAL & ELECTRONIC

IEEE Journal on Emerging and Selected Topics in Circuits and Systems

Pub Date : 2024-09-16 DOI: 10.1109/JETCAS.2024.3450053

引用次数: 0

Modeling the Effect of SEUs on the Configuration Memory of SRAM-FPGA-Based CNN Accelerators 模拟 SEU 对基于 SRAM-FPGA 的 CNN 加速器配置存储器的影响

IF 3.7 2区工程技术 Q2 ENGINEERING, ELECTRICAL & ELECTRONIC

IEEE Journal on Emerging and Selected Topics in Circuits and Systems

Pub Date : 2024-09-16 DOI: 10.1109/JETCAS.2024.3460792

Zhen Gao;Jiaqi Feng;Shihui Gao;Qiang Liu;Guangjun Ge;Yu Wang;Pedro Reviriego

Convolutional Neural Networks (CNNs) are widely used in computer vision applications. SRAM based Field Programmable Gate Arrays (SRAM-FPGAs) are popular for the acceleration of CNNs. Since SRAM-FPGAs are prone to soft errors, the reliability evaluation and efficient fault tolerance design become very important for the use of FPGA-based CNNs in safety critical scenarios. Hardware based fault injection is an effective approach for the reliability evaluation, and the results can provide valuable references for the fault tolerance design. However, the complexity of building a fault injection platform poses a big obstacle for researchers working on the fault tolerance design. To remove this obstacle, this paper first performs a complete reliability evaluation for errors on the configuration memory of the FPGA based CNN accelerators, and then studies the impact of errors on the output feature maps of each layer. Based on the statistical analysis, we propose several fault models for the effect of SEUs on the configuration memory of the FPGA based CNN accelerators, and build a software simulator based on the fault models. Experiments show that the evaluation results based on the software simulator are very close to those from the hardware fault injections. Therefore, the proposed fault models and simulator can facilitate the fault tolerance design and reliability evaluation of CNN accelerators.

卷积神经网络（cnn）在计算机视觉领域有着广泛的应用。基于SRAM的现场可编程门阵列（SRAM- fpga）被广泛用于cnn的加速。由于sram - fpga容易出现软错误，因此可靠性评估和高效容错设计对于在安全关键场景下使用基于fpga的cnn变得非常重要。基于硬件的故障注入是一种有效的可靠性评估方法，其结果可为容错设计提供有价值的参考。然而，构建故障注入平台的复杂性给容错设计带来了很大的障碍。为了消除这一障碍，本文首先对基于FPGA的CNN加速器配置存储器的错误进行了完整的可靠性评估，然后研究了错误对各层输出特征映射的影响。在统计分析的基础上，提出了几种seu对FPGA CNN加速器组态内存影响的故障模型，并基于这些故障模型构建了软件仿真器。实验表明，基于软件模拟器的评估结果与硬件故障注入的评估结果非常接近。因此，所提出的故障模型和仿真器可以方便地进行CNN加速器的容错设计和可靠性评估。

{"title":"Modeling the Effect of SEUs on the Configuration Memory of SRAM-FPGA-Based CNN Accelerators","authors":"Zhen Gao;Jiaqi Feng;Shihui Gao;Qiang Liu;Guangjun Ge;Yu Wang;Pedro Reviriego","doi":"10.1109/JETCAS.2024.3460792","DOIUrl":"https://doi.org/10.1109/JETCAS.2024.3460792","url":null,"abstract":"Convolutional Neural Networks (CNNs) are widely used in computer vision applications. SRAM based Field Programmable Gate Arrays (SRAM-FPGAs) are popular for the acceleration of CNNs. Since SRAM-FPGAs are prone to soft errors, the reliability evaluation and efficient fault tolerance design become very important for the use of FPGA-based CNNs in safety critical scenarios. Hardware based fault injection is an effective approach for the reliability evaluation, and the results can provide valuable references for the fault tolerance design. However, the complexity of building a fault injection platform poses a big obstacle for researchers working on the fault tolerance design. To remove this obstacle, this paper first performs a complete reliability evaluation for errors on the configuration memory of the FPGA based CNN accelerators, and then studies the impact of errors on the output feature maps of each layer. Based on the statistical analysis, we propose several fault models for the effect of SEUs on the configuration memory of the FPGA based CNN accelerators, and build a software simulator based on the fault models. Experiments show that the evaluation results based on the software simulator are very close to those from the hardware fault injections. Therefore, the proposed fault models and simulator can facilitate the fault tolerance design and reliability evaluation of CNN accelerators.","PeriodicalId":48827,"journal":{"name":"IEEE Journal on Emerging and Selected Topics in Circuits and Systems","volume":"14 4","pages":"799-810"},"PeriodicalIF":3.7,"publicationDate":"2024-09-16","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"142821273","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"工程技术","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}

引用次数: 0

Stealthy Backdoor Attack Against Federated Learning Through Frequency Domain by Backdoor Neuron Constraint and Model Camouflage 利用后门神经元约束和模型伪装，通过频域对联合学习进行隐形后门攻击

IF 3.7 2区工程技术 Q2 ENGINEERING, ELECTRICAL & ELECTRONIC

IEEE Journal on Emerging and Selected Topics in Circuits and Systems

Pub Date : 2024-08-27 DOI: 10.1109/JETCAS.2024.3450527

Yanqi Qiao;Dazhuang Liu;Rui Wang;Kaitai Liang

Federated Learning (FL) is a beneficial decentralized learning approach for preserving the privacy of local datasets of distributed agents. However, the distributed property of FL and untrustworthy data introducing the vulnerability to backdoor attacks. In this attack scenario, an adversary manipulates its local data with a specific trigger and trains a malicious local model to implant the backdoor. During inference, the global model would misbehave for any input with the trigger to the attacker-chosen prediction. Most existing backdoor attacks against FL focus on bypassing defense mechanisms, without considering the inspection of model parameters on the server. These attacks are susceptible to detection through dynamic clustering based on model parameter similarity. Besides, current methods provide limited imperceptibility of their trigger in the spatial domain. To address these limitations, we propose a stealthy backdoor attack called “Chironex” against FL with an imperceptible trigger in frequency space to deliver attack effectiveness, stealthiness and robustness against various countermeasures on FL. We first design a frequency trigger function to generate an imperceptible frequency trigger to evade human inspection. Then we fully exploit the attacker’s advantage to enhance attack robustness by estimating benign updates and analyzing the impact of the backdoor on model parameters through a task-sensitive neuron searcher. It disguises malicious updates as benign ones by reducing the impact of backdoor neurons that greatly contribute to the backdoor task based on activation value, and encouraging them to update towards benign model parameters trained by the attacker. We conduct extensive experiments on various image classifiers with real-world datasets to provide empirical evidence that Chironex can evade the most recent robust FL aggregation algorithms, and further achieve a distinctly higher attack success rate than existing attacks, without undermining the utility of the global model.

联邦学习（FL）是一种有益的分散学习方法，用于保护分布式代理的本地数据集的隐私性。然而，FL的分布式特性和不可信数据引入了后门攻击的漏洞。在此攻击场景中，攻击者使用特定触发器操纵其本地数据，并训练恶意本地模型来植入后门。在推理过程中，全局模型会对任何带有攻击者选择的预测触发器的输入产生错误行为。大多数针对FL的后门攻击都侧重于绕过防御机制，而没有考虑对服务器上的模型参数进行检查。基于模型参数相似度的动态聚类容易检测到这些攻击。此外，目前的方法提供了有限的不可感知的触发空间域。为了解决这些限制，我们提出了一种名为“Chironex”的针对FL的隐形后门攻击，该攻击在频率空间中具有不可察觉的触发，以提供针对FL的各种对策的攻击有效性，隐潜性和鲁棒性。我们首先设计了一个频率触发函数来生成不可察觉的频率触发以逃避人类检查。然后通过任务敏感神经元搜索器估计良性更新和分析后门对模型参数的影响，充分利用攻击者的优势增强攻击的鲁棒性。它通过减少基于激活值对后门任务贡献巨大的后门神经元的影响，并鼓励它们向攻击者训练的良性模型参数更新，将恶意更新伪装成良性更新。我们用真实世界的数据集对各种图像分类器进行了广泛的实验，以提供经验证据，证明Chironex可以逃避最新的鲁棒FL聚合算法，并进一步实现比现有攻击明显更高的攻击成功率，而不会破坏全局模型的效用。

{"title":"Stealthy Backdoor Attack Against Federated Learning Through Frequency Domain by Backdoor Neuron Constraint and Model Camouflage","authors":"Yanqi Qiao;Dazhuang Liu;Rui Wang;Kaitai Liang","doi":"10.1109/JETCAS.2024.3450527","DOIUrl":"10.1109/JETCAS.2024.3450527","url":null,"abstract":"Federated Learning (FL) is a beneficial decentralized learning approach for preserving the privacy of local datasets of distributed agents. However, the distributed property of FL and untrustworthy data introducing the vulnerability to backdoor attacks. In this attack scenario, an adversary manipulates its local data with a specific trigger and trains a malicious local model to implant the backdoor. During inference, the global model would misbehave for any input with the trigger to the attacker-chosen prediction. Most existing backdoor attacks against FL focus on bypassing defense mechanisms, without considering the inspection of model parameters on the server. These attacks are susceptible to detection through dynamic clustering based on model parameter similarity. Besides, current methods provide limited imperceptibility of their trigger in the spatial domain. To address these limitations, we propose a stealthy backdoor attack called “Chironex” against FL with an imperceptible trigger in frequency space to deliver attack effectiveness, stealthiness and robustness against various countermeasures on FL. We first design a frequency trigger function to generate an imperceptible frequency trigger to evade human inspection. Then we fully exploit the attacker’s advantage to enhance attack robustness by estimating benign updates and analyzing the impact of the backdoor on model parameters through a task-sensitive neuron searcher. It disguises malicious updates as benign ones by reducing the impact of backdoor neurons that greatly contribute to the backdoor task based on activation value, and encouraging them to update towards benign model parameters trained by the attacker. We conduct extensive experiments on various image classifiers with real-world datasets to provide empirical evidence that Chironex can evade the most recent robust FL aggregation algorithms, and further achieve a distinctly higher attack success rate than existing attacks, without undermining the utility of the global model.","PeriodicalId":48827,"journal":{"name":"IEEE Journal on Emerging and Selected Topics in Circuits and Systems","volume":"14 4","pages":"661-672"},"PeriodicalIF":3.7,"publicationDate":"2024-08-27","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"142197276","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"工程技术","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}

引用次数: 0

Chip and Package-Scale Interconnects for General-Purpose, Domain-Specific, and Quantum Computing Systems—Overview, Challenges, and Opportunities 通用、特定领域和量子计算系统的芯片和封装级互连 - 概述、挑战和机遇

IF 3.7 2区工程技术 Q2 ENGINEERING, ELECTRICAL & ELECTRONIC

IEEE Journal on Emerging and Selected Topics in Circuits and Systems

Pub Date : 2024-08-19 DOI: 10.1109/JETCAS.2024.3445829

Abhijit Das;Maurizio Palesi;John Kim;Partha Pratim Pande

The anticipated end of Moore’s law, coupled with the breakdown of Dennard scaling, compelled everyone to conceive forthcoming computing systems once transistors reach their limits. Three leading approaches to circumvent this situation are the chiplet paradigm, domain customisation and quantum computing. However, architectural and technological innovations have shifted the fundamental bottleneck from computation to communication. Hence, on-chip and on-package communication play a pivotal role in determining the performance, energy efficiency and scalability of general-purpose, domain-specific and quantum computing systems. This article reviews the recent advances in chip and package-scale interconnects due to the change in architecture, application and technology. The primary objective of this article is to present the current status, key challenges, and impact-worthy opportunities in this research area from the perspective of hardware architectures. The secondary objective of this article is to serve as a tutorial providing an overview of academic and industrial explorations in chip and package-scale communication infrastructure design for general-purpose, domain-specific and quantum computing systems.

摩尔定律的预期终结，再加上邓纳缩放的崩溃，迫使每个人都在构想晶体管达到极限后即将出现的计算系统。规避这种情况的三种主要方法是芯片范式、领域定制和量子计算。然而，架构和技术创新已将根本瓶颈从计算转移到通信。因此，片上和封装上通信在决定通用、特定领域和量子计算系统的性能、能效和可扩展性方面发挥着举足轻重的作用。本文回顾了由于架构、应用和技术的变化而导致的芯片级和封装级互连的最新进展。本文的主要目的是从硬件架构的角度介绍这一研究领域的现状、关键挑战和有影响的机遇。本文的次要目的是作为教程，概述学术界和工业界在通用、特定领域和量子计算系统的芯片和封装级通信基础设施设计方面的探索。

{"title":"Chip and Package-Scale Interconnects for General-Purpose, Domain-Specific, and Quantum Computing Systems—Overview, Challenges, and Opportunities","authors":"Abhijit Das;Maurizio Palesi;John Kim;Partha Pratim Pande","doi":"10.1109/JETCAS.2024.3445829","DOIUrl":"10.1109/JETCAS.2024.3445829","url":null,"abstract":"The anticipated end of Moore’s law, coupled with the breakdown of Dennard scaling, compelled everyone to conceive forthcoming computing systems once transistors reach their limits. Three leading approaches to circumvent this situation are the chiplet paradigm, domain customisation and quantum computing. However, architectural and technological innovations have shifted the fundamental bottleneck from computation to communication. Hence, on-chip and on-package communication play a pivotal role in determining the performance, energy efficiency and scalability of general-purpose, domain-specific and quantum computing systems. This article reviews the recent advances in chip and package-scale interconnects due to the change in architecture, application and technology. The primary objective of this article is to present the current status, key challenges, and impact-worthy opportunities in this research area from the perspective of hardware architectures. The secondary objective of this article is to serve as a tutorial providing an overview of academic and industrial explorations in chip and package-scale communication infrastructure design for general-purpose, domain-specific and quantum computing systems.","PeriodicalId":48827,"journal":{"name":"IEEE Journal on Emerging and Selected Topics in Circuits and Systems","volume":"14 3","pages":"354-370"},"PeriodicalIF":3.7,"publicationDate":"2024-08-19","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://ieeexplore.ieee.org/stamp/stamp.jsp?tp=&arnumber=10638543","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"142197237","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"工程技术","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"OA","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}

引用次数: 0

GNS: Graph-Based Network-on-Chip Shield for Early Defense Against Malicious Nodes in MPSoC GNS：基于图的片上网络屏蔽，用于早期防御 MPSoC 中的恶意节点

IF 3.7 2区工程技术 Q2 ENGINEERING, ELECTRICAL & ELECTRONIC

IEEE Journal on Emerging and Selected Topics in Circuits and Systems

Pub Date : 2024-08-05 DOI: 10.1109/JETCAS.2024.3438435

Haoyu Wang;Jianjie Ren;Basel Halak;Ahmad Atamli

In the rapidly evolving landscape of system design, Multi-Processor Systems-on-Chip (MPSoCs) have experienced significant growth in both scale and complexity, by integrating an array of Intellectual Properties (IPs) through Network-on-Chip (NoC) to execute complex parallel applications. However, this advancement has led to the emergence of security attacks caused by Malicious Third-Party IPs (M3PIPs), such as Denial-of-Service (DoS). Many current methods for detecting DoS attacks involve significant hardware overhead and are often inefficient in identifying anomalies at an early stage. Addressing this gap, we propose the Graph-based NoC Shield (GNS), a robust strategy meticulously crafted to detect, localize, and isolate malicious IPs at the very early stage of DoS appearance. Central to our approach is the use of a Graph Neural Network (GNN) and Long Short-Term Memory (LSTM) detection model. This combination capitalizes on network traffic data and routing dependency graphs to efficiently trace the source of network congestion and pinpoint attackers. Our extensive experimental analysis validates the effectiveness of the GNS framework, demonstrating a 98% detection accuracy and localization capabilities, achieved with minimal hardware overhead of 1.8% in each router, based on a pure 4*4 Mesh NoC system. The detection performance exceeds that of all other state-of-the-art works and most straightforward single machine learning inference models within the same context. Additionally, the hardware overhead is notably superior compared to other security schemes. Another key feature of our system is the implementation of a credit interposing mechanism. It was specifically designed to isolate M3PIPs engaging in Flooding-based DoS and effectively mitigate the spread of malicious traffic. This approach significantly enhances the security of NoC-based MPSoCs, offering early-stage detection with the superior accuracy compared to other models. Crucially, the GNS achieves this with up to 75% less hardware overhead than state-of-the-art solutions, thus striking a balance between efficiency and effectiveness in security implementation.

在快速发展的系统设计领域，多处理器片上系统（MPSoC）通过片上网络（NoC）集成了一系列知识产权（IP）以执行复杂的并行应用，其规模和复杂性都有了显著提高。然而，这一进步也导致了恶意第三方 IP（M3PIP）引起的安全攻击的出现，如拒绝服务（DoS）。目前许多检测 DoS 攻击的方法都涉及大量硬件开销，而且在早期识别异常情况方面往往效率低下。针对这一缺陷，我们提出了基于图形的 NoC 屏蔽（GNS），这是一种精心设计的强大策略，可在 DoS 出现的早期阶段检测、定位和隔离恶意 IP。我们方法的核心是使用图形神经网络（GNN）和长短期记忆（LSTM）检测模型。这一组合利用了网络流量数据和路由依赖图，可有效追踪网络拥塞的源头并精确定位攻击者。我们的大量实验分析验证了 GNS 框架的有效性，基于纯 4*4 网状 NoC 系统，每个路由器的硬件开销仅为 1.8%，却实现了 98% 的检测准确率和定位能力。其检测性能超过了所有其他最先进的研究成果和相同背景下最直接的单一机器学习推理模型。此外，硬件开销也明显优于其他安全方案。我们系统的另一个主要特点是实施了一种信用穿插机制。该机制专门用于隔离参与基于泛洪的 DoS 的 M3PIP，并有效缓解恶意流量的传播。这种方法大大增强了基于 NoC 的 MPSoC 的安全性，与其他模型相比，它能提供准确性更高的早期检测。最重要的是，与最先进的解决方案相比，GNS 可减少高达 75% 的硬件开销，从而在安全实施的效率和效果之间取得了平衡。

{"title":"GNS: Graph-Based Network-on-Chip Shield for Early Defense Against Malicious Nodes in MPSoC","authors":"Haoyu Wang;Jianjie Ren;Basel Halak;Ahmad Atamli","doi":"10.1109/JETCAS.2024.3438435","DOIUrl":"10.1109/JETCAS.2024.3438435","url":null,"abstract":"In the rapidly evolving landscape of system design, Multi-Processor Systems-on-Chip (MPSoCs) have experienced significant growth in both scale and complexity, by integrating an array of Intellectual Properties (IPs) through Network-on-Chip (NoC) to execute complex parallel applications. However, this advancement has led to the emergence of security attacks caused by Malicious Third-Party IPs (M3PIPs), such as Denial-of-Service (DoS). Many current methods for detecting DoS attacks involve significant hardware overhead and are often inefficient in identifying anomalies at an early stage. Addressing this gap, we propose the Graph-based NoC Shield (GNS), a robust strategy meticulously crafted to detect, localize, and isolate malicious IPs at the very early stage of DoS appearance. Central to our approach is the use of a Graph Neural Network (GNN) and Long Short-Term Memory (LSTM) detection model. This combination capitalizes on network traffic data and routing dependency graphs to efficiently trace the source of network congestion and pinpoint attackers. Our extensive experimental analysis validates the effectiveness of the GNS framework, demonstrating a 98% detection accuracy and localization capabilities, achieved with minimal hardware overhead of 1.8% in each router, based on a pure 4*4 Mesh NoC system. The detection performance exceeds that of all other state-of-the-art works and most straightforward single machine learning inference models within the same context. Additionally, the hardware overhead is notably superior compared to other security schemes. Another key feature of our system is the implementation of a credit interposing mechanism. It was specifically designed to isolate M3PIPs engaging in Flooding-based DoS and effectively mitigate the spread of malicious traffic. This approach significantly enhances the security of NoC-based MPSoCs, offering early-stage detection with the superior accuracy compared to other models. Crucially, the GNS achieves this with up to 75% less hardware overhead than state-of-the-art solutions, thus striking a balance between efficiency and effectiveness in security implementation.","PeriodicalId":48827,"journal":{"name":"IEEE Journal on Emerging and Selected Topics in Circuits and Systems","volume":"14 3","pages":"483-494"},"PeriodicalIF":3.7,"publicationDate":"2024-08-05","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"141934792","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"工程技术","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}

引用次数: 0

Investigating Register Cache Behavior: Implications for CUDA and Tensor Core Workloads on GPUs 研究寄存器缓存行为：对 GPU 上 CUDA 和张量核心工作负载的影响

IF 3.7 2区工程技术 Q2 ENGINEERING, ELECTRICAL & ELECTRONIC

IEEE Journal on Emerging and Selected Topics in Circuits and Systems

Pub Date : 2024-08-05 DOI: 10.1109/JETCAS.2024.3439193

Vahid Geraeinejad;Qiran Qian;Masoumeh Ebrahimi

GPUs are extensively employed as the primary devices for running a broad spectrum of applications, covering general-purpose applications as well as Artificial Intelligence (AI) applications. Register file, as the largest SRAM on the GPU die, accounts for over 20% of the total GPU energy consumption. Register cache has been introduced to reduce traffic from the register file and thus decrease total energy consumption when CUDA cores are utilized. However, the utilization of register cache has not been thoroughly investigated for Tensor Cores which are integrated into recent GPU architectures to meet AI workload demands. In this paper, we study the usage of register cache in both CUDA and Tensor Cores and conduct a thorough examination of their pros and cons. We have developed an open-source analytical simulator, called RFC-sim, to model and measure the energy consumption of both the register file and register cache. Our results show that while the register cache can reduce energy consumption by up to 40% in CUDA cores, it results in increased energy consumption by up to 23% in Tensor Cores. The main reason lies in the limited space of the register cache, which is not sufficient for the demand of Tensor cores to capture locality.

GPU 被广泛用作运行各种应用的主要设备，包括通用应用和人工智能（AI）应用。寄存器文件是 GPU 芯片上最大的 SRAM，占 GPU 总能耗的 20% 以上。引入寄存器缓存是为了减少寄存器文件的流量，从而降低使用 CUDA 内核时的总能耗。然而，对于集成到最近的 GPU 架构中以满足人工智能工作负载需求的张量核，寄存器缓存的利用率尚未得到深入研究。在本文中，我们研究了寄存器缓存在 CUDA 和 Tensor Cores 中的使用情况，并对它们的利弊进行了深入探讨。我们开发了一个名为 RFC-sim 的开源分析模拟器，对寄存器文件和寄存器缓存的能耗进行建模和测量。我们的结果表明，虽然寄存器缓存在 CUDA 内核中可以减少高达 40% 的能耗，但在 Tensor 内核中却会导致能耗增加高达 23%。主要原因在于寄存器缓存的空间有限，无法满足 Tensor 内核捕捉定位的需求。

{"title":"Investigating Register Cache Behavior: Implications for CUDA and Tensor Core Workloads on GPUs","authors":"Vahid Geraeinejad;Qiran Qian;Masoumeh Ebrahimi","doi":"10.1109/JETCAS.2024.3439193","DOIUrl":"10.1109/JETCAS.2024.3439193","url":null,"abstract":"GPUs are extensively employed as the primary devices for running a broad spectrum of applications, covering general-purpose applications as well as Artificial Intelligence (AI) applications. Register file, as the largest SRAM on the GPU die, accounts for over 20% of the total GPU energy consumption. Register cache has been introduced to reduce traffic from the register file and thus decrease total energy consumption when CUDA cores are utilized. However, the utilization of register cache has not been thoroughly investigated for Tensor Cores which are integrated into recent GPU architectures to meet AI workload demands. In this paper, we study the usage of register cache in both CUDA and Tensor Cores and conduct a thorough examination of their pros and cons. We have developed an open-source analytical simulator, called RFC-sim, to model and measure the energy consumption of both the register file and register cache. Our results show that while the register cache can reduce energy consumption by up to 40% in CUDA cores, it results in increased energy consumption by up to 23% in Tensor Cores. The main reason lies in the limited space of the register cache, which is not sufficient for the demand of Tensor cores to capture locality.","PeriodicalId":48827,"journal":{"name":"IEEE Journal on Emerging and Selected Topics in Circuits and Systems","volume":"14 3","pages":"469-482"},"PeriodicalIF":3.7,"publicationDate":"2024-08-05","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"141934967","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"工程技术","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}

引用次数: 0

Dynamic Fault Tolerance Approach for Network-on-Chip Architecture 片上网络架构的动态容错方法

IF 3.7 2区工程技术 Q2 ENGINEERING, ELECTRICAL & ELECTRONIC

IEEE Journal on Emerging and Selected Topics in Circuits and Systems

Pub Date : 2024-08-05 DOI: 10.1109/JETCAS.2024.3438250

Kasem Khalil;Ashok Kumar;Magdy Bayoumi

Network-on-Chip (NoC) architecture provides speed-efficient and scalable communication in complex integrated circuits. Attaining fault tolerance in NoC architectures is an ongoing research problem aiming to enhance the architecture’s reliability and performance. It seeks to mitigate the impact of router failures and enhance the overall system robustness. Fault tolerance is achieved by adding additional hardware, and the research challenge is to attain high reliability, high Mean Time To Failure (MTTF), and low Energy-Delay-Product (EDP) while sacrificing an acceptable area. It is particularly vital for applications with uninterrupted data flow. This paper proposes a fault-tolerance approach for NoC systems focusing on NoC routers to yield increased reliability and MTTF with an acceptable area overhead and low EDP. The proposed method proposes a dynamic reconfiguration mechanism by using a dynamic allocation of virtual channels and a bypass crossbar mechanism, ensuring uninterrupted data flow within the NoC. Evaluations of the proposed method are done on different mesh sizes using VHDL and an Altera 10GX FPGA, demonstrating the method’s superiority in reliability, reduced latency, and enhanced throughput. The results show that the proposed method has an acceptable area overhead of 25.3%, and its MTTF values are 3.7 times to 18 times higher than the traditional methods for varying network sizes, showing remarkable robustness against faults. The results show that the proposed method attains the best-reported reliability with the least EDP. Additionally, a layout of the circuit is also created and studied.

片上网络（NoC）架构可在复杂的集成电路中提供快速高效和可扩展的通信。在 NoC 架构中实现容错是一个持续研究的问题，目的是提高架构的可靠性和性能。它旨在减轻路由器故障的影响，提高整个系统的鲁棒性。容错是通过添加额外的硬件来实现的，而研究的挑战是在牺牲可接受的面积的同时，实现高可靠性、高平均故障时间（MTTF）和低能量-延迟-产品（EDP）。这对于需要不间断数据流的应用尤为重要。本文针对 NoC 系统提出了一种容错方法，重点关注 NoC 路由器，以提高可靠性和 MTTF，同时实现可接受的面积开销和低 EDP。所提出的方法通过使用虚拟通道的动态分配和旁路交叉条机制，提出了一种动态重新配置机制，以确保 NoC 内的数据流不中断。使用 VHDL 和 Altera 10GX FPGA 在不同网格尺寸上对所提方法进行了评估，证明了该方法在可靠性、降低延迟和提高吞吐量方面的优越性。结果表明，所提方法的面积开销为 25.3%，是可以接受的，而且在不同的网络规模下，其 MTTF 值是传统方法的 3.7 倍至 18 倍，显示了对故障的显著鲁棒性。结果表明，所提出的方法以最小的 EDP 达到了最佳的可靠性。此外，还创建并研究了电路布局。

{"title":"Dynamic Fault Tolerance Approach for Network-on-Chip Architecture","authors":"Kasem Khalil;Ashok Kumar;Magdy Bayoumi","doi":"10.1109/JETCAS.2024.3438250","DOIUrl":"10.1109/JETCAS.2024.3438250","url":null,"abstract":"Network-on-Chip (NoC) architecture provides speed-efficient and scalable communication in complex integrated circuits. Attaining fault tolerance in NoC architectures is an ongoing research problem aiming to enhance the architecture’s reliability and performance. It seeks to mitigate the impact of router failures and enhance the overall system robustness. Fault tolerance is achieved by adding additional hardware, and the research challenge is to attain high reliability, high Mean Time To Failure (MTTF), and low Energy-Delay-Product (EDP) while sacrificing an acceptable area. It is particularly vital for applications with uninterrupted data flow. This paper proposes a fault-tolerance approach for NoC systems focusing on NoC routers to yield increased reliability and MTTF with an acceptable area overhead and low EDP. The proposed method proposes a dynamic reconfiguration mechanism by using a dynamic allocation of virtual channels and a bypass crossbar mechanism, ensuring uninterrupted data flow within the NoC. Evaluations of the proposed method are done on different mesh sizes using VHDL and an Altera 10GX FPGA, demonstrating the method’s superiority in reliability, reduced latency, and enhanced throughput. The results show that the proposed method has an acceptable area overhead of 25.3%, and its MTTF values are 3.7 times to 18 times higher than the traditional methods for varying network sizes, showing remarkable robustness against faults. The results show that the proposed method attains the best-reported reliability with the least EDP. Additionally, a layout of the circuit is also created and studied.","PeriodicalId":48827,"journal":{"name":"IEEE Journal on Emerging and Selected Topics in Circuits and Systems","volume":"14 3","pages":"384-394"},"PeriodicalIF":3.7,"publicationDate":"2024-08-05","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":null,"resultStr":null,"platform":"Semanticscholar","paperid":"141968836","PeriodicalName":null,"FirstCategoryId":null,"ListUrlMain":null,"RegionNum":2,"RegionCategory":"工程技术","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":"","EPubDate":null,"PubModel":null,"JCR":null,"JCRName":null,"Score":null,"Total":0}

引用次数: 0

首页上一页

下一页尾页

类型

全部化学•材料生命科学医学物理工程技术环境•农林材料科学地球科学法学管理学化学环境科学与生态学计算机科学教育学经济学农林科学人文科学生物学数学物理与天体物理心理学综合性期刊其他工业工程理学历史学农学文学信息工程

数据库

全部 ACS Publications Elsevier ieeexplore Springer The Royal Society of Chemistry Wiley

期刊

IEEE Journal on Emerging and Selected Topics in Circuits and Systems

全部 Acc. Chem. Res. ACS Applied Bio Materials ACS Appl. Electron. Mater. ACS Appl. Energy Mater. ACS Appl. Mater. Interfaces ACS Appl. Nano Mater. ACS Appl. Polym. Mater. ACS BIOMATER-SCI ENG ACS Catal. ACS Cent. Sci. ACS Chem. Biol. ACS Chemical Health & Safety ACS Chem. Neurosci. ACS Comb. Sci. ACS Earth Space Chem. ACS Energy Lett. ACS Infect. Dis. ACS Macro Lett. ACS Mater. Lett. ACS Med. Chem. Lett. ACS Nano ACS Omega ACS Photonics ACS Sens. ACS Sustainable Chem. Eng. ACS Synth. Biol. Anal. Chem. BIOCHEMISTRY-US Bioconjugate Chem. BIOMACROMOLECULES Chem. Res. Toxicol. Chem. Rev. Chem. Mater. CRYST GROWTH DES ENERG FUEL Environ. Sci. Technol. Environ. Sci. Technol. Lett. Eur. J. Inorg. Chem. IND ENG CHEM RES Inorg. Chem. J. Agric. Food. Chem. J. Chem. Eng. Data J. Chem. Educ. J. Chem. Inf. Model. J. Chem. Theory Comput. J. Med. Chem. J. Nat. Prod. J PROTEOME RES J. Am. Chem. Soc. LANGMUIR MACROMOLECULES Mol. Pharmaceutics Nano Lett. Org. Lett. ORG PROCESS RES DEV ORGANOMETALLICS J. Org. Chem. J. Phys. Chem. J. Phys. Chem. A J. Phys. Chem. B J. Phys. Chem. C J. Phys. Chem. Lett. Analyst Anal. Methods Biomater. Sci. Catal. Sci. Technol. Chem. Commun. Chem. Soc. Rev. CHEM EDUC RES PRACT CRYSTENGCOMM Dalton Trans. Energy Environ. Sci. ENVIRON SCI-NANO ENVIRON SCI-PROC IMP ENVIRON SCI-WAT RES Faraday Discuss. Food Funct. Green Chem. Inorg. Chem. Front. Integr. Biol. J. Anal. At. Spectrom. J. Mater. Chem. A J. Mater. Chem. B J. Mater. Chem. C Lab Chip Mater. Chem. Front. Mater. Horiz. MEDCHEMCOMM Metallomics Mol. Biosyst. Mol. Syst. Des. Eng. Nanoscale Nanoscale Horiz. Nat. Prod. Rep. New J. Chem. Org. Biomol. Chem. Org. Chem. Front. PHOTOCH PHOTOBIO SCI PCCP Polym. Chem.

﹀