Layer-wise Knowledge Distillation for Cross-Device Federated Learning

2023 International Conference on Information Networking (ICOIN) Pub Date : 2023-01-11 DOI:10.1109/ICOIN56518.2023.10049011

Huy Q. Le, Loc X. Nguyen, Seong-Bae Park, C. Hong

{"title":"Layer-wise Knowledge Distillation for Cross-Device Federated Learning","authors":"Huy Q. Le, Loc X. Nguyen, Seong-Bae Park, C. Hong","doi":"10.1109/ICOIN56518.2023.10049011","DOIUrl":null,"url":null,"abstract":"Federated Learning (FL) has been proposed as a decentralized machine learning system where multiple clients jointly train the model without sharing private data. In FL, the statistical heterogeneity among devices has become a crucial challenge, which can cause degradation in generalization performance. Previous FL approaches have proven that leveraging the proximal regularization at the local training process can alleviate the divergence of parameter aggregation from biased local models. In this work, to address the heterogeneity issues in conventional FL, we propose a layer-wise knowledge distillation method in federated learning, namely, FedLKD, which regularizes the local training step via the knowledge distillation scheme between global and local models utilizing the small proxy dataset. Hence, FedLKD deploys the layer-wise knowledge distillation of the multiple devices and the global server as the clients’ regularized loss function. A layer-wise knowledge distillation mechanism is introduced to update the local model to exploit the common representation from different layers. Through extensive experiments, we demonstrate that FedLKD outperforms the vanilla FedAvg and FedProx on three federated datasets.","PeriodicalId":285763,"journal":{"name":"2023 International Conference on Information Networking (ICOIN)","volume":"1 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2023-01-11","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"2023 International Conference on Information Networking (ICOIN)","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1109/ICOIN56518.2023.10049011","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 0

Abstract

Federated Learning (FL) has been proposed as a decentralized machine learning system where multiple clients jointly train the model without sharing private data. In FL, the statistical heterogeneity among devices has become a crucial challenge, which can cause degradation in generalization performance. Previous FL approaches have proven that leveraging the proximal regularization at the local training process can alleviate the divergence of parameter aggregation from biased local models. In this work, to address the heterogeneity issues in conventional FL, we propose a layer-wise knowledge distillation method in federated learning, namely, FedLKD, which regularizes the local training step via the knowledge distillation scheme between global and local models utilizing the small proxy dataset. Hence, FedLKD deploys the layer-wise knowledge distillation of the multiple devices and the global server as the clients’ regularized loss function. A layer-wise knowledge distillation mechanism is introduced to update the local model to exploit the common representation from different layers. Through extensive experiments, we demonstrate that FedLKD outperforms the vanilla FedAvg and FedProx on three federated datasets.

查看原文

微信好友朋友圈 QQ好友复制链接

本刊更多论文

面向跨设备联邦学习的分层知识蒸馏

联邦学习(FL)是一种分散的机器学习系统，其中多个客户端在不共享私有数据的情况下共同训练模型。在FL中，设备之间的统计异质性已经成为一个关键的挑战，它可能导致泛化性能的下降。先前的FL方法已经证明，在局部训练过程中利用近端正则化可以减轻有偏差的局部模型的参数聚集的分歧。在这项工作中，为了解决传统FL中的异构问题，我们提出了一种分层知识蒸馏方法，即FedLKD，该方法利用小型代理数据集，通过全局模型和局部模型之间的知识蒸馏方案来正则化局部训练步骤。因此，FedLKD将多个设备和全局服务器的分层知识蒸馏部署为客户端的正则化损失函数。引入了一种分层知识蒸馏机制来更新局部模型，以利用来自不同层的通用表示。通过大量的实验，我们证明了fedkd在三个联邦数据集上优于传统的fedag和FedProx。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文去求助

来源期刊

2023 International Conference on Information Networking (ICOIN)

自引率

0.00%

发文量