多无人机网络中基于 RL 和 DRL 的分布式用户访问方案

IF 7.5 2区 计算机科学 Q1 ENGINEERING, ELECTRICAL & ELECTRONIC IEEE Transactions on Vehicular Technology Pub Date : 2024-11-18 DOI:10.1109/TVT.2024.3499976
Ming Cheng;Saifei He;Min Lin;Wei-Ping Zhu;Jiangzhou Wang
{"title":"多无人机网络中基于 RL 和 DRL 的分布式用户访问方案","authors":"Ming Cheng;Saifei He;Min Lin;Wei-Ping Zhu;Jiangzhou Wang","doi":"10.1109/TVT.2024.3499976","DOIUrl":null,"url":null,"abstract":"Unmanned aerial vehicles (UAVs) have been used as aerial platform to enhance the capacity and coverage of wireless networks. The user access is challenging due to the rapid varying channels between the moving UAVs and the ground users and the interference and conflict among users. This paper aims to investigate the user access and maximize the transmission rate while guaranteeing fairness and reducing handover. The multi-user access problem is formulated to a sequential decision problem in reinforcement learning (RL). A distributed multi-armed bandit (MAB) based algorithm is proposed to address this issue. The MAB based algorithm uses straightforward reward feedback to maintain a set of probabilistic weights, which help users to make decisions. Additionally, a multi-agent proximal policy optimization (MAPPO) based algorithm in deep reinforcement learning (DRL) is employed. The MAPPO based algorithm is centrally trained and executed in a distributed manner, and it is capable of efficiently handling multi-user access decisions. Simulation results show that the MAPPO based algorithm can achieve the highest system throughput and the distributed MAB based algorithm can reduce handover and enhancing fairness. The proposed distributed algorithms outperform benchmarks in throughput and robustness significantly.","PeriodicalId":13421,"journal":{"name":"IEEE Transactions on Vehicular Technology","volume":"74 3","pages":"5241-5246"},"PeriodicalIF":7.5000,"publicationDate":"2024-11-18","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":"{\"title\":\"RL and DRL Based Distributed User Access Schemes in Multi-UAV Networks\",\"authors\":\"Ming Cheng;Saifei He;Min Lin;Wei-Ping Zhu;Jiangzhou Wang\",\"doi\":\"10.1109/TVT.2024.3499976\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"Unmanned aerial vehicles (UAVs) have been used as aerial platform to enhance the capacity and coverage of wireless networks. The user access is challenging due to the rapid varying channels between the moving UAVs and the ground users and the interference and conflict among users. This paper aims to investigate the user access and maximize the transmission rate while guaranteeing fairness and reducing handover. The multi-user access problem is formulated to a sequential decision problem in reinforcement learning (RL). A distributed multi-armed bandit (MAB) based algorithm is proposed to address this issue. The MAB based algorithm uses straightforward reward feedback to maintain a set of probabilistic weights, which help users to make decisions. Additionally, a multi-agent proximal policy optimization (MAPPO) based algorithm in deep reinforcement learning (DRL) is employed. The MAPPO based algorithm is centrally trained and executed in a distributed manner, and it is capable of efficiently handling multi-user access decisions. Simulation results show that the MAPPO based algorithm can achieve the highest system throughput and the distributed MAB based algorithm can reduce handover and enhancing fairness. The proposed distributed algorithms outperform benchmarks in throughput and robustness significantly.\",\"PeriodicalId\":13421,\"journal\":{\"name\":\"IEEE Transactions on Vehicular Technology\",\"volume\":\"74 3\",\"pages\":\"5241-5246\"},\"PeriodicalIF\":7.5000,\"publicationDate\":\"2024-11-18\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"0\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"IEEE Transactions on Vehicular Technology\",\"FirstCategoryId\":\"94\",\"ListUrlMain\":\"https://ieeexplore.ieee.org/document/10755136/\",\"RegionNum\":2,\"RegionCategory\":\"计算机科学\",\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"Q1\",\"JCRName\":\"ENGINEERING, ELECTRICAL & ELECTRONIC\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"IEEE Transactions on Vehicular Technology","FirstCategoryId":"94","ListUrlMain":"https://ieeexplore.ieee.org/document/10755136/","RegionNum":2,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"Q1","JCRName":"ENGINEERING, ELECTRICAL & ELECTRONIC","Score":null,"Total":0}
引用次数: 0

摘要

无人机(uav)已被用作空中平台来增强无线网络的容量和覆盖范围。由于移动无人机与地面用户之间的信道快速变化以及用户之间的干扰和冲突,给用户接入带来了挑战。本文的目的是在保证公平性和减少切换的同时,研究用户访问和最大传输速率。将多用户访问问题形式化为强化学习(RL)中的顺序决策问题。针对这一问题,提出了一种基于分布式多臂强盗(MAB)算法。基于MAB的算法使用直接的奖励反馈来维持一组概率权重,这有助于用户做出决策。此外,在深度强化学习(DRL)中采用了基于多智能体近端策略优化(MAPPO)的算法。基于MAPPO的算法集中训练并以分布式方式执行,能够有效地处理多用户访问决策。仿真结果表明,基于MAPPO的算法可以获得最高的系统吞吐量,而基于分布式MAB的算法可以减少切换,提高公平性。所提出的分布式算法在吞吐量和鲁棒性方面明显优于基准测试。
本文章由计算机程序翻译,如有差异,请以英文原文为准。
查看原文
分享 分享
微信好友 朋友圈 QQ好友 复制链接
本刊更多论文
RL and DRL Based Distributed User Access Schemes in Multi-UAV Networks
Unmanned aerial vehicles (UAVs) have been used as aerial platform to enhance the capacity and coverage of wireless networks. The user access is challenging due to the rapid varying channels between the moving UAVs and the ground users and the interference and conflict among users. This paper aims to investigate the user access and maximize the transmission rate while guaranteeing fairness and reducing handover. The multi-user access problem is formulated to a sequential decision problem in reinforcement learning (RL). A distributed multi-armed bandit (MAB) based algorithm is proposed to address this issue. The MAB based algorithm uses straightforward reward feedback to maintain a set of probabilistic weights, which help users to make decisions. Additionally, a multi-agent proximal policy optimization (MAPPO) based algorithm in deep reinforcement learning (DRL) is employed. The MAPPO based algorithm is centrally trained and executed in a distributed manner, and it is capable of efficiently handling multi-user access decisions. Simulation results show that the MAPPO based algorithm can achieve the highest system throughput and the distributed MAB based algorithm can reduce handover and enhancing fairness. The proposed distributed algorithms outperform benchmarks in throughput and robustness significantly.
求助全文
通过发布文献求助,成功后即可免费获取论文全文。 去求助
来源期刊
CiteScore
6.00
自引率
8.80%
发文量
1245
审稿时长
6.3 months
期刊介绍: The scope of the Transactions is threefold (which was approved by the IEEE Periodicals Committee in 1967) and is published on the journal website as follows: Communications: The use of mobile radio on land, sea, and air, including cellular radio, two-way radio, and one-way radio, with applications to dispatch and control vehicles, mobile radiotelephone, radio paging, and status monitoring and reporting. Related areas include spectrum usage, component radio equipment such as cavities and antennas, compute control for radio systems, digital modulation and transmission techniques, mobile radio circuit design, radio propagation for vehicular communications, effects of ignition noise and radio frequency interference, and consideration of the vehicle as part of the radio operating environment. Transportation Systems: The use of electronic technology for the control of ground transportation systems including, but not limited to, traffic aid systems; traffic control systems; automatic vehicle identification, location, and monitoring systems; automated transport systems, with single and multiple vehicle control; and moving walkways or people-movers. Vehicular Electronics: The use of electronic or electrical components and systems for control, propulsion, or auxiliary functions, including but not limited to, electronic controls for engineer, drive train, convenience, safety, and other vehicle systems; sensors, actuators, and microprocessors for onboard use; electronic fuel control systems; vehicle electrical components and systems collision avoidance systems; electromagnetic compatibility in the vehicle environment; and electric vehicles and controls.
期刊最新文献
Optimization-Unfolded Physical Tokenizer for Noise-Robust Few-Shot Target Recognition via Physics-Guided LLMs Kendall's Tau-Based Spectrum Sensing with CFAR Property for Vehicular Communications Under Impulsive Noise Federated Edge Learning Based Multi-Drone Optimal Delivery Routing Approach Low-PAPR AFDM With Dual-Dimensional Index Modulation via Dynamic Pre-Chirp Selection Triggering Phantom Obstacles: Black-Box Physical Adversarial Attacks Against MDE for Autonomous Vehicles
×
引用
GB/T 7714-2015
复制
MLA
复制
APA
复制
导出至
BibTeX EndNote RefMan NoteFirst NoteExpress
×
×
提示
您的信息不完整,为了账户安全,请先补充。
现在去补充
×
提示
您因"违规操作"
具体请查看互助需知
我知道了
×
提示
现在去查看 取消
×
提示
确定
0
微信
客服QQ
Book学术公众号 扫码关注我们
反馈
×
意见反馈
请填写您的意见或建议
请填写您的手机或邮箱
已复制链接
已复制链接
快去分享给好友吧!
我知道了
×
扫码分享
扫码分享
Book学术官方微信
Book学术官方微信
Book学术文献互助
Book学术文献互助群
群 号:604180095
Book学术
文献互助 智能选刊 最新文献 互助须知 联系我们:info@booksci.cn
Book学术提供免费学术资源搜索服务,方便国内外学者检索中英文文献。致力于提供最便捷和优质的服务体验。
Copyright © 2023 Book学术 All rights reserved.
ghs 京公网安备 11010802042870号 京ICP备2023020795号-1