针对高不确定性蜂群对抗的分层强化学习

IF 7.9 2区 计算机科学 Q1 AUTOMATION & CONTROL SYSTEMS IEEE Transactions on Automation Science and Engineering Pub Date : 2024-11-05 DOI:10.1109/TASE.2024.3487219
Qizhen Wu;Kexin Liu;Lei Chen;Jinhu Lü
{"title":"针对高不确定性蜂群对抗的分层强化学习","authors":"Qizhen Wu;Kexin Liu;Lei Chen;Jinhu Lü","doi":"10.1109/TASE.2024.3487219","DOIUrl":null,"url":null,"abstract":"In swarm robotics, confrontation including the pursuit-evasion game is a key scenario. High uncertainty caused by unknown opponents’ strategies, dynamic obstacles, and insufficient training complicates the action space into a hybrid decision process. Although the deep reinforcement learning method is significant for swarm confrontation since it can handle various sizes, as an end-to–end implementation, it cannot deal with the hybrid process. Here, we propose a novel hierarchical reinforcement learning approach consisting of a target allocation layer, a path planning layer, and the underlying dynamic interaction mechanism between the two layers, which indicates the quantified uncertainty. It decouples the hybrid process into discrete allocation and continuous planning layers, with a probabilistic ensemble model to quantify the uncertainty and regulate the interaction frequency adaptively. Furthermore, to overcome the unstable training process introduced by the two layers, we design an integration training method including pre-training and cross-training, which enhances the training efficiency and stability. Experiment results in both comparison, ablation, and real-robot studies validate the effectiveness and generalization performance of our proposed approach. In our defined experiments with twenty to forty agents, the win rate of the proposed method reaches around ninety percent, outperforming other traditional methods. Note to Practitioners—With artificial intelligence rapidly developing, robots will play a significant role in the future. Especially, the swarm formed by many robots holds promising potential in civil and military applications. Promoting the swarm into games or battles is rather riveting. The reinforcement learning method provides a plausible solution to realize the battle of robotic swarms. There are still some issues that need to be addressed. On one hand, we focus on the uncertainty caused by the battlefield nature and the environment which limits our ability for the implementation of swarms. On the other hand, we solve the problem that the decision process combined with commands and actions is a hybrid system, which cannot be directly reflected in the confrontation of swarms. Overall, our approaches throw light on artificial general intelligence and also reveal a solution to interpretable intelligence.","PeriodicalId":51060,"journal":{"name":"IEEE Transactions on Automation Science and Engineering","volume":"22 ","pages":"8630-8644"},"PeriodicalIF":7.9000,"publicationDate":"2024-11-05","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":"{\"title\":\"Hierarchical Reinforcement Learning for Swarm Confrontation With High Uncertainty\",\"authors\":\"Qizhen Wu;Kexin Liu;Lei Chen;Jinhu Lü\",\"doi\":\"10.1109/TASE.2024.3487219\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"In swarm robotics, confrontation including the pursuit-evasion game is a key scenario. High uncertainty caused by unknown opponents’ strategies, dynamic obstacles, and insufficient training complicates the action space into a hybrid decision process. Although the deep reinforcement learning method is significant for swarm confrontation since it can handle various sizes, as an end-to–end implementation, it cannot deal with the hybrid process. Here, we propose a novel hierarchical reinforcement learning approach consisting of a target allocation layer, a path planning layer, and the underlying dynamic interaction mechanism between the two layers, which indicates the quantified uncertainty. It decouples the hybrid process into discrete allocation and continuous planning layers, with a probabilistic ensemble model to quantify the uncertainty and regulate the interaction frequency adaptively. Furthermore, to overcome the unstable training process introduced by the two layers, we design an integration training method including pre-training and cross-training, which enhances the training efficiency and stability. Experiment results in both comparison, ablation, and real-robot studies validate the effectiveness and generalization performance of our proposed approach. In our defined experiments with twenty to forty agents, the win rate of the proposed method reaches around ninety percent, outperforming other traditional methods. Note to Practitioners—With artificial intelligence rapidly developing, robots will play a significant role in the future. Especially, the swarm formed by many robots holds promising potential in civil and military applications. Promoting the swarm into games or battles is rather riveting. The reinforcement learning method provides a plausible solution to realize the battle of robotic swarms. There are still some issues that need to be addressed. On one hand, we focus on the uncertainty caused by the battlefield nature and the environment which limits our ability for the implementation of swarms. On the other hand, we solve the problem that the decision process combined with commands and actions is a hybrid system, which cannot be directly reflected in the confrontation of swarms. Overall, our approaches throw light on artificial general intelligence and also reveal a solution to interpretable intelligence.\",\"PeriodicalId\":51060,\"journal\":{\"name\":\"IEEE Transactions on Automation Science and Engineering\",\"volume\":\"22 \",\"pages\":\"8630-8644\"},\"PeriodicalIF\":7.9000,\"publicationDate\":\"2024-11-05\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"0\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"IEEE Transactions on Automation Science and Engineering\",\"FirstCategoryId\":\"94\",\"ListUrlMain\":\"https://ieeexplore.ieee.org/document/10744028/\",\"RegionNum\":2,\"RegionCategory\":\"计算机科学\",\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"Q1\",\"JCRName\":\"AUTOMATION & CONTROL SYSTEMS\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"IEEE Transactions on Automation Science and Engineering","FirstCategoryId":"94","ListUrlMain":"https://ieeexplore.ieee.org/document/10744028/","RegionNum":2,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"Q1","JCRName":"AUTOMATION & CONTROL SYSTEMS","Score":null,"Total":0}
引用次数: 0

摘要

在群体机器人中,包括追击-逃避游戏在内的对抗是一个关键的场景。由于未知对手的策略、动态障碍和训练不足导致的高度不确定性使行动空间成为一个混合决策过程。虽然深度强化学习方法可以处理各种规模的群体对抗,但作为端到端实现,它无法处理混合过程。在此,我们提出了一种新的分层强化学习方法,该方法由目标分配层、路径规划层和两层之间的动态交互机制组成,这表明了量化的不确定性。将混合过程解耦为离散分配层和连续规划层,采用概率集成模型量化不确定性,自适应调节交互频率。此外,为了克服两层训练引入的不稳定训练过程,我们设计了一种包含预训练和交叉训练的集成训练方法,提高了训练效率和稳定性。对比、消融和真实机器人研究的实验结果验证了我们提出的方法的有效性和泛化性能。在我们定义的20到40个智能体的实验中,该方法的胜率达到90%左右,优于其他传统方法。从业人员注意:随着人工智能的迅速发展,机器人将在未来发挥重要作用。特别是,由众多机器人组成的群体在民用和军事应用方面具有很大的潜力。将蜂群推广到游戏或战斗中是非常吸引人的。强化学习方法为实现机器人群作战提供了一种可行的解决方案。还有一些问题需要解决。一方面,我们关注由战场性质和环境造成的不确定性,这些不确定性限制了我们实施蜂群的能力。另一方面,解决了命令与行动相结合的决策过程是一个混合系统,不能直接反映在群体对抗中的问题。总的来说,我们的方法揭示了人工通用智能,也揭示了可解释智能的解决方案。
本文章由计算机程序翻译,如有差异,请以英文原文为准。
查看原文
分享 分享
微信好友 朋友圈 QQ好友 复制链接
本刊更多论文
Hierarchical Reinforcement Learning for Swarm Confrontation With High Uncertainty
In swarm robotics, confrontation including the pursuit-evasion game is a key scenario. High uncertainty caused by unknown opponents’ strategies, dynamic obstacles, and insufficient training complicates the action space into a hybrid decision process. Although the deep reinforcement learning method is significant for swarm confrontation since it can handle various sizes, as an end-to–end implementation, it cannot deal with the hybrid process. Here, we propose a novel hierarchical reinforcement learning approach consisting of a target allocation layer, a path planning layer, and the underlying dynamic interaction mechanism between the two layers, which indicates the quantified uncertainty. It decouples the hybrid process into discrete allocation and continuous planning layers, with a probabilistic ensemble model to quantify the uncertainty and regulate the interaction frequency adaptively. Furthermore, to overcome the unstable training process introduced by the two layers, we design an integration training method including pre-training and cross-training, which enhances the training efficiency and stability. Experiment results in both comparison, ablation, and real-robot studies validate the effectiveness and generalization performance of our proposed approach. In our defined experiments with twenty to forty agents, the win rate of the proposed method reaches around ninety percent, outperforming other traditional methods. Note to Practitioners—With artificial intelligence rapidly developing, robots will play a significant role in the future. Especially, the swarm formed by many robots holds promising potential in civil and military applications. Promoting the swarm into games or battles is rather riveting. The reinforcement learning method provides a plausible solution to realize the battle of robotic swarms. There are still some issues that need to be addressed. On one hand, we focus on the uncertainty caused by the battlefield nature and the environment which limits our ability for the implementation of swarms. On the other hand, we solve the problem that the decision process combined with commands and actions is a hybrid system, which cannot be directly reflected in the confrontation of swarms. Overall, our approaches throw light on artificial general intelligence and also reveal a solution to interpretable intelligence.
求助全文
通过发布文献求助,成功后即可免费获取论文全文。 去求助
来源期刊
IEEE Transactions on Automation Science and Engineering
IEEE Transactions on Automation Science and Engineering 工程技术-自动化与控制系统
CiteScore
12.50
自引率
14.30%
发文量
404
审稿时长
3.0 months
期刊介绍: The IEEE Transactions on Automation Science and Engineering (T-ASE) publishes fundamental papers on Automation, emphasizing scientific results that advance efficiency, quality, productivity, and reliability. T-ASE encourages interdisciplinary approaches from computer science, control systems, electrical engineering, mathematics, mechanical engineering, operations research, and other fields. T-ASE welcomes results relevant to industries such as agriculture, biotechnology, healthcare, home automation, maintenance, manufacturing, pharmaceuticals, retail, security, service, supply chains, and transportation. T-ASE addresses a research community willing to integrate knowledge across disciplines and industries. For this purpose, each paper includes a Note to Practitioners that summarizes how its results can be applied or how they might be extended to apply in practice.
期刊最新文献
Human-Robot Collaborative Disassembly Planning via Multi-Objective Proximal Policy Optimization with Teacher-Student Co-Evolution Mechanism Model-Free Fixed-Time Sliding Mode Control for Wearable Exoskeletons with Adaptive-Boundary Prescribed Performance DualSkill: Unifying Discrete Stability and Continuous Flexibility for Embodied Control Event-Triggered Observer-Based Secure Fault Estimation and Fault-Tolerant Control for Markov Jump Systems A Contrastive Learning Framework With Individualized Similarity for Industrial Soft Sensing
×
引用
GB/T 7714-2015
复制
MLA
复制
APA
复制
导出至
BibTeX EndNote RefMan NoteFirst NoteExpress
×
×
提示
您的信息不完整,为了账户安全,请先补充。
现在去补充
×
提示
您因"违规操作"
具体请查看互助需知
我知道了
×
提示
现在去查看 取消
×
提示
确定
0
微信
客服QQ
Book学术公众号 扫码关注我们
反馈
×
意见反馈
请填写您的意见或建议
请填写您的手机或邮箱
已复制链接
已复制链接
快去分享给好友吧!
我知道了
×
扫码分享
扫码分享
Book学术官方微信
Book学术官方微信
Book学术文献互助
Book学术文献互助群
群 号:604180095
Book学术
文献互助 智能选刊 最新文献 互助须知 联系我们:info@booksci.cn
Book学术提供免费学术资源搜索服务,方便国内外学者检索中英文文献。致力于提供最便捷和优质的服务体验。
Copyright © 2023 Book学术 All rights reserved.
ghs 京公网安备 11010802042870号 京ICP备2023020795号-1