MESA: Cooperative Meta-Exploration in Multi-Agent Learning through Exploiting State-Action Space Structure

Adaptive Agents and Multi-Agent Systems Pub Date : 2024-05-01 DOI:10.5555/3635637.3663073

Zhicheng Zhang, Yancheng Liang, Yi Wu, Fei Fang

引用次数: 1

Abstract

Multi-agent reinforcement learning (MARL) algorithms often struggle to find strategies close to Pareto optimal Nash Equilibrium, owing largely to the lack of efficient exploration. The problem is exacerbated in sparse-reward settings, caused by the larger variance exhibited in policy learning. This paper introduces MESA, a novel meta-exploration method for cooperative multi-agent learning. It learns to explore by first identifying the agents' high-rewarding joint state-action subspace from training tasks and then learning a set of diverse exploration policies to"cover"the subspace. These trained exploration policies can be integrated with any off-policy MARL algorithm for test-time tasks. We first showcase MESA's advantage in a multi-step matrix game. Furthermore, experiments show that with learned exploration policies, MESA achieves significantly better performance in sparse-reward tasks in several multi-agent particle environments and multi-agent MuJoCo environments, and exhibits the ability to generalize to more challenging tasks at test time.

查看原文

微信好友朋友圈 QQ好友复制链接

本刊更多论文

MESA：通过利用状态-行动空间结构在多代理学习中进行合作元探索

多代理强化学习（MARL）算法通常很难找到接近帕累托最优纳什均衡的策略，这主要是由于缺乏有效的探索。在奖励稀疏的环境中，由于策略学习中表现出的较大方差，这一问题更加严重。本文介绍了 MESA，一种用于多机器人合作学习的新型元探索方法。它首先从训练任务中识别出各代理的高回报联合状态-行动子空间，然后学习一系列不同的探索策略来 "覆盖 "该子空间，从而学会探索。这些经过训练的探索策略可与任何非策略 MARL 算法集成，用于测试时间任务。我们首先展示了 MESA 在多步骤矩阵博弈中的优势。此外，实验表明，利用学习到的探索策略，MESA 在多个多代理粒子环境和多代理 MuJoCo 环境中的稀疏奖励任务中取得了明显更好的性能，并表现出了在测试时间泛化到更具挑战性任务的能力。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文去求助

来源期刊

Adaptive Agents and Multi-Agent Systems

自引率

0.00%

发文量

期刊最新文献

Discovering Consistent Subelections Strategic Cost Selection in Participatory Budgeting Minimizing State Exploration While Searching Graphs with Unknown Obstacles vMFER: von Mises-Fisher Experience Resampling Based on Uncertainty of Gradient Directions for Policy Improvement of Actor-Critic Algorithms Reinforcement Nash Equilibrium Solver