Decoupled Prioritized Resampling for Offline RL

IF 9.7 1区 计算机科学 Q1 COMPUTER SCIENCE, ARTIFICIAL INTELLIGENCE IEEE transactions on neural networks and learning systems Pub Date : 2024-11-21 DOI:10.1109/TNNLS.2024.3488358
Yang Yue;Bingyi Kang;Xiao Ma;Qisen Yang;Gao Huang;Shiji Song;Shuicheng Yan
{"title":"Decoupled Prioritized Resampling for Offline RL","authors":"Yang Yue;Bingyi Kang;Xiao Ma;Qisen Yang;Gao Huang;Shiji Song;Shuicheng Yan","doi":"10.1109/TNNLS.2024.3488358","DOIUrl":null,"url":null,"abstract":"Offline reinforcement learning (RL) is challenged by the distributional shift problem. To tackle this issue, existing works mainly focus on designing sophisticated policy constraints between the learned policy and the behavior policy. However, these constraints are applied equally to well-performing and inferior actions through uniform sampling, which might negatively affect the learned policy. In this article, we propose offline decoupled prioritized resampling (ODPR), which designs specialized priority functions for the suboptimal policy constraint issue in offline RL and employs unique decoupled resampling for training stability. Through theoretical analysis, we show that the distinctive priority functions induce a provable improved behavior policy by modifying the distribution of the original behavior policy, and when constrained to this improved policy, a policy-constrained offline RL algorithm is likely to yield a better solution. We provide two practical implementations to balance computation and performance: one estimates priorities based on a fit value network [advantage-based ODPR (ODPR-A)] and the other utilizes trajectory returns [return-based ODPR (ODPR-R)] for quick computation. As a highly compatible plug-and-play component, ODPR is evaluated with five prevalent offline RL algorithms: behavior cloning (BC), twin delayed deep deterministic policy gradient + BC (TD3 + BC), OnestepRL, conservative Q-learning (CQL), and implicit Q-learning (IQL). Our experiments confirm that both ODPR-A and ODPR-R significantly improve performance across all baseline methods. Moreover, ODPR-A can be effective in some challenging settings, i.e., without trajectory information. Code and pretrained weights are available at <uri>https://github.com/yueyang130/ODPR</uri>.","PeriodicalId":13303,"journal":{"name":"IEEE transactions on neural networks and learning systems","volume":"36 7","pages":"13094-13108"},"PeriodicalIF":9.7000,"publicationDate":"2024-11-21","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"IEEE transactions on neural networks and learning systems","FirstCategoryId":"94","ListUrlMain":"https://ieeexplore.ieee.org/document/10759860/","RegionNum":1,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"Q1","JCRName":"COMPUTER SCIENCE, ARTIFICIAL INTELLIGENCE","Score":null,"Total":0}
引用次数: 0

Abstract

Offline reinforcement learning (RL) is challenged by the distributional shift problem. To tackle this issue, existing works mainly focus on designing sophisticated policy constraints between the learned policy and the behavior policy. However, these constraints are applied equally to well-performing and inferior actions through uniform sampling, which might negatively affect the learned policy. In this article, we propose offline decoupled prioritized resampling (ODPR), which designs specialized priority functions for the suboptimal policy constraint issue in offline RL and employs unique decoupled resampling for training stability. Through theoretical analysis, we show that the distinctive priority functions induce a provable improved behavior policy by modifying the distribution of the original behavior policy, and when constrained to this improved policy, a policy-constrained offline RL algorithm is likely to yield a better solution. We provide two practical implementations to balance computation and performance: one estimates priorities based on a fit value network [advantage-based ODPR (ODPR-A)] and the other utilizes trajectory returns [return-based ODPR (ODPR-R)] for quick computation. As a highly compatible plug-and-play component, ODPR is evaluated with five prevalent offline RL algorithms: behavior cloning (BC), twin delayed deep deterministic policy gradient + BC (TD3 + BC), OnestepRL, conservative Q-learning (CQL), and implicit Q-learning (IQL). Our experiments confirm that both ODPR-A and ODPR-R significantly improve performance across all baseline methods. Moreover, ODPR-A can be effective in some challenging settings, i.e., without trajectory information. Code and pretrained weights are available at https://github.com/yueyang130/ODPR.
查看原文
分享 分享
微信好友 朋友圈 QQ好友 复制链接
本刊更多论文
离线 RL 的解耦优先重采样
离线强化学习(RL)受到分布移位问题的挑战。为了解决这个问题,现有的工作主要集中在设计学习策略和行为策略之间的复杂策略约束。然而,通过统一采样,这些约束同样适用于表现良好和较差的动作,这可能会对学习策略产生负面影响。在本文中,我们提出了离线解耦优先重采样(ODPR),它为离线RL中的次优策略约束问题设计了专门的优先级函数,并采用独特的解耦重采样来提高训练稳定性。通过理论分析,我们证明了不同的优先级函数通过修改原始行为策略的分布来诱导可证明的改进行为策略,并且当约束于该改进策略时,策略约束的离线RL算法可能产生更好的解。我们提供了两种实际的实现来平衡计算和性能:一种是基于适合值网络[基于优势的ODPR (ODPR- a)]估计优先级,另一种是利用轨迹返回[基于返回的ODPR (ODPR- r)]进行快速计算。作为一个高度兼容的即插即用组件,ODPR使用五种流行的离线RL算法进行评估:行为克隆(BC)、双延迟深度确定性策略梯度+ BC (TD3 + BC)、OnestepRL、保守q -学习(CQL)和隐式q -学习(IQL)。我们的实验证实,ODPR-A和ODPR-R在所有基线方法中都显著提高了性能。此外,ODPR-A在一些具有挑战性的环境中是有效的,例如,没有轨迹信息。代码和预训练的权重可以在https://github.com/yueyang130/ODPR上获得。
本文章由计算机程序翻译,如有差异,请以英文原文为准。
求助全文
约1分钟内获得全文 去求助
来源期刊
IEEE transactions on neural networks and learning systems
IEEE transactions on neural networks and learning systems COMPUTER SCIENCE, ARTIFICIAL INTELLIGENCE-COMPUTER SCIENCE, HARDWARE & ARCHITECTURE
CiteScore
23.80
自引率
9.60%
发文量
2102
审稿时长
3-8 weeks
期刊介绍: The focus of IEEE Transactions on Neural Networks and Learning Systems is to present scholarly articles discussing the theory, design, and applications of neural networks as well as other learning systems. The journal primarily highlights technical and scientific research in this domain.
期刊最新文献
Stable and Accurate Robot Trajectory Tracking Using Variable-Stiffness Euclideanizing Flow Robust Sequential Recommendation With Decorrelation and Debiasing ChACo: Channel-Wise Adaptive Competitive Layer-Wise Learning A Latent Diffusion for Stable Frame Interpolation. Spike-EIFNet: Lightweight Spike-Driven Event-Image Fusion Network for Accurate and Efficient Semantic Segmentation.
×
引用
GB/T 7714-2015
复制
MLA
复制
APA
复制
导出至
BibTeX EndNote RefMan NoteFirst NoteExpress
×
×
提示
您的信息不完整,为了账户安全,请先补充。
现在去补充
×
提示
您因"违规操作"
具体请查看互助需知
我知道了
×
提示
现在去查看 取消
×
提示
确定
0
微信
客服QQ
Book学术公众号 扫码关注我们
反馈
×
意见反馈
请填写您的意见或建议
请填写您的手机或邮箱
已复制链接
已复制链接
快去分享给好友吧!
我知道了
×
扫码分享
扫码分享
Book学术官方微信
Book学术官方微信
Book学术文献互助
Book学术文献互助群
群 号:604180095
Book学术
文献互助 智能选刊 最新文献 互助须知 联系我们:info@booksci.cn
Book学术提供免费学术资源搜索服务,方便国内外学者检索中英文文献。致力于提供最便捷和优质的服务体验。
Copyright © 2023 Book学术 All rights reserved.
ghs 京公网安备 11010802042870号 京ICP备2023020795号-1