基于模型的复杂动态系统反馈控制策略优化算法

IF 3.9 2区工程技术 Q2 COMPUTER SCIENCE, INTERDISCIPLINARY APPLICATIONS Computers & Chemical Engineering Pub Date : 2025-04-01 Epub Date: 2025-02-06 DOI:10.1016/j.compchemeng.2025.109032

Lucky E. Yerimah, Christian Jorgensen, B. Wayne Bequette

{"title":"基于模型的复杂动态系统反馈控制策略优化算法","authors":"Lucky E. Yerimah, Christian Jorgensen, B. Wayne Bequette","doi":"10.1016/j.compchemeng.2025.109032","DOIUrl":null,"url":null,"abstract":"<div><div>Model-free Reinforcement learning (RL) has been successfully used in benchmark systems such as the Cart-Pole, Inverted-Pendulum, and Robotic arms. However, model-free RL algorithms have several limitations, including large data requirements and handling of state constraints. Model-based and hybrid RL algorithms offer opportunities to tackle these limitations. This research investigated the application of a model-based policy optimization algorithm (MBPO) for feedback control of the Van de Vusse reaction and the Quadruple tank system. MBPO-trained agents suffer from inaccuracies of the learned model and the computational burden of the online optimization neural network models and policy parameters. We propose a modified model-based policy optimization (MMBPO) algorithm that uses linear dynamic system models. This minimizes a learned model’s inaccuracies and eliminates the computational requirements of training the neural network models. Simulation results show that model-based policy optimization algorithms can track the setpoints of the dynamic systems studied.</div></div>","PeriodicalId":286,"journal":{"name":"Computers & Chemical Engineering","volume":"195 ","pages":"Article 109032"},"PeriodicalIF":3.9000,"publicationDate":"2025-04-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":"{\"title\":\"Model-based policy optimization algorithms for feedback control of complex dynamic systems\",\"authors\":\"Lucky E. Yerimah, Christian Jorgensen, B. Wayne Bequette\",\"doi\":\"10.1016/j.compchemeng.2025.109032\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"<div><div>Model-free Reinforcement learning (RL) has been successfully used in benchmark systems such as the Cart-Pole, Inverted-Pendulum, and Robotic arms. However, model-free RL algorithms have several limitations, including large data requirements and handling of state constraints. Model-based and hybrid RL algorithms offer opportunities to tackle these limitations. This research investigated the application of a model-based policy optimization algorithm (MBPO) for feedback control of the Van de Vusse reaction and the Quadruple tank system. MBPO-trained agents suffer from inaccuracies of the learned model and the computational burden of the online optimization neural network models and policy parameters. We propose a modified model-based policy optimization (MMBPO) algorithm that uses linear dynamic system models. This minimizes a learned model’s inaccuracies and eliminates the computational requirements of training the neural network models. Simulation results show that model-based policy optimization algorithms can track the setpoints of the dynamic systems studied.</div></div>\",\"PeriodicalId\":286,\"journal\":{\"name\":\"Computers & Chemical Engineering\",\"volume\":\"195 \",\"pages\":\"Article 109032\"},\"PeriodicalIF\":3.9000,\"publicationDate\":\"2025-04-01\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"0\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"Computers & Chemical Engineering\",\"FirstCategoryId\":\"5\",\"ListUrlMain\":\"https://www.sciencedirect.com/science/article/pii/S0098135425000365\",\"RegionNum\":2,\"RegionCategory\":\"工程技术\",\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"2025/2/6 0:00:00\",\"PubModel\":\"Epub\",\"JCR\":\"Q2\",\"JCRName\":\"COMPUTER SCIENCE, INTERDISCIPLINARY APPLICATIONS\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"Computers & Chemical Engineering","FirstCategoryId":"5","ListUrlMain":"https://www.sciencedirect.com/science/article/pii/S0098135425000365","RegionNum":2,"RegionCategory":"工程技术","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"2025/2/6 0:00:00","PubModel":"Epub","JCR":"Q2","JCRName":"COMPUTER SCIENCE, INTERDISCIPLINARY APPLICATIONS","Score":null,"Total":0}

引用次数: 0

摘要

无模型强化学习（RL）已经成功地应用于诸如Cart-Pole、倒摆和机械臂等基准系统中。然而，无模型强化学习算法有几个限制，包括大数据需求和状态约束的处理。基于模型和混合RL算法提供了解决这些限制的机会。本文研究了基于模型的策略优化算法（MBPO）在Van de Vusse反应和四缸系统反馈控制中的应用。mbpo训练的智能体存在学习模型的不准确性以及在线优化神经网络模型和策略参数的计算负担。我们提出了一种改进的基于模型的策略优化（MMBPO）算法，该算法使用线性动态系统模型。这最大限度地减少了学习模型的不准确性，消除了训练神经网络模型的计算需求。仿真结果表明，基于模型的策略优化算法能够跟踪所研究的动态系统的设定值。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

查看原文

微信好友朋友圈 QQ好友复制链接

本刊更多论文

Model-based policy optimization algorithms for feedback control of complex dynamic systems

Model-free Reinforcement learning (RL) has been successfully used in benchmark systems such as the Cart-Pole, Inverted-Pendulum, and Robotic arms. However, model-free RL algorithms have several limitations, including large data requirements and handling of state constraints. Model-based and hybrid RL algorithms offer opportunities to tackle these limitations. This research investigated the application of a model-based policy optimization algorithm (MBPO) for feedback control of the Van de Vusse reaction and the Quadruple tank system. MBPO-trained agents suffer from inaccuracies of the learned model and the computational burden of the online optimization neural network models and policy parameters. We propose a modified model-based policy optimization (MMBPO) algorithm that uses linear dynamic system models. This minimizes a learned model’s inaccuracies and eliminates the computational requirements of training the neural network models. Simulation results show that model-based policy optimization algorithms can track the setpoints of the dynamic systems studied.

求助全文

通过发布文献求助，成功后即可免费获取论文全文。去求助

来源期刊

Computers & Chemical Engineering 工程技术-工程：化工

CiteScore

8.70

自引率

14.00%

发文量

374

审稿时长

70 days

期刊介绍： Computers & Chemical Engineering is primarily a journal of record for new developments in the application of computing and systems technology to chemical engineering problems.