A New Deep Reinforcement Learning Run-to-Run Control Algorithm for Mixed-Product Production Mode in Semiconductor Manufacturing

IF 7.9 2区 计算机科学 Q1 AUTOMATION & CONTROL SYSTEMS IEEE Transactions on Automation Science and Engineering Pub Date : 2026-01-01 Epub Date: 2025-01-15 DOI:10.1109/TASE.2025.3526675
Shu-Kai S. Fan;Tzu-Jung Chen
{"title":"A New Deep Reinforcement Learning Run-to-Run Control Algorithm for Mixed-Product Production Mode in Semiconductor Manufacturing","authors":"Shu-Kai S. Fan;Tzu-Jung Chen","doi":"10.1109/TASE.2025.3526675","DOIUrl":null,"url":null,"abstract":"This paper proposes a new Run-to-Run (R2R) control framework based on deep deterministic policy gradient (DDPG) for the mixed-product production mode in semiconductor manufacturing. The DDPG algorithm is particularly developed to configure a deep reinforcement learning environment well suited to mixed-product production modes. To address the challenges posed in deep reinforcement learning, three enhanced mechanisms have been developed to improve the training of the proposed DDPG model for mixed-product R2R applications. These mechanisms include a piece-wise reward function, training with dynamic targets, and the new recall principle. It is demonstrated from the comprehensive simulation results that the proposed R2R control framework outperforms five noted mixed-product R2R control algorithms in the literature. The research outcome of this paper signifies a promising viability of deep reinforcement learning for highly complex and dynamic environments with continuous action spaces in the mixed-product R2R practice.Note to Practitioners—The mixed-product production schedule presents a significant challenge in managing process recipes across diverse products to address initial recipe bias, process shifts, and patterned disturbances occurring from run to run, product to product, and cycle to cycle. To tackle this issue, we propose a Run-to-Run controller for mixed-product operations, utilizing deep reinforcement learning strategies. This controller offers adaptability to evolving mixed-product environments and system dynamics, eliminating the need for explicit model updates. This adaptability is particularly crucial in real-world scenarios characterized by complex system dynamics and dynamic process disturbances, which may not be adequately addressed by conventional control methods such as EWMA-based controllers. The proposed controller is not only straightforward to implement but has also showcased its effectiveness in producing high-quality control outputs.","PeriodicalId":51060,"journal":{"name":"IEEE Transactions on Automation Science and Engineering","volume":"23 ","pages":"3299-3315"},"PeriodicalIF":7.9000,"publicationDate":"2026-01-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"IEEE Transactions on Automation Science and Engineering","FirstCategoryId":"94","ListUrlMain":"https://ieeexplore.ieee.org/document/10843296/","RegionNum":2,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"2025/1/15 0:00:00","PubModel":"Epub","JCR":"Q1","JCRName":"AUTOMATION & CONTROL SYSTEMS","Score":null,"Total":0}
引用次数: 0

Abstract

This paper proposes a new Run-to-Run (R2R) control framework based on deep deterministic policy gradient (DDPG) for the mixed-product production mode in semiconductor manufacturing. The DDPG algorithm is particularly developed to configure a deep reinforcement learning environment well suited to mixed-product production modes. To address the challenges posed in deep reinforcement learning, three enhanced mechanisms have been developed to improve the training of the proposed DDPG model for mixed-product R2R applications. These mechanisms include a piece-wise reward function, training with dynamic targets, and the new recall principle. It is demonstrated from the comprehensive simulation results that the proposed R2R control framework outperforms five noted mixed-product R2R control algorithms in the literature. The research outcome of this paper signifies a promising viability of deep reinforcement learning for highly complex and dynamic environments with continuous action spaces in the mixed-product R2R practice.Note to Practitioners—The mixed-product production schedule presents a significant challenge in managing process recipes across diverse products to address initial recipe bias, process shifts, and patterned disturbances occurring from run to run, product to product, and cycle to cycle. To tackle this issue, we propose a Run-to-Run controller for mixed-product operations, utilizing deep reinforcement learning strategies. This controller offers adaptability to evolving mixed-product environments and system dynamics, eliminating the need for explicit model updates. This adaptability is particularly crucial in real-world scenarios characterized by complex system dynamics and dynamic process disturbances, which may not be adequately addressed by conventional control methods such as EWMA-based controllers. The proposed controller is not only straightforward to implement but has also showcased its effectiveness in producing high-quality control outputs.
查看原文
分享 分享
微信好友 朋友圈 QQ好友 复制链接
本刊更多论文
半导体制造混合产品生产模式下一种新的深度强化学习运行控制算法
针对半导体制造中的混合产品生产模式,提出了一种基于深度确定性策略梯度(DDPG)的R2R控制框架。DDPG算法专门用于配置深度强化学习环境,非常适合混合产品生产模式。为了解决深度强化学习带来的挑战,研究人员开发了三种增强机制,以改进混合产品R2R应用中提出的DDPG模型的训练。这些机制包括分段奖励函数、动态目标训练和新的回忆原则。综合仿真结果表明,所提出的R2R控制框架优于文献中五种著名的混合产品R2R控制算法。本文的研究结果表明,深度强化学习在混合产品R2R实践中具有连续动作空间的高度复杂和动态环境中具有良好的可行性。从业人员注意事项——混合产品生产计划在管理不同产品的过程配方方面提出了重大挑战,以解决初始配方偏差、过程转移,以及从运行到运行、从产品到产品、从周期到周期发生的模式干扰。为了解决这个问题,我们提出了一个用于混合产品操作的Run-to-Run控制器,利用深度强化学习策略。该控制器提供了适应不断发展的混合产品环境和系统动力学,消除了明确的模型更新的需要。这种适应性在以复杂系统动力学和动态过程干扰为特征的现实场景中尤为重要,而传统的控制方法(如基于ewma的控制器)可能无法充分解决这些问题。所提出的控制器不仅易于实现,而且在产生高质量控制输出方面也显示出其有效性。
本文章由计算机程序翻译,如有差异,请以英文原文为准。
求助全文
约1分钟内获得全文 去求助
来源期刊
IEEE Transactions on Automation Science and Engineering
IEEE Transactions on Automation Science and Engineering 工程技术-自动化与控制系统
CiteScore
12.50
自引率
14.30%
发文量
404
审稿时长
3.0 months
期刊介绍: The IEEE Transactions on Automation Science and Engineering (T-ASE) publishes fundamental papers on Automation, emphasizing scientific results that advance efficiency, quality, productivity, and reliability. T-ASE encourages interdisciplinary approaches from computer science, control systems, electrical engineering, mathematics, mechanical engineering, operations research, and other fields. T-ASE welcomes results relevant to industries such as agriculture, biotechnology, healthcare, home automation, maintenance, manufacturing, pharmaceuticals, retail, security, service, supply chains, and transportation. T-ASE addresses a research community willing to integrate knowledge across disciplines and industries. For this purpose, each paper includes a Note to Practitioners that summarizes how its results can be applied or how they might be extended to apply in practice.
期刊最新文献
Non-model-based Finite-Time Adaptive Neural Output Feedback Control for an Active Ankle-Foot Orthosis Optimal K -Wafer Cyclic Sequence of Dual-Armed Cluster Tool with Purge Operation Advanced Multilayer Neurolearning of Disturbed Nonlinear Systems With State Safety Constraints Tripartite Hybrid Game-Theoretic Optimization for Integrated Vehicle-Station-Grid System with Charging Station Heterogeneity Decentralized Path Planning with Supervisory Coordination for Multi-AGV Systems in Non-Standardized Automated Warehouses
×
引用
GB/T 7714-2015
复制
MLA
复制
APA
复制
导出至
BibTeX EndNote RefMan NoteFirst NoteExpress
×
×
提示
您的信息不完整,为了账户安全,请先补充。
现在去补充
×
提示
您因"违规操作"
具体请查看互助需知
我知道了
×
提示
现在去查看 取消
×
提示
确定
0
微信
客服QQ
Book学术公众号 扫码关注我们
反馈
×
意见反馈
请填写您的意见或建议
请填写您的手机或邮箱
已复制链接
已复制链接
快去分享给好友吧!
我知道了
×
扫码分享
扫码分享
Book学术官方微信
Book学术官方微信
Book学术文献互助
Book学术文献互助群
群 号:604180095
Book学术
文献互助 智能选刊 最新文献 互助须知 联系我们:info@booksci.cn
Book学术提供免费学术资源搜索服务,方便国内外学者检索中英文文献。致力于提供最便捷和优质的服务体验。
Copyright © 2023 Book学术 All rights reserved.
ghs 京公网安备 11010802042870号 京ICP备2023020795号-1