{"title":"A New Deep Reinforcement Learning Run-to-Run Control Algorithm for Mixed-Product Production Mode in Semiconductor Manufacturing","authors":"Shu-Kai S. Fan;Tzu-Jung Chen","doi":"10.1109/TASE.2025.3526675","DOIUrl":null,"url":null,"abstract":"This paper proposes a new Run-to-Run (R2R) control framework based on deep deterministic policy gradient (DDPG) for the mixed-product production mode in semiconductor manufacturing. The DDPG algorithm is particularly developed to configure a deep reinforcement learning environment well suited to mixed-product production modes. To address the challenges posed in deep reinforcement learning, three enhanced mechanisms have been developed to improve the training of the proposed DDPG model for mixed-product R2R applications. These mechanisms include a piece-wise reward function, training with dynamic targets, and the new recall principle. It is demonstrated from the comprehensive simulation results that the proposed R2R control framework outperforms five noted mixed-product R2R control algorithms in the literature. The research outcome of this paper signifies a promising viability of deep reinforcement learning for highly complex and dynamic environments with continuous action spaces in the mixed-product R2R practice.Note to Practitioners—The mixed-product production schedule presents a significant challenge in managing process recipes across diverse products to address initial recipe bias, process shifts, and patterned disturbances occurring from run to run, product to product, and cycle to cycle. To tackle this issue, we propose a Run-to-Run controller for mixed-product operations, utilizing deep reinforcement learning strategies. This controller offers adaptability to evolving mixed-product environments and system dynamics, eliminating the need for explicit model updates. This adaptability is particularly crucial in real-world scenarios characterized by complex system dynamics and dynamic process disturbances, which may not be adequately addressed by conventional control methods such as EWMA-based controllers. The proposed controller is not only straightforward to implement but has also showcased its effectiveness in producing high-quality control outputs.","PeriodicalId":51060,"journal":{"name":"IEEE Transactions on Automation Science and Engineering","volume":"23 ","pages":"3299-3315"},"PeriodicalIF":7.9000,"publicationDate":"2026-01-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"IEEE Transactions on Automation Science and Engineering","FirstCategoryId":"94","ListUrlMain":"https://ieeexplore.ieee.org/document/10843296/","RegionNum":2,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"2025/1/15 0:00:00","PubModel":"Epub","JCR":"Q1","JCRName":"AUTOMATION & CONTROL SYSTEMS","Score":null,"Total":0}
引用次数: 0
Abstract
This paper proposes a new Run-to-Run (R2R) control framework based on deep deterministic policy gradient (DDPG) for the mixed-product production mode in semiconductor manufacturing. The DDPG algorithm is particularly developed to configure a deep reinforcement learning environment well suited to mixed-product production modes. To address the challenges posed in deep reinforcement learning, three enhanced mechanisms have been developed to improve the training of the proposed DDPG model for mixed-product R2R applications. These mechanisms include a piece-wise reward function, training with dynamic targets, and the new recall principle. It is demonstrated from the comprehensive simulation results that the proposed R2R control framework outperforms five noted mixed-product R2R control algorithms in the literature. The research outcome of this paper signifies a promising viability of deep reinforcement learning for highly complex and dynamic environments with continuous action spaces in the mixed-product R2R practice.Note to Practitioners—The mixed-product production schedule presents a significant challenge in managing process recipes across diverse products to address initial recipe bias, process shifts, and patterned disturbances occurring from run to run, product to product, and cycle to cycle. To tackle this issue, we propose a Run-to-Run controller for mixed-product operations, utilizing deep reinforcement learning strategies. This controller offers adaptability to evolving mixed-product environments and system dynamics, eliminating the need for explicit model updates. This adaptability is particularly crucial in real-world scenarios characterized by complex system dynamics and dynamic process disturbances, which may not be adequately addressed by conventional control methods such as EWMA-based controllers. The proposed controller is not only straightforward to implement but has also showcased its effectiveness in producing high-quality control outputs.
期刊介绍:
The IEEE Transactions on Automation Science and Engineering (T-ASE) publishes fundamental papers on Automation, emphasizing scientific results that advance efficiency, quality, productivity, and reliability. T-ASE encourages interdisciplinary approaches from computer science, control systems, electrical engineering, mathematics, mechanical engineering, operations research, and other fields. T-ASE welcomes results relevant to industries such as agriculture, biotechnology, healthcare, home automation, maintenance, manufacturing, pharmaceuticals, retail, security, service, supply chains, and transportation. T-ASE addresses a research community willing to integrate knowledge across disciplines and industries. For this purpose, each paper includes a Note to Practitioners that summarizes how its results can be applied or how they might be extended to apply in practice.