Incomplete Information Competition Strategy Based on Improved Asynchronous Advantage Actor Critical Model

Proceedings of the 2020 4th International Conference on Deep Learning Technologies Pub Date : 2020-07-10 DOI:10.1145/3417188.3417189

Cong Zhao, Bing Xiao, Lin Zha

{"title":"Incomplete Information Competition Strategy Based on Improved Asynchronous Advantage Actor Critical Model","authors":"Cong Zhao, Bing Xiao, Lin Zha","doi":"10.1145/3417188.3417189","DOIUrl":null,"url":null,"abstract":"In recent years, game theory has been widely used in the field of deep learning, mainly including intelligent competition strategies of complete information games and incomplete information games. This paper focuses on incomplete information games, and proposes a low-dimensional semantic feature based on category coding and an incomplete information competition strategy based on the improved Asynchronous Advantage Actor-Critic (A3C) network model. First, the A3C network model in deep reinforcement learning is adopted in the competition strategy, and its network structure is improved according to the semantic features based on category coding. The improved A3C model is implemented in parallel by a series of \"workers\". The \"workers\" is a new deep learning model structure proposed in this paper. Secondly, this article combines supervised learning and Deep Reinforcement Learning (DRL) to propose a new competitive strategy. Through conducting a large number of real-time experiments with human players on online competitive websites, the comparison with the existing methods in terms of the ratio of winning and losing and the ranking rate, the experimental results indicate the superiority of the new competitive strategy.","PeriodicalId":373913,"journal":{"name":"Proceedings of the 2020 4th International Conference on Deep Learning Technologies","volume":"2015 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2020-07-10","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"Proceedings of the 2020 4th International Conference on Deep Learning Technologies","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1145/3417188.3417189","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 0

Abstract

In recent years, game theory has been widely used in the field of deep learning, mainly including intelligent competition strategies of complete information games and incomplete information games. This paper focuses on incomplete information games, and proposes a low-dimensional semantic feature based on category coding and an incomplete information competition strategy based on the improved Asynchronous Advantage Actor-Critic (A3C) network model. First, the A3C network model in deep reinforcement learning is adopted in the competition strategy, and its network structure is improved according to the semantic features based on category coding. The improved A3C model is implemented in parallel by a series of "workers". The "workers" is a new deep learning model structure proposed in this paper. Secondly, this article combines supervised learning and Deep Reinforcement Learning (DRL) to propose a new competitive strategy. Through conducting a large number of real-time experiments with human players on online competitive websites, the comparison with the existing methods in terms of the ratio of winning and losing and the ranking rate, the experimental results indicate the superiority of the new competitive strategy.

查看原文

微信好友朋友圈 QQ好友复制链接

本刊更多论文

基于改进异步优势参与者关键模型的不完全信息竞争策略

近年来，博弈论在深度学习领域得到了广泛的应用，主要包括完全信息博弈和不完全信息博弈的智能竞争策略。本文以不完全信息博弈为研究对象，提出了一种基于类别编码的低维语义特征和一种基于改进的异步优势参与者-批评者(A3C)网络模型的不完全信息竞争策略。首先，在竞争策略中采用深度强化学习中的A3C网络模型，并根据基于类别编码的语义特征对其网络结构进行改进。改进的A3C模型由一系列“工人”并行实现。“工人”是本文提出的一种新的深度学习模型结构。其次，本文将监督学习与深度强化学习(DRL)相结合，提出一种新的竞争策略。通过在在线竞技网站上对人类棋手进行大量的实时实验，并与现有方法在胜败比和排名率方面进行比较，实验结果表明了新竞技策略的优越性。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文去求助

来源期刊

Proceedings of the 2020 4th International Conference on Deep Learning Technologies

自引率

0.00%

发文量