Xiali Li, Bo Liu, Zhi Wei, Zhaoqi Wang, Licheng Wu
{"title":"Tjong: A transformer-based Mahjong AI via hierarchical decision-making and fan backward","authors":"Xiali Li, Bo Liu, Zhi Wei, Zhaoqi Wang, Licheng Wu","doi":"10.1049/cit2.12298","DOIUrl":null,"url":null,"abstract":"<p>Mahjong, a complex game with hidden information and sparse rewards, poses significant challenges. Existing Mahjong AIs require substantial hardware resources and extensive datasets to enhance AI capabilities. The authors propose a transformer-based Mahjong AI (Tjong) via hierarchical decision-making. By utilising self-attention mechanisms, Tjong effectively captures tile patterns and game dynamics, and it decouples the decision process into two distinct stages: action decision and tile decision. This design reduces decision complexity considerably. Additionally, a fan backward technique is proposed to address the sparse rewards by allocating reversed rewards for actions based on winning hands. Tjong consists of 15M parameters and is trained using approximately 0.5 M data over 7 days of supervised learning on a single server with 2 GPUs. The action decision achieved an accuracy of 94.63%, while the claim decision attained 98.55% and the discard decision reached 81.51%. In a tournament format, Tjong outperformed AIs (CNN, MLP, RNN, ResNet, VIT), achieving scores up to 230% higher than its opponents. Furthermore, after 3 days of reinforcement learning training, it ranked within the top 1% on the leaderboard on the Botzone platform.</p>","PeriodicalId":46211,"journal":{"name":"CAAI Transactions on Intelligence Technology","volume":"9 4","pages":"982-995"},"PeriodicalIF":8.4000,"publicationDate":"2024-03-21","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"https://onlinelibrary.wiley.com/doi/epdf/10.1049/cit2.12298","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"CAAI Transactions on Intelligence Technology","FirstCategoryId":"94","ListUrlMain":"https://onlinelibrary.wiley.com/doi/10.1049/cit2.12298","RegionNum":2,"RegionCategory":"计算机科学","ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"Q1","JCRName":"COMPUTER SCIENCE, ARTIFICIAL INTELLIGENCE","Score":null,"Total":0}
引用次数: 0
Abstract
Mahjong, a complex game with hidden information and sparse rewards, poses significant challenges. Existing Mahjong AIs require substantial hardware resources and extensive datasets to enhance AI capabilities. The authors propose a transformer-based Mahjong AI (Tjong) via hierarchical decision-making. By utilising self-attention mechanisms, Tjong effectively captures tile patterns and game dynamics, and it decouples the decision process into two distinct stages: action decision and tile decision. This design reduces decision complexity considerably. Additionally, a fan backward technique is proposed to address the sparse rewards by allocating reversed rewards for actions based on winning hands. Tjong consists of 15M parameters and is trained using approximately 0.5 M data over 7 days of supervised learning on a single server with 2 GPUs. The action decision achieved an accuracy of 94.63%, while the claim decision attained 98.55% and the discard decision reached 81.51%. In a tournament format, Tjong outperformed AIs (CNN, MLP, RNN, ResNet, VIT), achieving scores up to 230% higher than its opponents. Furthermore, after 3 days of reinforcement learning training, it ranked within the top 1% on the leaderboard on the Botzone platform.
期刊介绍:
CAAI Transactions on Intelligence Technology is a leading venue for original research on the theoretical and experimental aspects of artificial intelligence technology. We are a fully open access journal co-published by the Institution of Engineering and Technology (IET) and the Chinese Association for Artificial Intelligence (CAAI) providing research which is openly accessible to read and share worldwide.