A Transformer-Based Longer Entity Attention Model for Chinese Named Entity Recognition in Aerospace

2022 5th International Conference on Advanced Electronic Materials, Computers and Software Engineering (AEMCSE) Pub Date : 2022-04-01 DOI:10.1109/AEMCSE55572.2022.00077

Shuai Gong, Xiong Xiong, Yunfei Liu, Shengyang Li, Anqi Liu

{"title":"A Transformer-Based Longer Entity Attention Model for Chinese Named Entity Recognition in Aerospace","authors":"Shuai Gong, Xiong Xiong, Yunfei Liu, Shengyang Li, Anqi Liu","doi":"10.1109/AEMCSE55572.2022.00077","DOIUrl":null,"url":null,"abstract":"Chinese aerospace knowledge includes many long entities, such as professional terms, equipment names, and cabinets. However, current Named Entity Recognition (NER) algorithms typically address these longer and shorter entities uniformly. In this paper, a Longer Entity Attention (LEA) model based on the transformer is proposed. After the transformer encoding layer, LEA integrates sentence tags, sets thresholds according to the length of entities, and processes the hidden layer features of entities larger than the defined threshold to enhance the ability of the model to recognize longer entities. In addition, we construct an Aerospace Chinese NER dataset (ACNE) containing rich entity categories and domain knowledge. Experimental results demonstrate that LEA outperforms previous state-of-the-art models on ACNE, and shows a significant improvement on longer entities in each threshold range on OntoNotes 5.0 and ACNE datasets.","PeriodicalId":309096,"journal":{"name":"2022 5th International Conference on Advanced Electronic Materials, Computers and Software Engineering (AEMCSE)","volume":"111 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2022-04-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"2022 5th International Conference on Advanced Electronic Materials, Computers and Software Engineering (AEMCSE)","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1109/AEMCSE55572.2022.00077","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 0

Abstract

Chinese aerospace knowledge includes many long entities, such as professional terms, equipment names, and cabinets. However, current Named Entity Recognition (NER) algorithms typically address these longer and shorter entities uniformly. In this paper, a Longer Entity Attention (LEA) model based on the transformer is proposed. After the transformer encoding layer, LEA integrates sentence tags, sets thresholds according to the length of entities, and processes the hidden layer features of entities larger than the defined threshold to enhance the ability of the model to recognize longer entities. In addition, we construct an Aerospace Chinese NER dataset (ACNE) containing rich entity categories and domain knowledge. Experimental results demonstrate that LEA outperforms previous state-of-the-art models on ACNE, and shows a significant improvement on longer entities in each threshold range on OntoNotes 5.0 and ACNE datasets.

查看原文

微信好友朋友圈 QQ好友复制链接

本刊更多论文

基于变换的航天中文命名实体识别长实体注意模型

中国的航天知识包括许多长实体，如专业术语、设备名称和机柜。然而，当前的命名实体识别(NER)算法通常统一地处理这些较长和较短的实体。本文提出了一种基于变压器的长实体注意(LEA)模型。在转换编码层之后，LEA集成句子标签，根据实体的长度设置阈值，对大于定义阈值的实体的隐藏层特征进行处理，增强模型对较长实体的识别能力。此外，我们构建了一个包含丰富实体类别和领域知识的航空航天中文NER数据集(ACNE)。实验结果表明，LEA在痤疮上优于以前的最先进的模型，并且在OntoNotes 5.0和痤疮数据集上，在每个阈值范围内对较长的实体都有显着改善。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文去求助

来源期刊

2022 5th International Conference on Advanced Electronic Materials, Computers and Software Engineering (AEMCSE)

自引率

0.00%

发文量