基于词向量表示的税务相关领域新词发现

網際網路技術學刊 Pub Date : 2023-07-01 DOI:10.53106/160792642023072404010

Wei Wei Wei Wei, Wei Liu Wei Wei, Beibei Zhang Wei Liu, Rafał Scherer Beibei Zhang, Robertas Damaševičius Rafal Scherer

{"title":"基于词向量表示的税务相关领域新词发现","authors":"Wei Wei Wei Wei, Wei Liu Wei Wei, Beibei Zhang Wei Liu, Rafał Scherer Beibei Zhang, Robertas Damaševičius Rafal Scherer","doi":"10.53106/160792642023072404010","DOIUrl":null,"url":null,"abstract":"\n New words detection, as basic research in natural language processing, has gained extensive concern from academic and business communities. When the existing Chinese word segmentation technology is applied in the specific field of tax-related finance, because it cannot correctly identify new words in the field, it will have an impact on subsequent information extraction and entity recognition. Aiming at the current problems in new word discovery, it proposed a new word detection method using statistical features that are based on the inner measurement and branch entropy and then combined with word vector representation. First, perform word segmentation preprocessing on the corpus, calculate the internal cohesion degree of words through statistics of scattered string mutual information, filter out candidate two-tuples, and then filter and expand the two-tuples; next, it locks the boundaries of new words through calculate the branch entropy. Finally, expand the new vocabulary dictionary according to the cosine similarity principle of word vector representation. The unsupervised neologism discovery proposed in this paper allows for automatic growth of the neologism lexicon, experimental results on large-scale corpus verify the effectiveness of this method.\n \n","PeriodicalId":442331,"journal":{"name":"網際網路技術學刊","volume":"1 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2023-07-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":"{\"title\":\"Discovery of New Words in Tax-related Fields Based on Word Vector Representation\",\"authors\":\"Wei Wei Wei Wei, Wei Liu Wei Wei, Beibei Zhang Wei Liu, Rafał Scherer Beibei Zhang, Robertas Damaševičius Rafal Scherer\",\"doi\":\"10.53106/160792642023072404010\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"\\n New words detection, as basic research in natural language processing, has gained extensive concern from academic and business communities. When the existing Chinese word segmentation technology is applied in the specific field of tax-related finance, because it cannot correctly identify new words in the field, it will have an impact on subsequent information extraction and entity recognition. Aiming at the current problems in new word discovery, it proposed a new word detection method using statistical features that are based on the inner measurement and branch entropy and then combined with word vector representation. First, perform word segmentation preprocessing on the corpus, calculate the internal cohesion degree of words through statistics of scattered string mutual information, filter out candidate two-tuples, and then filter and expand the two-tuples; next, it locks the boundaries of new words through calculate the branch entropy. Finally, expand the new vocabulary dictionary according to the cosine similarity principle of word vector representation. The unsupervised neologism discovery proposed in this paper allows for automatic growth of the neologism lexicon, experimental results on large-scale corpus verify the effectiveness of this method.\\n \\n\",\"PeriodicalId\":442331,\"journal\":{\"name\":\"網際網路技術學刊\",\"volume\":\"1 1\",\"pages\":\"0\"},\"PeriodicalIF\":0.0000,\"publicationDate\":\"2023-07-01\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"0\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"網際網路技術學刊\",\"FirstCategoryId\":\"1085\",\"ListUrlMain\":\"https://doi.org/10.53106/160792642023072404010\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"\",\"JCRName\":\"\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"網際網路技術學刊","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.53106/160792642023072404010","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 0

摘要

新词检测作为自然语言处理的基础研究，受到了学术界和企业界的广泛关注。现有的中文分词技术在涉税金融特定领域应用时，由于无法正确识别该领域的新词，会对后续的信息提取和实体识别产生影响。针对当前新词发现中存在的问题，提出了一种基于内度量和分支熵的统计特征与词向量表示相结合的新词检测方法。首先对语料库进行分词预处理，通过统计分散的字符串互信息计算词的内部衔接度，过滤出候选双元组，然后对双元组进行过滤和扩展;其次，通过计算分支熵来锁定新词的边界。最后，根据词向量表示的余弦相似原理扩展新词汇字典。本文提出的无监督新词发现方法实现了新词词典的自动增长，在大规模语料库上的实验结果验证了该方法的有效性。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

查看原文

微信好友朋友圈 QQ好友复制链接

本刊更多论文

Discovery of New Words in Tax-related Fields Based on Word Vector Representation

New words detection, as basic research in natural language processing, has gained extensive concern from academic and business communities. When the existing Chinese word segmentation technology is applied in the specific field of tax-related finance, because it cannot correctly identify new words in the field, it will have an impact on subsequent information extraction and entity recognition. Aiming at the current problems in new word discovery, it proposed a new word detection method using statistical features that are based on the inner measurement and branch entropy and then combined with word vector representation. First, perform word segmentation preprocessing on the corpus, calculate the internal cohesion degree of words through statistics of scattered string mutual information, filter out candidate two-tuples, and then filter and expand the two-tuples; next, it locks the boundaries of new words through calculate the branch entropy. Finally, expand the new vocabulary dictionary according to the cosine similarity principle of word vector representation. The unsupervised neologism discovery proposed in this paper allows for automatic growth of the neologism lexicon, experimental results on large-scale corpus verify the effectiveness of this method.

求助全文

通过发布文献求助，成功后即可免费获取论文全文。去求助

来源期刊

網際網路技術學刊

自引率

0.00%

发文量