BIOptimus: Pre-training an Optimal Biomedical Language Model with Curriculum Learning for Named Entity Recognition

Workshop on Biomedical Natural Language Processing Pub Date : 2023-08-16 DOI:10.18653/v1/2023.bionlp-1.31

Vera Pavlova, M. Makhlouf

{"title":"BIOptimus: Pre-training an Optimal Biomedical Language Model with Curriculum Learning for Named Entity Recognition","authors":"Vera Pavlova, M. Makhlouf","doi":"10.18653/v1/2023.bionlp-1.31","DOIUrl":null,"url":null,"abstract":"Using language models (LMs) pre-trained in a self-supervised setting on large corpora and then fine-tuning for a downstream task has helped to deal with the problem of limited label data for supervised learning tasks such as Named Entity Recognition (NER). Recent research in biomedical language processing has offered a number of biomedical LMs pre-trained using different methods and techniques that advance results on many BioNLP tasks, including NER. However, there is still a lack of a comprehensive comparison of pre-training approaches that would work more optimally in the biomedical domain. This paper aims to investigate different pre-training methods, such as pre-training the biomedical LM from scratch and pre-training it in a continued fashion. We compare existing methods with our proposed pre-training method of initializing weights for new tokens by distilling existing weights from the BERT model inside the context where the tokens were found. The method helps to speed up the pre-training stage and improve performance on NER. In addition, we compare how masking rate, corruption strategy, and masking strategies impact the performance of the biomedical LM. Finally, using the insights from our experiments, we introduce a new biomedical LM (BIOptimus), which is pre-trained using Curriculum Learning (CL) and contextualized weight distillation method. Our model sets new states of the art on several biomedical Named Entity Recognition (NER) tasks. We release our code and all pre-trained models.","PeriodicalId":200974,"journal":{"name":"Workshop on Biomedical Natural Language Processing","volume":"24 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2023-08-16","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"Workshop on Biomedical Natural Language Processing","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.18653/v1/2023.bionlp-1.31","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 0

Abstract

Using language models (LMs) pre-trained in a self-supervised setting on large corpora and then fine-tuning for a downstream task has helped to deal with the problem of limited label data for supervised learning tasks such as Named Entity Recognition (NER). Recent research in biomedical language processing has offered a number of biomedical LMs pre-trained using different methods and techniques that advance results on many BioNLP tasks, including NER. However, there is still a lack of a comprehensive comparison of pre-training approaches that would work more optimally in the biomedical domain. This paper aims to investigate different pre-training methods, such as pre-training the biomedical LM from scratch and pre-training it in a continued fashion. We compare existing methods with our proposed pre-training method of initializing weights for new tokens by distilling existing weights from the BERT model inside the context where the tokens were found. The method helps to speed up the pre-training stage and improve performance on NER. In addition, we compare how masking rate, corruption strategy, and masking strategies impact the performance of the biomedical LM. Finally, using the insights from our experiments, we introduce a new biomedical LM (BIOptimus), which is pre-trained using Curriculum Learning (CL) and contextualized weight distillation method. Our model sets new states of the art on several biomedical Named Entity Recognition (NER) tasks. We release our code and all pre-trained models.

查看原文

微信好友朋友圈 QQ好友复制链接

本刊更多论文

基于课程学习的生物医学语言模型的命名实体识别预训练

在大型语料库的自监督设置中使用预训练的语言模型(LMs)，然后对下游任务进行微调，有助于处理有监督学习任务(如命名实体识别(NER))中标签数据有限的问题。最近在生物医学语言处理方面的研究已经提供了一些使用不同方法和技术进行预训练的生物医学LMs，这些方法和技术在许多BioNLP任务(包括NER)上取得了进步。然而，对于在生物医学领域更有效的预训练方法，仍然缺乏一个全面的比较。本文旨在探讨不同的预训练方法，如对生物医学LM进行从头预训练和持续预训练。我们将现有方法与我们提出的预训练方法进行比较，该方法通过在发现标记的上下文中从BERT模型中提取现有权重来初始化新标记的权重。该方法有助于加快预训练阶段，提高NER的性能。此外，我们比较了掩蔽率、腐败策略和掩蔽策略对生物医学LM性能的影响。最后，基于实验结果，我们引入了一种新的生物医学LM (BIOptimus)，该LM使用课程学习(CL)和情境化权重蒸馏方法进行预训练。我们的模型在几个生物医学命名实体识别(NER)任务上设置了新的技术状态。我们发布我们的代码和所有预先训练的模型。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文去求助

来源期刊

Workshop on Biomedical Natural Language Processing

自引率

0.00%

发文量