On the Effectiveness of Pretrained Models for API Learning

2022 IEEE/ACM 30th International Conference on Program Comprehension (ICPC) Pub Date : 2022-04-05 DOI:10.1145/3524610.3527886

M. Hadi, Imam Nur Bani Yusuf, Ferdian Thung, K. Luong, Lingxiao Jiang, F. H. Fard, David Lo

{"title":"On the Effectiveness of Pretrained Models for API Learning","authors":"M. Hadi, Imam Nur Bani Yusuf, Ferdian Thung, K. Luong, Lingxiao Jiang, F. H. Fard, David Lo","doi":"10.1145/3524610.3527886","DOIUrl":null,"url":null,"abstract":"Developers frequently use APIs to implement certain functionalities, such as parsing Excel Files, reading and writing text files line by line, etc. Developers can greatly benefit from automatic API usage sequence generation based on natural language queries for building applications in a faster and cleaner manner. Existing approaches utilize information retrieval models to search for matching API sequences given a query or use RNN-based encoder-decoder to generate API sequences. As it stands, the first approach treats queries and API names as bags of words. It lacks deep comprehension of the semantics of the queries. The latter approach adapts a neural language model to encode a user query into a fixed-length context vector and generate API sequences from the context vector. We want to understand the effectiveness of recent Pre-trained Transformer based Models (PTMs) for the API learning task. These PTMs are trained on large natural language corpora in an unsupervised manner to retain contextual knowledge about the language and have found success in solving similar Natural Language Processing (NLP) problems. However, the applicability of PTMs has not yet been explored for the API sequence generation task. We use a dataset that contains 7 million annotations collected from GitHub to evaluate the PTMs empirically. This dataset was also used to assess previous approaches. Based on our results, PTMs generate more accurate API sequences and outperform other related methods by ∼11%. We have also identified two different tokenization approaches that can contribute to a significant boost in PTMs' performance for the API sequence generation task.","PeriodicalId":426634,"journal":{"name":"2022 IEEE/ACM 30th International Conference on Program Comprehension (ICPC)","volume":"80 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2022-04-05","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"7","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"2022 IEEE/ACM 30th International Conference on Program Comprehension (ICPC)","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1145/3524610.3527886","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 7

Abstract

Developers frequently use APIs to implement certain functionalities, such as parsing Excel Files, reading and writing text files line by line, etc. Developers can greatly benefit from automatic API usage sequence generation based on natural language queries for building applications in a faster and cleaner manner. Existing approaches utilize information retrieval models to search for matching API sequences given a query or use RNN-based encoder-decoder to generate API sequences. As it stands, the first approach treats queries and API names as bags of words. It lacks deep comprehension of the semantics of the queries. The latter approach adapts a neural language model to encode a user query into a fixed-length context vector and generate API sequences from the context vector. We want to understand the effectiveness of recent Pre-trained Transformer based Models (PTMs) for the API learning task. These PTMs are trained on large natural language corpora in an unsupervised manner to retain contextual knowledge about the language and have found success in solving similar Natural Language Processing (NLP) problems. However, the applicability of PTMs has not yet been explored for the API sequence generation task. We use a dataset that contains 7 million annotations collected from GitHub to evaluate the PTMs empirically. This dataset was also used to assess previous approaches. Based on our results, PTMs generate more accurate API sequences and outperform other related methods by ∼11%. We have also identified two different tokenization approaches that can contribute to a significant boost in PTMs' performance for the API sequence generation task.

查看原文

微信好友朋友圈 QQ好友复制链接

本刊更多论文

预训练模型在API学习中的有效性研究

开发人员经常使用api来实现某些功能，例如解析Excel文件，逐行读写文本文件等。开发人员可以从基于自然语言查询的自动API使用序列生成中受益匪浅，从而以更快、更清晰的方式构建应用程序。现有的方法利用信息检索模型来搜索给定查询的匹配API序列或使用基于rnn的编码器-解码器来生成API序列。目前，第一种方法将查询和API名称视为单词包。它缺乏对查询语义的深刻理解。后一种方法采用神经语言模型将用户查询编码为固定长度的上下文向量，并从上下文向量生成API序列。我们想了解最近基于预训练的变压器模型(ptm)在API学习任务中的有效性。这些ptm以无监督的方式在大型自然语言语料库上进行训练，以保留有关语言的上下文知识，并在解决类似的自然语言处理(NLP)问题上取得了成功。然而，ptm在API序列生成任务中的适用性尚未得到探讨。我们使用从GitHub收集的包含700万个注释的数据集来对ptm进行经验评估。该数据集也用于评估以前的方法。根据我们的研究结果，PTMs生成更准确的API序列，并且比其他相关方法高出约11%。我们还确定了两种不同的标记化方法，它们可以显著提高ptm在API序列生成任务中的性能。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文去求助

来源期刊

2022 IEEE/ACM 30th International Conference on Program Comprehension (ICPC)

自引率

0.00%

发文量

期刊最新文献

Context-based Cluster Fault Localization Fine-Grained Code-Comment Semantic Interaction Analysis Find Bugs in Static Bug Finders Self-Supervised Learning of Smart Contract Representations An Exploratory Study of Analyzing JavaScript Online Code Clones