Evaluating Commit Message Generation: To BLEU Or Not To BLEU?

2022 IEEE/ACM 44th International Conference on Software Engineering: New Ideas and Emerging Results (ICSE-NIER) Pub Date : 2022-04-20 DOI:10.1145/3510455.3512790

Samanta Dey, Venkatesh Vinayakarao, Monika Gupta, Sampath Dechu

{"title":"Evaluating Commit Message Generation: To BLEU Or Not To BLEU?","authors":"Samanta Dey, Venkatesh Vinayakarao, Monika Gupta, Sampath Dechu","doi":"10.1145/3510455.3512790","DOIUrl":null,"url":null,"abstract":"Commit messages play an important role in several software engineering tasks such as program comprehension and understanding program evolution. However, programmers neglect to write good commit messages. Hence, several Commit Message Generation (CMG) tools have been proposed. We observe that the recent state of the art CMG tools use simple and easy to compute automated evaluation metrics such as BLEU4 or its variants. The advances in the field of Machine Translation (MT) indicate several weaknesses of BLEU4 and its variants. They also propose several other metrics for evaluating Natural Language Generation (NLG) tools. In this work, we discuss the suitability of various MT metrics for the CMG task. Based on the insights from our experiments, we propose a new variant specifically for evaluating the CMG task. We re-evaluate the state of the art CMG tools on our new metric. We believe that our work fixes an important gap that exists in the understanding of evaluation metrics for CMG research. CCS CONCEPTS• Software and its engineering $\\rightarrow$Software verification and validation.","PeriodicalId":416186,"journal":{"name":"2022 IEEE/ACM 44th International Conference on Software Engineering: New Ideas and Emerging Results (ICSE-NIER)","volume":"12 6 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2022-04-20","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"1","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"2022 IEEE/ACM 44th International Conference on Software Engineering: New Ideas and Emerging Results (ICSE-NIER)","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1145/3510455.3512790","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 1

Abstract

Commit messages play an important role in several software engineering tasks such as program comprehension and understanding program evolution. However, programmers neglect to write good commit messages. Hence, several Commit Message Generation (CMG) tools have been proposed. We observe that the recent state of the art CMG tools use simple and easy to compute automated evaluation metrics such as BLEU4 or its variants. The advances in the field of Machine Translation (MT) indicate several weaknesses of BLEU4 and its variants. They also propose several other metrics for evaluating Natural Language Generation (NLG) tools. In this work, we discuss the suitability of various MT metrics for the CMG task. Based on the insights from our experiments, we propose a new variant specifically for evaluating the CMG task. We re-evaluate the state of the art CMG tools on our new metric. We believe that our work fixes an important gap that exists in the understanding of evaluation metrics for CMG research. CCS CONCEPTS• Software and its engineering $\rightarrow$Software verification and validation.

查看原文

微信好友朋友圈 QQ好友复制链接

本刊更多论文

评估提交消息的生成:去BLEU还是不去BLEU?

提交消息在一些软件工程任务中扮演着重要的角色，例如程序理解和理解程序演化。然而，程序员忽略了编写好的提交消息。因此，提出了几种提交消息生成(Commit Message Generation, CMG)工具。我们观察到，最新的最先进的CMG工具使用简单和容易的方法来计算自动评估指标，如BLEU4或它的变体。机器翻译(MT)领域的进步表明了BLEU4及其变体的一些弱点。他们还提出了评估自然语言生成(NLG)工具的其他几个指标。在这项工作中，我们讨论了各种MT指标对CMG任务的适用性。基于我们实验的见解，我们提出了一个新的变体，专门用于评估CMG任务。我们根据新指标重新评估CMG工具的现状。我们认为，我们的工作弥补了对CMG研究评估指标理解上的一个重要空白。CCS CONCEPTS•软件及其工程$\右箭头$软件验证和确认。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文去求助

来源期刊

2022 IEEE/ACM 44th International Conference on Software Engineering: New Ideas and Emerging Results (ICSE-NIER)

自引率

0.00%

发文量

期刊最新文献

Investigating User Perceptions of Conversational Agents for Software-related ExploratoryWeb Search Black Box Technique to Reduce Energy Consumption of Android Apps Utilizing Persistence for Post Facto Suppression of Invalid Anomalies Using System Logs Statistical Reasoning About Programs Title Page iii