On the Efficacy of Dynamic Behavior Comparison for Judging Functional Equivalence

2019 19th International Working Conference on Source Code Analysis and Manipulation (SCAM) Pub Date : 2019-09-01 DOI:10.1109/SCAM.2019.00030

Marcus Kessel C. Atkinson

{"title":"On the Efficacy of Dynamic Behavior Comparison for Judging Functional Equivalence","authors":"Marcus Kessel, C. Atkinson","doi":"10.1109/SCAM.2019.00030","DOIUrl":null,"url":null,"abstract":"Since it was first proposed in 1992 under the name of \"behavior sampling\", the idea of judging whether software systems are functionally equivalent by observing their responses to common stimuli (i.e. tests) has been used for a range of tasks such as software retrieval, functional redundancy measurement and semantic clone detection. However, its efficacy has only been studied in one small experiment, with limited generalizability, described in the original paper proposing the approach. The results of that experiment suggest that a relatively small number of randomly generated tests (i.e. 4) is sufficient to recognize non-functional-equivalent software 85% of the time. This number has therefore been adopted as \"sufficient\" in numerous applications of the approach. In this paper we present a much larger study which suggests at least 39 randomly generated tests are actually needed to achieve this level of effectiveness, but that a far fewer number of tests generated using coverage-based heuristics are sufficient. Since these results are much more generalizable, they have implications for future applications of behavioral sampling for dynamic behavior comparison.","PeriodicalId":431316,"journal":{"name":"2019 19th International Working Conference on Source Code Analysis and Manipulation (SCAM)","volume":"151 ","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2019-09-01","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"6","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"2019 19th International Working Conference on Source Code Analysis and Manipulation (SCAM)","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1109/SCAM.2019.00030","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 6

Abstract

Since it was first proposed in 1992 under the name of "behavior sampling", the idea of judging whether software systems are functionally equivalent by observing their responses to common stimuli (i.e. tests) has been used for a range of tasks such as software retrieval, functional redundancy measurement and semantic clone detection. However, its efficacy has only been studied in one small experiment, with limited generalizability, described in the original paper proposing the approach. The results of that experiment suggest that a relatively small number of randomly generated tests (i.e. 4) is sufficient to recognize non-functional-equivalent software 85% of the time. This number has therefore been adopted as "sufficient" in numerous applications of the approach. In this paper we present a much larger study which suggests at least 39 randomly generated tests are actually needed to achieve this level of effectiveness, but that a far fewer number of tests generated using coverage-based heuristics are sufficient. Since these results are much more generalizable, they have implications for future applications of behavioral sampling for dynamic behavior comparison.

查看原文

微信好友朋友圈 QQ好友复制链接

本刊更多论文

论动态行为比较对判断功能对等的有效性

自1992年以“行为抽样”的名义首次提出以来，通过观察软件系统对共同刺激(即测试)的反应来判断软件系统是否在功能上等同的想法已被用于一系列任务，如软件检索、功能冗余测量和语义克隆检测。然而，它的功效只在一个小实验中进行了研究，具有有限的普遍性，在提出该方法的原始论文中进行了描述。该实验的结果表明，相对少量的随机生成的测试(即4个)足以在85%的时间内识别非功能等效的软件。因此，在该方法的许多应用中，这个数字被认为是“足够的”。在这篇论文中，我们提出了一个更大的研究，它表明至少需要39个随机生成的测试来达到这个水平的有效性，但是使用基于覆盖率的启发式生成的测试数量要少得多就足够了。由于这些结果更具普遍性，它们对动态行为比较的行为抽样的未来应用具有启示意义。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文去求助