Group-based corpus scheduling for parallel fuzzing

软件产业与工程 Pub Date : 2022-11-07 DOI:10.1145/3540250.3560885

Taotao Gu, Xiang Li, Shuaibing Lu, Jianwen Tian, Yuanping Nie, Xiaohui Kuang, Zhechao Lin, Chenyifan Liu, Jie Liang, Yu Jiang

{"title":"Group-based corpus scheduling for parallel fuzzing","authors":"Taotao Gu, Xiang Li, Shuaibing Lu, Jianwen Tian, Yuanping Nie, Xiaohui Kuang, Zhechao Lin, Chenyifan Liu, Jie Liang, Yu Jiang","doi":"10.1145/3540250.3560885","DOIUrl":null,"url":null,"abstract":"Parallel fuzzing relies on hardware resources to guarantee test throughput and efficiency. In industrial practice, it is well known that parallel fuzzing faces the challenge of task division, but most works neglect the important process of corpus allocation. In this paper, we proposed a group-based corpus scheduling strategy to address these two issues, which has been accepted by the LLVM community. And we implement a parallel fuzzer based on this strategy called glibFuzzer. glibFuzzer first groups the global corpus into different subsets and then assigns different energy scores and different scores to them. The energy scores were mainly determined by the seed size and the length of coverage information, and the difference score can describe the degree of difference in the code covered by different subsets of seeds. In each round of key local corpus construction, the master node selects high-quality seeds by combining the two scores to improve test efficiency and avoid task conflict. To prove the effectiveness of the strategy, we conducted an extensive evaluation on the real-world programs and FuzzBench. After 4×24 CPU-hours, glibFuzzer covered 22.02% more branches and executed 19.42 times more test cases than libFuzzer in 18 real-world programs. glibFuzzer showed an average branch coverage increase of 73.02%, 55.02%, 55.86% over AFL, PAFL, UniFuzz, respectively. More importantly, glibFuzzer found over 100 unique vulnerabilities.","PeriodicalId":68155,"journal":{"name":"软件产业与工程","volume":null,"pages":null},"PeriodicalIF":0.0000,"publicationDate":"2022-11-07","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"软件产业与工程","FirstCategoryId":"1089","ListUrlMain":"https://doi.org/10.1145/3540250.3560885","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 0

Abstract

Parallel fuzzing relies on hardware resources to guarantee test throughput and efficiency. In industrial practice, it is well known that parallel fuzzing faces the challenge of task division, but most works neglect the important process of corpus allocation. In this paper, we proposed a group-based corpus scheduling strategy to address these two issues, which has been accepted by the LLVM community. And we implement a parallel fuzzer based on this strategy called glibFuzzer. glibFuzzer first groups the global corpus into different subsets and then assigns different energy scores and different scores to them. The energy scores were mainly determined by the seed size and the length of coverage information, and the difference score can describe the degree of difference in the code covered by different subsets of seeds. In each round of key local corpus construction, the master node selects high-quality seeds by combining the two scores to improve test efficiency and avoid task conflict. To prove the effectiveness of the strategy, we conducted an extensive evaluation on the real-world programs and FuzzBench. After 4×24 CPU-hours, glibFuzzer covered 22.02% more branches and executed 19.42 times more test cases than libFuzzer in 18 real-world programs. glibFuzzer showed an average branch coverage increase of 73.02%, 55.02%, 55.86% over AFL, PAFL, UniFuzz, respectively. More importantly, glibFuzzer found over 100 unique vulnerabilities.

查看原文

微信好友朋友圈 QQ好友复制链接

本刊更多论文

基于组的语料库并行模糊调度

并行模糊测试依靠硬件资源来保证测试的吞吐量和效率。在工业实践中，并行模糊算法面临着任务划分的挑战，但大多数研究忽略了语料库分配这一重要过程。在本文中，我们提出了一种基于组的语料库调度策略来解决这两个问题，该策略已被LLVM社区所接受。我们基于这个策略实现了一个并行模糊器叫做glibFuzzer。glibFuzzer首先将全局语料库分成不同的子集，然后为它们分配不同的能量分数和不同的分数。能量分数主要由种子的大小和覆盖信息的长度决定，差异分数可以描述不同种子子集所覆盖代码的差异程度。在每一轮关键局部语料库构建中，主节点结合两个分数选择优质种子，提高测试效率，避免任务冲突。为了证明该策略的有效性，我们对现实世界的程序和FuzzBench进行了广泛的评估。在使用4×24 cpu小时后，在18个实际程序中，glibFuzzer覆盖的分支比libFuzzer多22.02%，执行的测试用例比libFuzzer多19.42倍。glibFuzzer比AFL、PAFL、unifuzzer的平均枝覆盖率分别提高了73.02%、55.02%、55.86%。更重要的是，glibFuzzer发现了100多个独特的漏洞。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文去求助

来源期刊

软件产业与工程

自引率

0.00%

发文量

676