Scholarly big data quality assessment: a case study of document linking and conflation with S2ORC

Jian Wu, Ryan Hiltabrand, Dominik Soós, C. Lee Giles
{"title":"Scholarly big data quality assessment: a case study of document linking and conflation with S2ORC","authors":"Jian Wu, Ryan Hiltabrand, Dominik Soós, C. Lee Giles","doi":"10.1145/3558100.3563850","DOIUrl":null,"url":null,"abstract":"Recently, the Allen Institute for Artificial Intelligence released the Semantic Scholar Open Research Corpus (S2ORC), one of the largest open-access scholarly big datasets with more than 130 million scholarly paper records. S2ORC contains a significant portion of automatically generated metadata. The metadata quality could impact downstream tasks such as citation analysis, citation prediction, and link analysis. In this project, we assess the document linking quality and estimate the document conflation rate for the S2ORC dataset. Using semi-automatically curated ground truth corpora, we estimated that the overall document linking quality is high, with 92.6% of documents correctly linking to six major databases, but the linking quality varies depending on subject domains. The document conflation rate is around 2.6%, meaning that about 97.4% of documents are unique. We further quantitatively compared three near-duplicate detection methods using the ground truth created from S2ORC. The experiments indicated that locality-sensitive hashing was the best method in terms of effectiveness and scalability, achieving high performance (F1=0.960) and a much reduced runtime. Our code and data are available at https://github.com/lamps-lab/docconflation.","PeriodicalId":146244,"journal":{"name":"Proceedings of the 22nd ACM Symposium on Document Engineering","volume":"55 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2022-09-20","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"1","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"Proceedings of the 22nd ACM Symposium on Document Engineering","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1145/3558100.3563850","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}
引用次数: 1

Abstract

Recently, the Allen Institute for Artificial Intelligence released the Semantic Scholar Open Research Corpus (S2ORC), one of the largest open-access scholarly big datasets with more than 130 million scholarly paper records. S2ORC contains a significant portion of automatically generated metadata. The metadata quality could impact downstream tasks such as citation analysis, citation prediction, and link analysis. In this project, we assess the document linking quality and estimate the document conflation rate for the S2ORC dataset. Using semi-automatically curated ground truth corpora, we estimated that the overall document linking quality is high, with 92.6% of documents correctly linking to six major databases, but the linking quality varies depending on subject domains. The document conflation rate is around 2.6%, meaning that about 97.4% of documents are unique. We further quantitatively compared three near-duplicate detection methods using the ground truth created from S2ORC. The experiments indicated that locality-sensitive hashing was the best method in terms of effectiveness and scalability, achieving high performance (F1=0.960) and a much reduced runtime. Our code and data are available at https://github.com/lamps-lab/docconflation.
查看原文
分享 分享
微信好友 朋友圈 QQ好友 复制链接
本刊更多论文
学术大数据质量评估:以S2ORC文件链接与合并为例
最近,艾伦人工智能研究所发布了语义学者开放研究语料库(S2ORC),这是最大的开放获取学术大数据集之一,拥有超过1.3亿篇学术论文记录。S2ORC包含大量自动生成的元数据。元数据质量会影响下游任务,如引文分析、引文预测和链接分析。在这个项目中,我们评估了S2ORC数据集的文档链接质量并估计了文档合并率。使用半自动整理的真实语料库,我们估计总体文档链接质量很高,92.6%的文档正确链接到六个主要数据库,但链接质量因主题领域而异。文档合并率约为2.6%,这意味着大约97.4%的文档是唯一的。我们进一步使用S2ORC创建的地面真值定量比较了三种近重复检测方法。实验表明,在有效性和可扩展性方面,位置敏感哈希是最好的方法,实现了高性能(F1=0.960)和更短的运行时间。我们的代码和数据可在https://github.com/lamps-lab/docconflation上获得。
本文章由计算机程序翻译,如有差异,请以英文原文为准。
求助全文
约1分钟内获得全文 去求助
来源期刊
自引率
0.00%
发文量
0
期刊最新文献
How did dennis ritchie produce his PhD thesis?: a typographical mystery From print to online newspapers on small displays: a layout generation approach aimed at preserving entry points Binarization of photographed documents image quality, processing time and size assessment Tab this folder of documents: page stream segmentation of business documents Graphical document representation for french newsletters analysis
×
引用
GB/T 7714-2015
复制
MLA
复制
APA
复制
导出至
BibTeX EndNote RefMan NoteFirst NoteExpress
×
×
提示
您的信息不完整,为了账户安全,请先补充。
现在去补充
×
提示
您因"违规操作"
具体请查看互助需知
我知道了
×
提示
现在去查看 取消
×
提示
确定
0
微信
客服QQ
Book学术公众号 扫码关注我们
反馈
×
意见反馈
请填写您的意见或建议
请填写您的手机或邮箱
已复制链接
已复制链接
快去分享给好友吧!
我知道了
×
扫码分享
扫码分享
Book学术官方微信
Book学术文献互助
Book学术文献互助群
群 号:481959085
Book学术
文献互助 智能选刊 最新文献 互助须知 联系我们:info@booksci.cn
Book学术提供免费学术资源搜索服务,方便国内外学者检索中英文文献。致力于提供最便捷和优质的服务体验。
Copyright © 2023 Book学术 All rights reserved.
ghs 京公网安备 11010802042870号 京ICP备2023020795号-1