{"title":"Integrating Biomedical Publications with Existing Metadata","authors":"Nikolay Nikolov, P. Stoehr","doi":"10.1109/CBMS.2008.127","DOIUrl":null,"url":null,"abstract":"Currently biomedical literature is largely disconnected from its metadata. While there are freely accessible centralised metadata repositories the publications themselves are split among a large number of repositories. We address this problem by harvesting freely accessible biomedical publications from the Web and integrating them with the corresponding metadata. The system involves title recognition applied on the harvested publications using knowledge-based algorithm and a fuzzy match between the extracted title and the metadata records using edit distance metric. So far we were able to locate +300.000 publications on the Web and achieve +96% precision and nearly 85% recall on a random sample of 250 documents harvested from the Web.","PeriodicalId":377855,"journal":{"name":"2008 21st IEEE International Symposium on Computer-Based Medical Systems","volume":"9 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2008-06-17","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"1","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"2008 21st IEEE International Symposium on Computer-Based Medical Systems","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1109/CBMS.2008.127","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}
引用次数: 1
Abstract
Currently biomedical literature is largely disconnected from its metadata. While there are freely accessible centralised metadata repositories the publications themselves are split among a large number of repositories. We address this problem by harvesting freely accessible biomedical publications from the Web and integrating them with the corresponding metadata. The system involves title recognition applied on the harvested publications using knowledge-based algorithm and a fuzzy match between the extracted title and the metadata records using edit distance metric. So far we were able to locate +300.000 publications on the Web and achieve +96% precision and nearly 85% recall on a random sample of 250 documents harvested from the Web.