基于有限书目元数据的文本分类

2009 Fourth International Conference on Digital Information Management Pub Date : 2009-12-18 DOI:10.1109/ICDIM.2009.5356767

K. Denecke, T. Risse, Thomas Baehr

{"title":"基于有限书目元数据的文本分类","authors":"K. Denecke, T. Risse, Thomas Baehr","doi":"10.1109/ICDIM.2009.5356767","DOIUrl":null,"url":null,"abstract":"In this paper, we introduce a method for categorizing digital items according to their topic, only relying on the document's metadata, such as author name and title information. The proposed approach is based on a set of lexical resources constructed for our purposes (e.g., journal titles, conference names) and on a traditional machine-learning classifier that assigns one category to each document based on identified core features. The system is evaluated on a real-world data set and the influence of different feature combinations and settings is studied. Although the available information is limited, the results show that the approach is capable to efficiently classify data items representing documents.","PeriodicalId":300287,"journal":{"name":"2009 Fourth International Conference on Digital Information Management","volume":"1 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2009-12-18","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"6","resultStr":"{\"title\":\"Text classification based on limited bibliographic metadata\",\"authors\":\"K. Denecke, T. Risse, Thomas Baehr\",\"doi\":\"10.1109/ICDIM.2009.5356767\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"In this paper, we introduce a method for categorizing digital items according to their topic, only relying on the document's metadata, such as author name and title information. The proposed approach is based on a set of lexical resources constructed for our purposes (e.g., journal titles, conference names) and on a traditional machine-learning classifier that assigns one category to each document based on identified core features. The system is evaluated on a real-world data set and the influence of different feature combinations and settings is studied. Although the available information is limited, the results show that the approach is capable to efficiently classify data items representing documents.\",\"PeriodicalId\":300287,\"journal\":{\"name\":\"2009 Fourth International Conference on Digital Information Management\",\"volume\":\"1 1\",\"pages\":\"0\"},\"PeriodicalIF\":0.0000,\"publicationDate\":\"2009-12-18\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"6\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"2009 Fourth International Conference on Digital Information Management\",\"FirstCategoryId\":\"1085\",\"ListUrlMain\":\"https://doi.org/10.1109/ICDIM.2009.5356767\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"\",\"JCRName\":\"\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"2009 Fourth International Conference on Digital Information Management","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1109/ICDIM.2009.5356767","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 6

摘要

在本文中，我们介绍了一种根据主题对数字项目进行分类的方法，该方法仅依赖于文档的元数据，如作者姓名和标题信息。提出的方法基于为我们的目的而构建的一组词汇资源(例如，期刊标题，会议名称)和传统的机器学习分类器，该分类器根据识别的核心特征为每个文档分配一个类别。在实际数据集上对系统进行了评估，并研究了不同特征组合和设置的影响。尽管可用的信息有限，但结果表明，该方法能够有效地对表示文档的数据项进行分类。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

查看原文

微信好友朋友圈 QQ好友复制链接

本刊更多论文

Text classification based on limited bibliographic metadata

In this paper, we introduce a method for categorizing digital items according to their topic, only relying on the document's metadata, such as author name and title information. The proposed approach is based on a set of lexical resources constructed for our purposes (e.g., journal titles, conference names) and on a traditional machine-learning classifier that assigns one category to each document based on identified core features. The system is evaluated on a real-world data set and the influence of different feature combinations and settings is studied. Although the available information is limited, the results show that the approach is capable to efficiently classify data items representing documents.

求助全文

通过发布文献求助，成功后即可免费获取论文全文。去求助

来源期刊

2009 Fourth International Conference on Digital Information Management

自引率

0.00%

发文量

期刊最新文献

Ontology based entity disambiguation with natural language patterns Tiles — A model for classifying and using contextual information for context-aware applications Effectively and efficiently detect web page duplication From state-based to event-based contextual security policies P2P applied in CMS for advertising