基于层次模糊集的深度Web源分类

2010 International Conference On Computer Design and Applications Pub Date : 2010-06-25 DOI:10.1109/ICCDA.2010.5541144

Hai-Long Wang, Liang Yue, Pengpeng Zhao, Zhi-ming Cui

{"title":"基于层次模糊集的深度Web源分类","authors":"Hai-Long Wang, Liang Yue, Pengpeng Zhao, Zhi-ming Cui","doi":"10.1109/ICCDA.2010.5541144","DOIUrl":null,"url":null,"abstract":"This paper presents a classification method of data source using fuzzy set and probabilistic model. The words of each domain are classified into characteristic words and general words according to their contribution to the current domain. The fuzzy set is introduced into the simplification process of characteristic words and the common words as the normalized glossary tool, which can be able to find more precise glossary in the homepage text. And a vocabulary probabilistic model is build after the normalized process in various domains, these words are classified by calculating the distance between the data source form vector and each domain vector.","PeriodicalId":190625,"journal":{"name":"2010 International Conference On Computer Design and Applications","volume":"81 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2010-06-25","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":"{\"title\":\"Hierachical fuzzy set-based deep Web source classification\",\"authors\":\"Hai-Long Wang, Liang Yue, Pengpeng Zhao, Zhi-ming Cui\",\"doi\":\"10.1109/ICCDA.2010.5541144\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"This paper presents a classification method of data source using fuzzy set and probabilistic model. The words of each domain are classified into characteristic words and general words according to their contribution to the current domain. The fuzzy set is introduced into the simplification process of characteristic words and the common words as the normalized glossary tool, which can be able to find more precise glossary in the homepage text. And a vocabulary probabilistic model is build after the normalized process in various domains, these words are classified by calculating the distance between the data source form vector and each domain vector.\",\"PeriodicalId\":190625,\"journal\":{\"name\":\"2010 International Conference On Computer Design and Applications\",\"volume\":\"81 1\",\"pages\":\"0\"},\"PeriodicalIF\":0.0000,\"publicationDate\":\"2010-06-25\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"0\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"2010 International Conference On Computer Design and Applications\",\"FirstCategoryId\":\"1085\",\"ListUrlMain\":\"https://doi.org/10.1109/ICCDA.2010.5541144\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"\",\"JCRName\":\"\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"2010 International Conference On Computer Design and Applications","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1109/ICCDA.2010.5541144","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 0

摘要

提出了一种基于模糊集和概率模型的数据源分类方法。每个领域的词根据其对当前领域的贡献分为特征词和一般词。将模糊集作为规范化词汇表工具引入到特征词和常用词的简化过程中，可以在主页文本中找到更精确的词汇表。在各领域进行归一化处理后，建立词汇概率模型，通过计算数据源形式向量与各领域向量之间的距离对词汇进行分类。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

查看原文

微信好友朋友圈 QQ好友复制链接

本刊更多论文

Hierachical fuzzy set-based deep Web source classification

This paper presents a classification method of data source using fuzzy set and probabilistic model. The words of each domain are classified into characteristic words and general words according to their contribution to the current domain. The fuzzy set is introduced into the simplification process of characteristic words and the common words as the normalized glossary tool, which can be able to find more precise glossary in the homepage text. And a vocabulary probabilistic model is build after the normalized process in various domains, these words are classified by calculating the distance between the data source form vector and each domain vector.

求助全文

通过发布文献求助，成功后即可免费获取论文全文。去求助

来源期刊

2010 International Conference On Computer Design and Applications

自引率

0.00%

发文量