基于语言特征的文本聚类方法

2011 IEEE International Conference on Cloud Computing and Intelligence Systems Pub Date : 2011-10-13 DOI:10.1109/CCIS.2011.6045042

Kansheng Shi, Lemin Li, Jie He, Haitao Liu, Naitong Zhang, Wentao Song

{"title":"基于语言特征的文本聚类方法","authors":"Kansheng Shi, Lemin Li, Jie He, Haitao Liu, Naitong Zhang, Wentao Song","doi":"10.1109/CCIS.2011.6045042","DOIUrl":null,"url":null,"abstract":"The traditional K-means algorithm is sensitive to the initial point, easy to fall into local optimum. In order to avoid this kind of flaw, an improved K-means text clustering method WIKTCM is proposed. The new method creates an innovative initial centers selection method and accommodates the contribution of characteristics of different parts of speech to the text. In addition, the impact of outliers is considered. Experimental results show that the new method has better clustering results.","PeriodicalId":128504,"journal":{"name":"2011 IEEE International Conference on Cloud Computing and Intelligence Systems","volume":"14 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2011-10-13","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"4","resultStr":"{\"title\":\"A linguistic feature based text clustering method\",\"authors\":\"Kansheng Shi, Lemin Li, Jie He, Haitao Liu, Naitong Zhang, Wentao Song\",\"doi\":\"10.1109/CCIS.2011.6045042\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"The traditional K-means algorithm is sensitive to the initial point, easy to fall into local optimum. In order to avoid this kind of flaw, an improved K-means text clustering method WIKTCM is proposed. The new method creates an innovative initial centers selection method and accommodates the contribution of characteristics of different parts of speech to the text. In addition, the impact of outliers is considered. Experimental results show that the new method has better clustering results.\",\"PeriodicalId\":128504,\"journal\":{\"name\":\"2011 IEEE International Conference on Cloud Computing and Intelligence Systems\",\"volume\":\"14 1\",\"pages\":\"0\"},\"PeriodicalIF\":0.0000,\"publicationDate\":\"2011-10-13\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"4\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"2011 IEEE International Conference on Cloud Computing and Intelligence Systems\",\"FirstCategoryId\":\"1085\",\"ListUrlMain\":\"https://doi.org/10.1109/CCIS.2011.6045042\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"\",\"JCRName\":\"\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"2011 IEEE International Conference on Cloud Computing and Intelligence Systems","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1109/CCIS.2011.6045042","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 4

摘要

传统的K-means算法对初始点敏感，容易陷入局部最优。为了避免这种缺陷，提出了一种改进的k均值文本聚类方法WIKTCM。新方法创造了一种创新的初始中心选择方法，并适应了不同词类特征对文本的贡献。此外，还考虑了异常值的影响。实验结果表明，新方法具有较好的聚类效果。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

查看原文

微信好友朋友圈 QQ好友复制链接

本刊更多论文

A linguistic feature based text clustering method

The traditional K-means algorithm is sensitive to the initial point, easy to fall into local optimum. In order to avoid this kind of flaw, an improved K-means text clustering method WIKTCM is proposed. The new method creates an innovative initial centers selection method and accommodates the contribution of characteristics of different parts of speech to the text. In addition, the impact of outliers is considered. Experimental results show that the new method has better clustering results.

求助全文

通过发布文献求助，成功后即可免费获取论文全文。去求助

来源期刊

2011 IEEE International Conference on Cloud Computing and Intelligence Systems

自引率

0.00%

发文量

期刊最新文献

A dynamic and integrated load-balancing scheduling algorithm for Cloud datacenters A CPU-GPU hybrid computing framework for real-time clothing animation The communication of CAN bus used in synchronization control of multi-motor based on DSP An improved dynamic provable data possession model Ensuring the data integrity in cloud data storage