{"title":"基于生成模型和判别模型的中文垃圾邮件过滤比较","authors":"Yong Han, Yingying Wang, Huafu Ding, Haoliang Qi","doi":"10.1109/IALP.2011.64","DOIUrl":null,"url":null,"abstract":"Previous studies have shown that discriminative model is better than generative model for spam filtering, which is tested on the English dataset. But the study on Chinese Spam Filter is rare. So we compared the performance of Bogo: a classical generative model, Logistic Regression (LR) and Relaxed Online SVM (ROSVM): two typical discriminative models on the Chinese dataset. Bogo system adopts a generative model, which is based on Bayesian algorithm. We choose the public Chinese datasets: TREC06c, SEWM 2008, SEWM 2010, SEWM 2011, as the test dataset with immediate feedback. The discriminative model gives the better results than the generative model based on spam filter. ROSVM gives the best performance on Chinese spam filter.","PeriodicalId":297167,"journal":{"name":"2011 International Conference on Asian Language Processing","volume":"62 1","pages":"0"},"PeriodicalIF":0.0000,"publicationDate":"2011-11-15","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"1","resultStr":"{\"title\":\"The Comparison of Chinese Spam Filter Based on Generative Model and Discriminative Model\",\"authors\":\"Yong Han, Yingying Wang, Huafu Ding, Haoliang Qi\",\"doi\":\"10.1109/IALP.2011.64\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"Previous studies have shown that discriminative model is better than generative model for spam filtering, which is tested on the English dataset. But the study on Chinese Spam Filter is rare. So we compared the performance of Bogo: a classical generative model, Logistic Regression (LR) and Relaxed Online SVM (ROSVM): two typical discriminative models on the Chinese dataset. Bogo system adopts a generative model, which is based on Bayesian algorithm. We choose the public Chinese datasets: TREC06c, SEWM 2008, SEWM 2010, SEWM 2011, as the test dataset with immediate feedback. The discriminative model gives the better results than the generative model based on spam filter. ROSVM gives the best performance on Chinese spam filter.\",\"PeriodicalId\":297167,\"journal\":{\"name\":\"2011 International Conference on Asian Language Processing\",\"volume\":\"62 1\",\"pages\":\"0\"},\"PeriodicalIF\":0.0000,\"publicationDate\":\"2011-11-15\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"1\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"2011 International Conference on Asian Language Processing\",\"FirstCategoryId\":\"1085\",\"ListUrlMain\":\"https://doi.org/10.1109/IALP.2011.64\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"\",\"JCRName\":\"\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"2011 International Conference on Asian Language Processing","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/10.1109/IALP.2011.64","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}
The Comparison of Chinese Spam Filter Based on Generative Model and Discriminative Model
Previous studies have shown that discriminative model is better than generative model for spam filtering, which is tested on the English dataset. But the study on Chinese Spam Filter is rare. So we compared the performance of Bogo: a classical generative model, Logistic Regression (LR) and Relaxed Online SVM (ROSVM): two typical discriminative models on the Chinese dataset. Bogo system adopts a generative model, which is based on Bayesian algorithm. We choose the public Chinese datasets: TREC06c, SEWM 2008, SEWM 2010, SEWM 2011, as the test dataset with immediate feedback. The discriminative model gives the better results than the generative model based on spam filter. ROSVM gives the best performance on Chinese spam filter.