首页 | 本学科首页   官方微博 | 高级检索  
     检索      


A generalized cluster centroid based classifier for text categorization
Authors:Guansong Pang  Shengyi Jiang
Institution:1. School of Management, Guangdong University of Foreign Studies, Guangzhou 510006, PR China;2. School of Informatics, Guangdong University of Foreign Studies, Guangzhou 510420, PR China
Abstract:In this paper, a Generalized Cluster Centroid based Classifier (GCCC) and its variants for text categorization are proposed by utilizing a clustering algorithm to integrate two well-known classifiers, i.e., the K-nearest-neighbor (KNN) classifier and the Rocchio classifier. KNN, a lazy learning method, suffers from inefficiency in online categorization while achieving remarkable effectiveness. Rocchio, which has efficient categorization performance, fails to obtain an expressive categorization model due to its inherent linear separability assumption. Our proposed method mainly focuses on two points: one point is that we use a clustering algorithm to strengthen the expressiveness of the Rocchio model; another one is that we employ the improved Rocchio model to speed up the categorization process of KNN. Extensive experiments conducted on both English and Chinese corpora show that GCCC and its variants have better categorization ability than some state-of-the-art classifiers, i.e., Rocchio, KNN and Support Vector Machine (SVM).
Keywords:Text categorization  KNN  Rocchio  Clustering  Generalized cluster centroid
本文献已被 ScienceDirect 等数据库收录!
设为首页 | 免责声明 | 关于勤云 | 加入收藏

Copyright©北京勤云科技发展有限公司  京ICP备09084417号