文本聚类关键是有效解决特征词向量选择及特征词权重计算方法、文本相似度计算方法、聚类中心确定等三个问题。针对相关算法在三个关键环节上存在的问题,提出了适合自由文本特点的特征词权重计算方法和文本相似度计算方法;在此基础上提出了改进的CBC算法,从全局上自适应地确定文本集中的各个聚类中心。算法在实验中准确地确定了各个聚类中心,并在两个文本集上分别获得88.50%和94.00%的聚类准确率。
The three key points of text clustering are feature selection and weight calculation,texts similarity calculation and cluster center determination.This paper proposes two new methods based on the characteristic of free texts for feature-weight calculation and texts similarity calculation separately.Then an improved CBC algorithm is proposed to determine the cluster centers adaptively and globally.This algorithm produces all cluster center correctly,and obtains precision of 88.50% and 94.00% for two different text-set separately.