针对新闻领域的专题组织进行了研究,提出了一种基于时序窗口的动态热点话题提取模型。该模型整合了热点话题的两个特点。一方面关注主题词在新闻文本中的广泛性,衡量标准为多频道播报特征项的频率综合,词频越高其广泛性越高;另一方面考虑新闻流主题词的突发性,表现为特定时间段内主题词出现频率显著异常于其它时间段。引入时序窗口进行上升和下降突发模式提取,并结合TF-DF作为主题词赋权值依据。实验结果表明,这种基于时序窗口的动态热点话题提取模型对新闻文本进行主题抽取具有很好的性能。
This paper gives a description of a study of topic organization in the news domain, and presents a novel dynamic hot topic extraction model based on the time window. The model combines two characteristics of hot topics together. One is the pervasiveness of topic terms in news texts, which is evaluated by the occurrences of the topic terms reported by different channels, and the more frequent the occurrence of the topic terms reported, the higher the pervasiveness of topic terms. The other one is the burst of topic terms in the news stream, which can be assessed by the abnormal occurrence frequencies of topic terms in a specific interval compared with other different time intervals. The time window is introduced to make burst detection and the term frequency-proportional document freqency (TF-PDF) is combined to weigh the terms. The experimental results demonstrate that this model is effective in topic extraction for news texts.