由于主题的缺失,传统的网页噪音去除算法均是通过一些启发式的规则判断哪些是有用信息,哪些是噪音信息。而在主题爬行的环境下,由于有了明确的主题,可以使用一些不同的方法来发现网页噪音。提出了一种基于主题的网页噪音去除算法,通过构造网页DOM树的一个变种,即内容块树,利用分类器判断网页的噪音块。实验结果表明,该方法噪音去除精度是87%,而以前的方法仅有42%。
In the absence of topic, traditional web page noise removal algorithm judges content block which one is noise and which one is not with some heuristic rules. But within the environment of focused crawling, clear topic presents, higher precision and better effect is achieved in a different way. A noise removal algorithm based on focused topic is proposed. After a variation of DOM (doCument object module) tree of web pages is constructed, i.e. content block tree, noise segment will be judged by a trained classifier. Experimental results demonstrate that the precision of our method is 87%, which is much better than previous method whose precision is 42%.