位置:成果数据库 > 期刊 > 期刊详情页
基于版块的论坛增量搜集策略
  • 期刊名称:中文信息学报
  • 时间:0
  • 页码:62-68
  • 分类:TP391[自动化与计算机技术—计算机应用技术;自动化与计算机技术—计算机科学与技术]
  • 作者机构:[1]山东大学计算机科学与技术学院,山东济南250101
  • 相关基金:国家自然科学基金资助项目(60970047);山东省科技攻关资助项目(2007GG10001002,2008GG10001026);山东省自然科学基金资助项目(Y2008G19)
  • 相关项目:Web图像的语义表示及在聚类与排序中的应用
作者: 杜言琦|马军|
中文摘要:

该文研究论坛的增量搜集问题。由于在论坛中同一主题通常分布在多个页面上,而传统增量搜集技术的抓取策略通常是基于单个页面,因此这些技术并不适于对论坛增量搜集。该文通过对许多论坛中版块变化规律的统计分析,提出了基于版块的论坛增量搜集策略。该策略将属于同一版块的所有页面看做一个整体,以它做为抓取的基本单位。同时该策略利用版块权重和局部时间规律确定抓取频率和抓取时间点。实验结果表明本策略对新增和新回复帖子的平均召回率为99.3%,并且与平均调度方法相比系统总延迟最高可减小42%。

英文摘要:

This paper studies the problem of incremental crawling of forums. Since a topic in a forum is usually distributed in more than one page and the revisiting strategy of traditional incremental technologies is centered on the individual page, these technologies are not suitable for crawling forum sites incrementally. Based on the statistical analysis on the evolution of board in many Web forums, a novel and board-based incremental crawling strategy is proposed. The main idea of the approach is to define the pages of the same board as the basic unit for re-crawling. In detail, this approach leverages the board weights and local time discipline to allocate crawl resources and determine the crawl time. Experimental results show that the recall for the newly published and updated discussion threads is close to 99.3% for our method strategy, and the overall system delay is maximally decreased by 42% as compared with even seheduling method.

同期刊论文项目
同项目期刊论文