通过观察网站呈现网页的规律及网页本身的结构特点,提出基于URL类型及网页链接变化规律的入口页面识别算法,优先抓取入口页面.在实际应用中,取得了较好的更新效果.
The refreshment algorithm based on URL type and outlink change is proposed by observing the page orderliness of Web sites and the structural characteristics of the page. This algorithm is used for fetching the entry pages,and a perfect effect in real application is obtained.