针对深层网查询结果页面中噪音信息对数据区域识别的干扰问题,提出一种自动识别深层网查询结果数据区域的方法。该方法利用网页的重复结构和相似URL,将页面划分成不同的语义块,依据不同页面块之间URL的相似性识别出数据区域。实验结果表明,该方法能够提高数据区域识别的召回率和准确率。
Aiming at the problem that the noise information may interfere with the identification of the data region in Deep Web search result pages.This paper proposes an automatic approach to identify data region in Deep Web search result list pages.It employs continuous repetitive structure and similar URL to divide the sample pages into different semantic blocks,and identifies the block where the data region locates.Experimental results show the approzch can imprave the recall rate and accuracy of the date region identification.