东篱科研大数据发现系统（DRDS）

位置：成果数据库 > 期刊 > 期刊详情页

Web页面细粒度数据抽取方法研究

ISSN号：1000-7024
期刊名称：《计算机工程与设计》
时间：0
分类：TP391.3[自动化与计算机技术—计算机应用技术;自动化与计算机技术—计算机科学与技术]
作者机构：[1]首都师范大学信息工程学院,北京100048, [2]北京理工大学图书馆,北京100081, [3]西南大学计算机与信息科学学院,重庆400715
相关基金：国家自然科学基金项目（61272446）;北京市属高等学校人才强教深化计划基金项目（PHR201008083）

作者：王旭仁[1], 杨硕[1], 何发镁[2], 王彦丽[1], 张为群[3]

关键词：信息抽取, WEB挖掘, 包装器, 自动关联, information extraction, Web data mining, wrapper, automatically correlating

中文摘要：

尽管有很多方法用于Web页面信息抽取，对细粒度信息如数据项等的抽取需求仍然很迫切。提出了一个用于结构化数据抽取的解决方案，将Web页面上的信息以更细的粒度抽取出来。对包装器（wrapper）生成时所依据的信息进行了基于稳定性的分类，实现了模板和种子之间多对多的自动关联（automatically correlating），并按照信息稳定性的高低为每个字段生成多个抽取规则，在抽取信息时根据多个抽取规则进行抽取，只有在所有规则失效时才会导致抽取失败，提高了抽取系统的鲁棒性。实验结果表明，该方法具有良好的抽取功率和准确率。

英文摘要：

Although there are many approaches for data extraction from web pages, demand for finer-grained information, such as item information, is still urging especially in oriented domains applications. A solution is proposed for structured data extrac- tion from web pages. System characteristics are in the following aspects., generating the wrapper on the basis of information based on stability classification. The templates and the seeds of the many-to-many relationships in automatic way are realized. According to the information stability level for each field, multiple extraction rules are generated. Only when all rules fail, it is regarded as extraction failure. All above features improve extraction system robustness. Experimental results show that the method has good extraction successful rate and accurate rate.

同期刊论文项目