尽管有很多方法用于Web页面信息抽取,对细粒度信息如数据项等的抽取需求仍然很迫切。提出了一个用于结构化数据抽取的解决方案,将Web页面上的信息以更细的粒度抽取出来。对包装器(wrapper)生成时所依据的信息进行了基于稳定性的分类,实现了模板和种子之间多对多的自动关联(automatically correlating),并按照信息稳定性的高低为每个字段生成多个抽取规则,在抽取信息时根据多个抽取规则进行抽取,只有在所有规则失效时才会导致抽取失败,提高了抽取系统的鲁棒性。实验结果表明,该方法具有良好的抽取功率和准确率。
Although there are many approaches for data extraction from web pages, demand for finer-grained information, such as item information, is still urging especially in oriented domains applications. A solution is proposed for structured data extrac- tion from web pages. System characteristics are in the following aspects., generating the wrapper on the basis of information based on stability classification. The templates and the seeds of the many-to-many relationships in automatic way are realized. According to the information stability level for each field, multiple extraction rules are generated. Only when all rules fail, it is regarded as extraction failure. All above features improve extraction system robustness. Experimental results show that the method has good extraction successful rate and accurate rate.