东篱科研大数据发现系统（DRDS）

位置：成果数据库 > 期刊 > 期刊详情页

基于HTML树和模板的文献信息提取方法研究

ISSN号：1001-3695
期刊名称：《计算机应用研究》
时间：0
分类：TP311.13[自动化与计算机技术—计算机软件与理论;自动化与计算机技术—计算机科学与技术]
作者机构：[1]大连理工大学管理学院系统工程研究所,辽宁大连116024
相关基金：国家自然科学基金资助项目（70572099）; 辽宁省自然科学基金资助项目（1050349）

关键词：网页信息提取, 文档对象模型树, 模板, 文献信息搜集, Web information extraction, DOM tree, template, document information extraction

中文摘要：

教师科研文献信息的自动搜集是科研成果有效管理的重要手段,将网页信息的提取方法用于网络数据库中文献信息的自动搜集有广大的应用前景。提出基于DOM树和模板的文献信息提取方法,利用HTML标记间的嵌套关系将Web网页表示成一棵DOM树,将DOM树结构用于网页相似度的度量和自动分类,相似度高的网页应用同一模板进行信息提取。实验结果表明该方法在提取网络数据库中文献信息的准确率在94%以上。

英文摘要：

The automatic collection of the teacher research paper information is an important means of effective management of scientific research,there is a broad application prospects to apply the method of Web page information extraction to the paper information collection. This paper proposed a method of paper information collection based on the HTML tree and template. This method would represent the Web page into a DOM tree using the hierarchy relationship of the HTML tags,then the DOM tree would be used to the measure of the page similarity and the classification of Web pages. The information of Web pages with high similarity would be extracted using the same template. The experiment result shows that the accuracy of this method is above 94% in collecting the paper information from the Web database.

同期刊论文项目