东篱科研大数据发现系统（DRDS）

位置：成果数据库 > 期刊 > 期刊详情页

基于语义的互联网药品信息抽取算法

ISSN号：1003-3254
期刊名称：计算机系统应用
时间：0
页码：41-47
分类：TP311.1[自动化与计算机技术—计算机软件与理论;自动化与计算机技术—计算机科学与技术]
作者机构：[1]复旦大学软件学院,上海201203
相关基金：国家科技支撑项目（2006BAH02A05-06）;国家自然科学基金（60903078,60973025）
相关项目：结合描述逻辑和霍恩规则的不确定推理算法

关键词： WEB信息抽取, 语义词典, DOM, 信息熵, XPATH, 医药电子商务, Web information extraction, semantic dictionary, DOM, information entropy, XPath, medical E-business

中文摘要：

针对现有互联网信息抽取技术存在准确率不高、覆盖率低、人工干预多等诸多缺陷，提出了一种新的互联网药品信息抽取算法，通过引入语义技术构建三维语义词典，屏蔽不同药品信息网页在内容和结构上的异构性，同时利用所需抽取的目标药品属性信息具有一定聚集度的特征，基于信息熵的理论设计出对目标信息智能定位和抽取的方法。实验证明该算法既能降低人工干预，又具备较高的准确率和召回率。应用该算法能实时自动全面准确地获取互联网药品信息，为政府药监部门提供丰富的监管依据，对规范医药电子商务市场，保证人们的用药安全具有重要的现实意义。

英文摘要：

This article addresses defects of current Web information extraction technology such as low accuracy, low coverage, and manual intervention required, proposes a novel extraction algorithm of web medicine information. The algorithm sets up a three-dimentional semantic dictionary by introduction of the semantics technology, masks the isomerisms of the web page contents and structures, and at the same time, taking advantage of the fact that the attributes of the target medicine tend to have a character of aggregation, designs a way of intellectually locating and extracting the target information based on the theory of information entropy. Through related experiments proves that the algorithm is able to reduce the requirement of manual intervention of the information extraction, and has a high accuracy and recall rate. The application of this algorithm can automatically, comprehensively, and accurately obtain Internet medicine information in real time, offers abundant basis of supervision for the medicine supervision department, and therefore has a significant practical meaning of normalizing medical e-business and ensuring secure medication.

同期刊论文项目