东篱科研大数据发现系统（DRDS）

位置：成果数据库 > 期刊 > 期刊详情页

基于扩展特征向量空间模型的多源数据融合

ISSN号：1671-9352
期刊名称：山东大学学报(理学版)
时间：2013.10.21
页码：87-92
分类：TP391[自动化与计算机技术—计算机应用技术;自动化与计算机技术—计算机科学与技术]
作者机构：[1]河南财经政法大学计算机与信息工程学院,河南郑州450002
相关基金：国家自然科学基金资助项目（61202285）;国家级星火计划项目（2012GA750007）;河南省教育厅科学技术研究重点项目（12A120002）;河南省科技厅基础与前沿技术研究项目（122300410378）;河南财经政法大学学术创新骨干支持计划资助
相关项目：基于邻近局部切空间相似性的多流形学习研究

作者：陈珂锐|潘君|

关键词：自然语言处理, 本体, 多源数据融合, 语义判歧, natural language processing, ontology, multi-source data fusion, semantic role matching

中文摘要：

本体资源的扩充是自然语言处理的关键问题之一。传统的从单一数据源获取的信息其覆盖率较低，亟需建立一个整体的数据管理平台，对数据资源分类存储与整理。为此提出了AVP数据平台，构建AVP平台所面临的重要问题是多源数据的融合，即将不同来源的网站数据进行语义角色标注，对歧义词条进行识别判断，并最终归并到以义项为基本单位的数据仓库中；为解决多源数据融合的语义角色标注问题，给出了一种自动语义判歧方法。其基本思想是利用词条中的属性值对作为特征模板，并借助于属性值的共现概率，应用扩展向量空间模型对词条进行歧义识别。通过大量的实验对比可知，该系统在各方面均取得优异的成绩，所提出的算法能够很好地解决多源数据融合中的语义判歧问题。

英文摘要：

The expansion of ontology resource is one of the key for the whole natural language processing. Since the in- formation obtained traditionally from single data source could not reflect the overall picture and the coverage rate doesn＇ t reach targeted one, the construction of an integrated data management platform would be required to store and organize data sources by classification. The AVP data platform was proposed firstly. In the process of data construction on AVP platform, the most important issue is to integrate multi-source data, in other words, to perform semantic role labeling on web data coming from different sources, to identify ambiguous entries, and to eventually merge into data warehouses which use sense as the basic unit. An automated method of semantic role matching has been suggested, and it would solve the problem of semantic role matching resulted from multi-source data fusion. The basic idea is to use at- tribute-values of entries as the feature template, and then apply expand vector space model to identity ambiguity for en- tries while assisted by the co-occurrence probability of attribute values. Through the massive experimental contrast, the system mentioned above performed very well in all respects. The theory and algorithm proposed in this paper could solve the problem of semantic role matching existed in multi-source data fusion effectively.

同期刊论文项目