东篱科研大数据发现系统（DRDS）

位置：成果数据库 > 期刊 > 期刊详情页

基于非主属性值的实体匹配

ISSN号：0254-4164
期刊名称：《计算机学报》
时间：0
分类：TP392[自动化与计算机技术—计算机应用技术;自动化与计算机技术—计算机科学与技术]
作者机构：[1]苏州大学计算机科学与技术学院,江苏苏州215006, [2]昆士兰大学信息技术与电子工程学院,布里斯班澳大利亚4067
相关基金：国家自然科学基金（61402313,61472263,61303019,61572336）、江苏省博士后科研基金（1501090B）、中国博士后第58批面上基金（2015M581859）和江苏软件新技术与产业化协同创新中心的资助.

作者：杨强[1], 李直旭[1], 蒋俊[1], 赵朋朋[1], 刘冠峰[1], 刘安[1], 周晓方[1,2]

关键词：实体匹配, 非主属性, 数据质量, 性能, 算法, record matching, non-key attribute, data quality, performance, algorithm

中文摘要：

实体匹配旨在找出不同数据源中指代同一实体的实例.已有的实体匹配方法大都基于实体主属性值的相似度进行匹配,而很少有工作考虑到使用实体的非主属性值来辅助实体匹配.然而,当两条指代同一实体的主属性值差异较大的时候,这两个实体可能不会被认为是匹配的实体.另一方面,这两个实体很可能共享一些特别的非主属性值,而这些非主属性值恰好可以反映出两个实体的匹配关系.基于这种思想,文中提出了一种新颖的基于非主属性值的实体匹配算法.该算法以类似于决策树的结构为基础,通过使用这种结构,不仅可以解决噪声值和空缺值带来的问题,而且可以极大地提高发现匹配记录以及尽可能早地排除不匹配记录的效率.多个数据集上的实验结果表明我们的方法比现有的实体匹配方法具有更高的准确率和召回率.此外,使用我们提出的基于决策树的匹配算法等有关技术较Baseline匹配算法在匹配效率上高出10倍多.

英文摘要：

Record Matching （RM） finds out instances referring to the same entity between different data sources. Existing work mainly uses the similarity between the key attribute values of instances for RM, while seldom work employs non-key attribute values. As a result, when two instances referring to the same entity do not have similar key attribute values, they might be missed as a matching pair. On the other hand, some particular non-key attribute values shared by the two instances might reflect the relationship between them. Based on the intuition, we propose a novel RM method based on non-key attribute values. Compared to key attribute, non-key attributes can be more noisy and inconsistent. Besides, there are usually a lot more non-key attributes than key attributes, thus RM based on non-key attributes faces a significant efficiency problem. To deal with these challenges, we propose a rule-based algorithm based on a tree-like structure. With this tree-like structure, we can not only deal with noisy and missing values, but also greatly improve the efficiency of the method by finding out matched instances or filtering unmatched instances as early as possible. The experimental results based on several data sets demonstrate that our method outperforms existing RM methods by reaching a higher precision and recall. Besides, the proposed techniques can greatly improve the efficiency of a baseline algorithm beyond 10 times.

同期刊论文项目

大规模社交网络中基于复杂社会行为的社会信任关系挖掘

期刊论文 4

基于互联网海量信息的数据库文本类型数据清洗研究

期刊论文 4

安全可信的社会化服务推荐研究

期刊论文 2

面向用户的数据质量管理方法研究

期刊论文 2

同项目期刊论文

一种高效的保护隐私的轨迹相似度计算框架

数据隐私保护的社会化推荐协议

Trip Oriented Search on Activity Trajectory

一种高效的保护隐私的轨迹相似度计算框架

数据隐私保护的社会化推荐协议

Trip Oriented Search on Activity Trajectory

数据隐私保护的社会化推荐协议

Trip Oriented Search on Activity Trajectory

期刊信息

《计算机学报》
北大核心期刊（2011版）

主管单位:中国科学院
主办单位:中国计算机学会中国科学院计算技术研究所
主编：孙凝晖
地址：北京中关村科学院南路6号
邮编：100190
邮箱：cjc@ict.ac.cn
电话：010-62620695

国际标准刊号：ISSN：0254-4164
国内统一刊号：ISSN：11-1826/TP
邮发代号:2-833

获奖情况:
中国期刊方阵“双效”期刊

国内外数据库收录:
美国数学评论（网络版）,荷兰文摘与引文数据库,美国工程索引,美国剑桥科学文摘,日本日本科学技术振兴机构数据库,中国中国科技核心期刊,中国北大核心期刊（2004版）,中国北大核心期刊（2008版）,中国北大核心期刊（2011版）,中国北大核心期刊（2014版）,中国北大核心期刊（2000版）

被引量:48433