东篱科研大数据发现系统（DRDS）

位置：成果数据库 > 期刊 > 期刊详情页

基于语义构词的汉语词语语义相似度计算

ISSN号：1003-0077
期刊名称：《中文信息学报》
时间：0
分类：TP391[自动化与计算机技术—计算机应用技术;自动化与计算机技术—计算机科学与技术]
作者机构：[1]北京大学中国语言文学系,北京100871, [2]北京大学计算语言学研究所,北京100871, [3]北京大学计算语言教育部重点实验室,北京100871
相关基金：国家社科基金（16BYY137）;国家重点基础研究发展计划资助项目（2014CB340504）;国家社科基金（12＆ZD119）

关键词：词语语义相似度计算, 语义构词, 词义知识表示, 语素概念, Chinese word similarity computing, Chinese semantic word-formation, lexical knowledge representation, morphemic concepts

中文摘要：

汉语词语语义相似度计算,在中文信息处理的多种应用中扮演至关重要的角色。基于汉语字本位的思想,我们采用词类、构词结构、语素义等汉语语义构词知识,以“语素概念”为基础,计算汉语词语语义相似度。这种词义知识表示简单、直观、易于拓展,计算模型简洁、易懂,采用了尽可能少的特征和参数。实验表明,该文方法在典型“取样词对”上的表现突出,其数值更符合人类的感性认知,且在全局数据上也表现出了合理的分布规律。

英文摘要：

Chinese word similarity computing plays an important role in the Chinese information processing. Based on the notion of character-orientation, Chinese semantic word-formation knowledge, including word POS, word-formation pattern and morphemic concepts, is employed to compute Chinese word similarity. This lexical knowledge rep resentation is simple, intuitive and easy to expand and the model is straight-forward, with characteristics and param eters adopted as less as possible. Experimental results show that the approach is promising for the typical sampling word pair. Also, the numerical values of similarity are more in line with human cognition and present a reasonable distribution of the global data.

同期刊论文项目