东篱科研大数据发现系统（DRDS）

位置：成果数据库 > 期刊 > 期刊详情页

基于统计和浅层语言分析的维吾尔文语义串快速抽取

ISSN号：1003-0077
期刊名称：《中文信息学报》
时间：0
分类：TP391[自动化与计算机技术—计算机应用技术;自动化与计算机技术—计算机科学与技术]
作者机构：新疆大学信息科学与工程学院,新疆乌鲁木齐830046
相关基金：国家自然科学基金（61562083,61262062,61262063）; 新疆维吾尔自治区高校科研计划重点项目（XJEDU2012I11）

关键词：维吾尔文, n元递增算法, 语义串抽取, 主题相似度, 文本分类, Uyghur language, frequent pattern-growth algorithm, semantic string extraction, topic similarity, text classification

中文摘要：

该文研究一种改进的n元递增算法来抽取维吾尔文本中表达关键信息的语义串,并用带权语义串集来刻画文本主题,提出了一种类似于Jaccard相似度的文本和类主题相似度度量方法,并实现了相应的维吾尔文分类算法。实验结果表明,该文提出的文本模型简单有效,分类算法计算量不高,而且还能达到或超过经典分类器的分类综合性能。

英文摘要：

This paper proposes an improved frequent pattern-growth approach to discover and extract the semantic strings which express key information in Uyghur texts.Then the topics are described by these weighted semantic strings.Based on these features,the Uyghur text classification is conducted by a new-designed Jaccard-like similarity measure.Experimental results show that the proposed method achieves comparable performance with a reasonable computation cost with regard to two traditional classifiers.

同期刊论文项目