东篱科研大数据发现系统（DRDS）

位置：成果数据库 > 期刊 > 期刊详情页

基于最小编辑距离的维语词语检错与纠错研究

ISSN号：1003-0077
期刊名称：《中文信息学报》
时间：0
分类：TP391[自动化与计算机技术—计算机应用技术;自动化与计算机技术—计算机科学与技术]
作者机构：[1]新疆大学信息科学与工程学院,新疆乌鲁木齐830046
相关基金：国家自然科学基金资助项目（60662002）;新疆维吾尔自治区高校科研计划资助项目（XJEDU2005S02）

作者：玛依热·依布拉音[1], 米吉提·阿不里米提[1], 艾斯卡尔·艾木都拉[1]

关键词：计算机应用, 中文信息处理, 维语尔语, 词法分析, 纠错, 最小编辑距离, computer application, Chinese information processing, Uighur, morphological analysis, spelling check, minimum edit distance

中文摘要：

拼写错误的发现和候选词选取是文本分析中的一个重要的技术问题。本文结合维吾尔语的语音和词语结构特点，列出了文本中常见的拼写错误类型，详细分析了解决方法，利用最小编辑距离（minimum edit distance）算法实现了维吾尔语文本拼写错误分析中的查错和纠错功能，并以此为基础，结合维吾尔语构词规则，进一步提高了建议候选词的准确率和速度。该算法已被成功地应用到了维吾尔语文字自动校对和多文种文本检索等领域中。在以新疆高校学报为语料的测试中，词语查纠率达到85％以上。

英文摘要：

Error detection and ranking are important issues in language analyzing. This paper summarizes the common spelling errors according to the phonetic and lexical features of Uighur and discusses the corresponding solution. It also presents and implements a minimum edit distance based approach for Uighur spelling check and correction, integrating the Uighur morphological structure to improve accuracy and speed of the correction ranking. The method is already applied in the areas such as automatic Uygur proofreading and multi-lingual text retrieval. Experiment on the texts from University Journals published in Xinjiang reaches the accuracy of 85 %.

同期刊论文项目