拼写错误的发现和候选词选取是文本分析中的一个重要的技术问题。本文结合维吾尔语的语音和词语结构特点,列出了文本中常见的拼写错误类型,详细分析了解决方法,利用最小编辑距离(minimum edit distance)算法实现了维吾尔语文本拼写错误分析中的查错和纠错功能,并以此为基础,结合维吾尔语构词规则,进一步提高了建议候选词的准确率和速度。该算法已被成功地应用到了维吾尔语文字自动校对和多文种文本检索等领域中。在以新疆高校学报为语料的测试中,词语查纠率达到85%以上。
Error detection and ranking are important issues in language analyzing. This paper summarizes the common spelling errors according to the phonetic and lexical features of Uighur and discusses the corresponding solution. It also presents and implements a minimum edit distance based approach for Uighur spelling check and correction, integrating the Uighur morphological structure to improve accuracy and speed of the correction ranking. The method is already applied in the areas such as automatic Uygur proofreading and multi-lingual text retrieval. Experiment on the texts from University Journals published in Xinjiang reaches the accuracy of 85 %.