东篱科研大数据发现系统（DRDS）

位置：成果数据库 > 期刊 > 期刊详情页

两种基于双向比较的最长公共子串算法

ISSN号：1000-1239
期刊名称：《计算机研究与发展》
时间：0
分类：TP301.6[自动化与计算机技术—计算机系统结构;自动化与计算机技术—计算机科学与技术]
作者机构：[1]中国工程物理研究院计算机应用研究所,四川绵阳621900, [2]中国工程物理研究院电子工程研究所,四川绵阳621900
相关基金：国家“九七三”重点基础研究发展计划基金项目（2010CB328104）;中国工程物理研究院科学技术发展基金项目（200980403049）

作者：王开云[1], 孔思淇[1], 付云生[1], 潘泽友[1], 马卫东[2], 赵强[1]

关键词：最长公共子串, 双向比较, 连续同值区间, 跨越, 亚线性, longest common substring, bi-directional comparison, continuous same value segment, skip , sublinear

中文摘要：

查找两个给定字符串的最长公共子串（LCSstr）是一类重要字符串分析问题，在字符串近似匹配、计算机病毒特征码对比等方面有着广泛的用途．最长公共子串算法目前主要包括动态规划算法（LCSstrDP）和后缀数组算法（LCSstrSA），分别用于短串和长串的最长公共子串计算．前者代码简洁，但计算速度较慢，后者速度很快但算法非常复杂．提出两种基于双向比较的最长公共子串算法，即LCSstrSeL和LCSstrSCeL．LCSstrSeI。跨越已有的最长公共子串长度，与LCSstrDP相比，代码同样简洁，平均计算效率提高近一个数量级，并且不需要额外的存储空间．LCSstrSCeL是在LCSstrSeL的基础上，增加字符跨越、连续同值区间跨越等机制，平均效率较LCSstrSeI。亦有一定程度的提高，内存开销与LCSstrDP相近，在中小长度的字符串LCSstr计算中，平均计算效率高于LCSstrSA，某些情况下的计算效率可达到亚线性的速度．

英文摘要：

Finding the longest common substring （LCSstr） for two given strings is an important problem in string analysis. It can be used in many applications such as approximate string matching, biological sequences analysis, plagiarism detection and computer virus signature detection. There are two algorithms to solve the longest common substring problem： Dynamic programming （LCSstrDP） and the suffix array （LCSstrSA）. LCSstrDP solves the LCSstr problem by calculating the longest common suffix of the prefix （comparing from right to left）. Its code is simple, but of low efficientcy LCSstrSA calculates the longest common prefix of the suffix （comparing from left to right）. LCSstrSA~s time complexity is linear, though it is more complex. In this paper, we propose two LCSstr algorithms based on hi-directional comparison, named LCSstrSeL and LCSstrSCeL. LCSstrSeL skips the existing length of LCSstr with simple code and significantly improved efficiency, compared with LCSstrDP. On the basis of LCSstrSeL, LCSstrSCeL adds several mechanisms, such as character and continuous same value segment spanning. The test results show that not only the memory overhead of the algorithm is lower than that of LCSstrSA, but also the average efficiency is higher for small and medium strings. In some case the computation efficiency can be sublinear.

同期刊论文项目