随着基因测序技术和人类基因组计划的发展,从大量的生物数据中寻找相似的序列就越来越成为当前研究的热点问题.本文提出了一种聚类的多解析度字符串索引结构,用于解决生物序列的相似性查询问题.首先,以较小容量的MBR(最小绑定矩形)构造基因序列的多解析度字符串索引结构,然后通过对MBR的聚类以夏保序技术的应用,减小索引中MBR的平均体积,从而增加了查询向量到索引的空间距离,提高了索引的过滤能力.还给出了一种新的后处理方法,通过大量的减少编辑距离的计算,提高索引的性能.文中给出了该索引结构并详细介绍了索引的相关算法.实验表明,该索引结构是一种有效的处理生物数据的相似性查询的索引结构.
With the development of gene sequencing and human gene project, it became one of the research focuses to search similar sequences from vast biological data. This paper proposed a Clustered Multi Resolution String index structure to solve the problem of sequence similarity query. First, Multi Resolution String index structure is constructed with MBRs (Minimum Bounding Rectangles) of relatively small capacity. Then with the cluster of MBRs and the technique to keep the order of MBRs, the average volume of MBRs is reduced, so the distance between the query vector and the index is increased. Thus the filter ability of the index is greatly increased. This paper also proposed a new post process method to significantly increase index performance by greatly reducing the hum of ED computation. We describe the algorithm for the new index structure in detail and present experimental results, which show it is really an effective index structure for similarity query in the domain of biological data.