目前国内外在深层网络方面的研究几乎都围绕英文环境进行,还没有针对中文深层网络的研究.提出了对中文深层网络进行模式匹配和接口集成的方法.该方法首先创建一个用来存储同义词、超义词和子义词的字典,然后使用基于规则的分词算法将从接口中抽取的属性分成词.对于每一个属性,从定义的字典中找到其对应的所有同义词、超义词和子义词,生成一条相应的记录并存储到列表中,再从每条记录中选取出现次数最多的属性作为联合接口的属性.
Many researches about deep web focus on the deep web with English language, ignoring that with Chinese. In this paper, we present our work in schema matching and interface integration for Chinese deep web. We create a dictionary, which stores synonyms, hypernyms and hyponyms, at the very beginning. After interface extracting, we use Principle-based Segmentation algorithm to segment each attribute into words. Then, for each attribute, we look up the pre-created dictionary to find all its synonyms, hypernyms and hyponyms, form a record and store them in a list. Furthermore, we keep a counter for each attribute in the list to record times it appearing in the local interfaces. At last, we choose from each record a synonym with the largest count number as the attribute of union interface.