通用搜索引擎数据量庞大,但查询结果不够准确。分类目录正好相反。为了综合两者优势,对垂直搜索引擎进行了研究和分析。着重研究了垂直搜索引擎的核心模块——智能网络搜索蜘蛛。提出了搜索分析的新概念——规则。研究了蜘蛛中定义支持同义词的语义词典的方法,给出了按照规则分析和检索的实现方法和流程。程序需要定义多种规则,让蜘蛛依照规则进行网页爬行和信息采集。最后给出一个项目实例,证明了上述方法的可行性。
General search engine has large volume of data, but its search results are not accurate enough. Directories classification is on the contrary. In order to integrate advantages of the two, vertical search engine is studied and analyzed. The core module--intelligent search spider is mainly focused on. A new concept about searching and analyzing is brought forward: Rules. The method is researched that defining semantic dictionary which supports synonyms. The algorithm and flow that realize searching and analyzing according rules are afforded. Kinds of rules must be defined in search spider program, depending on which the function web pages crawling and information data extracting work. At last a project example is presented to prove the feasibility of these methods.