东篱科研大数据发现系统（DRDS）

位置：成果数据库 > 期刊 > 期刊详情页

基于DIV标签分段的藏文网页正文提取研究

ISSN号：1003-0077
期刊名称：《中文信息学报》
时间：0
分类：TP391.1[自动化与计算机技术—计算机应用技术;自动化与计算机技术—计算机科学与技术]
作者机构：西藏大学藏文信息技术研究中心,西藏拉萨850000
相关基金：2015年度西藏自治区自然科学基金项目“藏文搜索引擎关键技术研究”（项目号：2015ZR-14-9）;2015年度西藏自治区自然科学基金项目“基于逐字匹配的藏文分词技术与未登录词研究”（项目号：2015ZR-14-10）; 2013年度国家自然科学基金重大项目“跨语言社会舆情分析基础理论与关键技术研究”（项目号：61331013）阶段性成果

关键词：藏文网页, 分段, 正文信息, DIV元素, 标签, Tibetan webpage, paragraphing, main body text information, DIV element, tagging

中文摘要：

文章针对藏文电子文献资源匮乏、文本资源不规整、收集困难等问题,提出了基于DIV标签分段的藏文网页正文提取算法,该算法将原始网页信息分割为页面信息中与DIV元素等量的信息段,再对段中标签等非正文信息进行删除,最终形成该页正文。实验表明,正文提取结果准确、通用性强,适用于互联网上不同模型的藏文网页。

英文摘要：

The arrangement and collection of Tibetan electronic resources is the one of the important proceduresof constructing Tibetan information processing resource. Extracting algorithm of main body text from Tibetanwebpage was proposed based on DIV tagging paragraph in connection with scarce of Tibetan electronic literatureresources, unstructured text resources and its difficulties in collecting, and other issues. This algorithm can sepa-rate original webpage information as message segments equal to DIV elements information section, and then de-lete the non-main-body information such as the labeled paragraph using some strategy, finally forming the mainbody text of the page. Experimental result demonstrated that arithmetic method can archive accurate result oftext extraction with strong usability and it can be applied in various Tibetan webpage models on internet.

同期刊论文项目