企业财务报告中存在大量蕴含着许多重要财务信息的非结构化文本信息.这类信息难以被计算机识别、分析和处理,也难以通过数据库技术进行管理.本文结合本体相关理论和自然语言处理(Natural Language Processing,NLP)技术,从词语属性描述、词语关系组织和相关知识链接3个维度构建财务报告领域本体,利用NLP工具对中文财务报告中的文本信息进行处理,将非结构化文本信息转化为结构化信息并使用XBRL表示,在一定程度上实现了文本信息的数据库存储与计算机分析处理.
Significant financial information can be retrieved from the vast amount of textual data provided in Chinese business accounting reports(annual reports).Nevertheless,due to the unstructured nature,this textual information usually is difficult to be obtained and analyzed via traditional computer and database techniques.To address this issue,a set of unified domain-specific ontology is presented,combined with Chinese Natural language processing(NLP),which transforms accounting reports in unstructured text into a structured XBRL-based form via three different dimensions,namely word attribute description,word relation organization,and related knowledge links respectively.