树和模板的文献信息提取方法研究* Method of paper information extraction based on HTML tree and template期刊界 All Journals 搜尽天下杂志传播学术成果专业期刊搜索期刊信息化学术搜索

树和模板的文献信息提取方法研究*

引用本文：	李文立,王乐超,宋春雷. 树和模板的文献信息提取方法研究*[J]. 计算机应用研究, 2010, 27(12): 4615-4617. DOI: 10.3969/j.issn.1001-3695.2010.12.064

作者姓名：	李文立王乐超宋春雷

作者单位：	大连理工大学,管理学院,系统工程研究所,辽宁,大连,116024

基金项目：	国家自然科学基金资助项目(70572099)；辽宁省自然科学基金资助项目(1050349)

摘要：	教师科研文献信息的自动搜集是科研成果有效管理的重要手段，将网页信息的提取方法用于网络数据库中文献信息的自动搜集有广大的应用前景。提出基于DOM树和模板的文献信息提取方法，利用HTML标记间的嵌套关系将Web网页表示成一棵DOM树，将DOM树结构用于网页相似度的度量和自动分类，相似度高的网页应用同一模板进行信息提取。实验结果表明该方法在提取网络数据库中文献信息的准确率在94%以上。
关键词：	网页信息提取；文档对象模型树；模板；文献信息搜集
Method of paper information extraction based on HTML tree and template

LI Wen-li,WANG Le-chao,SONG Chun-lei. Method of paper information extraction based on HTML tree and template[J]. Application Research of Computers, 2010, 27(12): 4615-4617. DOI: 10.3969/j.issn.1001-3695.2010.12.064

Authors:	LI Wen-li WANG Le-chao SONG Chun-lei

Abstract:	The automatic collection of the teacher research paper information is an important means of effective management of scientific research, there is a broad application prospects to apply the method of Web page information extraction to the paper information collection. This paper proposed a method of paper information collection based on the HTML tree and template. This method would represent the Web page into a DOM tree using the hierarchy relationship of the HTML tags, then the DOM tree would be used to the measure of the page similarity and the classification of Web pages. The information of Web pages with high similarity would be extracted using the same template. The experiment result shows that the accuracy of this method is above 94% in collecting the paper information from the Web database.

Keywords:	Web information extraction DOM tree template document information extraction
本文献已被万方数据等数据库收录！
	点击此处可从《计算机应用研究》浏览原始摘要信息
	点击此处可从《计算机应用研究》下载全文

设为首页 | 免责声明 | 关于勤云 | 加入收藏