有效HTML文本信息抽取方法的研究* Research on methods for extracting text information from HTML pages期刊界 All Journals 搜尽天下杂志传播学术成果专业期刊搜索期刊信息化学术搜索

有效HTML文本信息抽取方法的研究*

引用本文：	韩忠明,李文正,莫倩.有效HTML文本信息抽取方法的研究*[J].计算机应用研究,2008,25(12):3568-3571.

作者姓名：	韩忠明李文正莫倩

作者单位：	北京工商大学,计算机学院,北京,100037

基金项目：	北京市教委科技计划面上资助项目（KM200810011008）

摘要：	从新闻网页和博客网页中抽取出正文内容是一个非常有意义的研究问题,但是多数网页中含有大量与正文无关的噪声内容，导致很难从网页中获取正确的文本信息。分析了中文新闻与博客网页的正文特征，用实验表明了利用HTML与文本的密度比可以进行文本的识别与抽取。提出了机器学习、统计估计以及FDR三种HLML正文抽取方法，并作了大量的实验比较和分析。实验结果表明，该算法可以有效地过滤噪声而且算法的复杂度很低，效率与效果均达到一个很好的平衡。
关键词：	网页信息抽取机器学习统计
Research on methods for extracting text information from HTML pages

HAN Zhong ming,LI Wen zheng,MO Qian.Research on methods for extracting text information from HTML pages[J].Application Research of Computers,2008,25(12):3568-3571.

Authors:	HAN Zhong ming LI Wen zheng MO Qian

Affiliation:	(College of Computer Science & Technology, Beijing Technology & Business University, Beijing 100037, China)

Abstract:	Extracting text information from news and blog HTML pages is a very important and interesting research problem. There are too many noises to extract precise texts. The paper analyzed statistic characterizes of HTML pages and showed it was possible to extract texts using information about the density of text vs. HTML code. Proposed three methods based on machine leaning, statistic and false discovery rate(FDR). Conducted comprehensive experiments and the result show these methods can effectively extract texts and balance effectivity and efficiency.

Keywords:	HTML pages information extracting machine learning statistic
本文献已被 CNKI 维普万方数据等数据库收录！
	点击此处可从《计算机应用研究》浏览原始摘要信息
	点击此处可从《计算机应用研究》下载全文

设为首页 | 免责声明 | 关于勤云 | 加入收藏