首页 | 本学科首页   官方微博 | 高级检索  
     

基于HTMLParser的Web信息抽取系统的设计与实现
引用本文:李彦刚,魏海平,侯兴华.基于HTMLParser的Web信息抽取系统的设计与实现[J].辽宁石油化工大学学报,2006,26(2):83-86.
作者姓名:李彦刚  魏海平  侯兴华
作者单位:辽宁石油化工大学计算机与通信工程学院,辽宁,抚顺,113001
摘    要:互联网上信息量的激增,迫切需要一些自动化的工具帮助人们在海量信息源中迅速找到真正需要的信息,如标题、链接e、mail和图片等,而HTML语言所表述的Web页面经浏览器分析后只适合浏览,不适合作为一种数据交换的方式由机器处理。介绍了HTMLParser的原理和java正则表达式相关知识,基于HTMLParser包和正则表达式。以提取网站内部email信息为例,提出了Web信息抽取系统设计方案,阐述了email信息抽取的工作原理和关键技术,给出了email抽取算法,并详细介绍了系统的抽取URL、email和存储模块,抽取结果保存于数据库中,供机器检索利用。

关 键 词:信息抽取  正则表达式  HTMLParser包  Java
文章编号:1672-6952(2006)02-0083-04
修稿时间:2005年12月2日

Design and Implementation of Web Information Extraction System Based on HTMLParser
LI Yan-gang,WEI Hai-ping,HOU Xing-hua.Design and Implementation of Web Information Extraction System Based on HTMLParser[J].Journal of Liaoning University of Petroleum & Chemical Technology,2006,26(2):83-86.
Authors:LI Yan-gang  WEI Hai-ping  HOU Xing-hua
Abstract:The rapid growth of the Web contents increases the need for some automatic tools to help to find the exact information among the magnanimous information sources such as titles,links,emails,pictures etc.The Web pages expressed by HTML,after analyzed by Internet Explorer,are suitable for browse,but not for machine processing as the way of data exchange.The principle of HTMLParser and related knowledge of regular expression,package HTMLParser and regular expression were introduced.Taking extracting email information inside websites as an example,the scheme of design was proposed.The principle of email extraction and key technique were presented.The algorithm of email extraction was given.URL extraction module,email extraction module and storage module were described in detail.The result of extraction is stored in database for the use of data retrieval.
Keywords:Java
本文献已被 CNKI 万方数据 等数据库收录!
设为首页 | 免责声明 | 关于勤云 | 加入收藏

Copyright©北京勤云科技发展有限公司  京ICP备09084417号