主题Deep Web爬虫框架研究 Research for framework of subject deep web crawler期刊界 All Journals 搜尽天下杂志传播学术成果专业期刊搜索期刊信息化学术搜索

主题Deep Web爬虫框架研究

引用本文：	黄聪会,张水平,胡洋.主题Deep Web爬虫框架研究[J].计算机工程与设计,2010,31(5).

作者姓名：	黄聪会张水平胡洋

作者单位：	空军工程大学电讯工程学院,陕西,西安,710077

基金项目：	陕西省自然科学基金项目

摘要：	为满足用户精确化和个性化获取信息的需要,通过分析Deep Web信息的特点,提出了一个可搜索不同主题Deep Web 信息的爬虫框架.针对爬虫框架中Deep Web数据库发现和Deep Web爬虫爬行策略两个难题,分别提出了使用通用搜索引擎以加快发现不同主题的Deep Web数据库和采用常用字最大限度下载Deep Web信息的技术.实验结果表明了该框架采用的技术是可行的.
关键词：	深网爬虫搜索引擎信息抽取常用字
Research for framework of subject deep web crawler

HUANG Cong-hui,ZHANG Shui-ping,HU Yang.Research for framework of subject deep web crawler[J].Computer Engineering and Design,2010,31(5).

Authors:	HUANG Cong-hui ZHANG Shui-ping HU Yang

Affiliation:	HUANG Cong-hui,ZHANG Shui-ping,HU Yang (Institute of Telecommunication Engineering,Air Force Engineering University,Xi\'an 710077,China)

Abstract:	To satisfy people's demand for getting precise and personal information,characteristics of deep web information are analyzed,and a framework of crawler for searching different subject information in deep web is put forward. To solve the difficult problems of deep web database discovery and deep web crawler crawling strategy,the technologies of discovering different subject deep web database quickly to use the universal search engine and downloading deep web information to the utmost by adopting the commonly...

Keywords:	deep web crawler search engine information extraction commonly used Chinese characters
本文献已被 CNKI 万方数据等数据库收录！

设为首页 | 免责声明 | 关于勤云 | 加入收藏