首页 | 本学科首页   官方微博 | 高级检索  
     

基于决策树和链接相似的Deep Web查询接口判定*
引用本文:李雪玲,施化吉,兰均,李星毅. 基于决策树和链接相似的Deep Web查询接口判定*[J]. 计算机应用研究, 2011, 28(11): 4086-4088. DOI: 10.3969/j.issn.1001-3695.2011.11.022
作者姓名:李雪玲  施化吉  兰均  李星毅
作者单位:江苏大学计算机科学与通信工程学院,江苏镇江,212013
基金项目:江苏省高校自然科学重大基金资助项目(08KJA520001);国家自然科学基金资助项目(70971067)
摘    要:针对现有Deep Web查询接口判定方法误判较多、无法有效区分搜索引擎类接口的不足,提出了基于决策树和链接相似的Deep Web查询接口判定方法。该方法利用信息增益率选取重要属性,并构建决策树对接口表单进行预判定,识别特征较为明显的接口;然后利用基于链接相似的判定方法对未识别出的接口进行二次判定,准确识别真正查询接口,排除搜索引擎类接口。结果表明,该方法能有效区分搜索引擎类接口,提高了分类的准确率和查全率。

关 键 词:Deep Web; 查询接口; 决策树; 链接相似

Deep Web query interface identification based on decision tree and link-similar
LI Xue-ling,SHI Hua-ji,LAN Jun,LI Xing-yi. Deep Web query interface identification based on decision tree and link-similar[J]. Application Research of Computers, 2011, 28(11): 4086-4088. DOI: 10.3969/j.issn.1001-3695.2011.11.022
Authors:LI Xue-ling  SHI Hua-ji  LAN Jun  LI Xing-yi
Affiliation:LI Xue-ling,SHI Hua-ji,LAN Jun,LI Xing-yi(School of Computer Science & Telecommunications Engineering,Jiangsu University,Zhenjiang Jiangsu 212013,China)
Abstract:In order to solve the problems existed in the traditional method that Deep Web query interfaces are more false positives and search engine class interface can not be effectively distinguished, this paper proposed a Deep Web query interface identification method based on decision tree and link-similar. This method used attribute information gain ratio as selection level, built a decision tree to pre-determine the form of the interfaces to identify the most interfaces which had some distinct features, and then used a new method based on link-similar to identify these unidentified again, distinguishing between Deep Web query interface and the interface of search engines. The result of experiment shows that it can enhance the accuracy and proves that it is better than the traditional methods.
Keywords:Deep Web   query interface   decision tree   link-similar
本文献已被 CNKI 万方数据 等数据库收录!
点击此处可从《计算机应用研究》浏览原始摘要信息
点击此处可从《计算机应用研究》下载全文
设为首页 | 免责声明 | 关于勤云 | 加入收藏

Copyright©北京勤云科技发展有限公司  京ICP备09084417号