首页 | 本学科首页   官方微博 | 高级检索  
     


THESUS: Organizing Web document collections based on link semantics
Authors:Maria?Halkidi  author-information"  >  author-information__contact u-icon-before"  >  mailto:mhalk@aueb.gr"   title="  mhalk@aueb.gr"   itemprop="  email"   data-track="  click"   data-track-action="  Email author"   data-track-label="  "  >Email author,Benjamin?Nguyen,Iraklis?Varlamis,Michalis?Vazirgiannis
Affiliation:(1) 76 Patision Street, Athens University of Economics and Business, Athens, Greece;(2) Domaine de Voluceau, INRIA, 78153 Le Chesnay, France
Abstract:The requirements for effective search and management of the WWW are stronger than ever. Currently Web documents are classified based on their content not taking into account the fact that these documents are connected to each other by links. We claim that a pagersquos classification is enriched by the detection of its incoming linksrsquo semantics. This would enable effective browsing and enhance the validity of search results in the WWW context. Another aspect that is underaddressed and strictly related to the tasks of browsing and searching is the similarity of documents at the semantic level. The above observations lead us to the adoption of a hierarchy of concepts (ontology) and a thesaurus to exploit links and provide a better characterization of Web documents. The enhancement of document characterization makes operations such as clustering and labeling very interesting. To this end, we devised a system called THESUS. The system deals with an initial sets of Web documents, extracts keywords from all pagesrsquo incoming links, and converts them to semantics by mapping them to a domainrsquos ontology. Then a clustering algorithm is applied to discover groups of Web documents. The effectiveness of the clustering process is based on the use of a novel similarity measure between documents characterized by sets of terms. Web documents are organized into thematic subsets based on their semantics. The subsets are then labeled, thereby enabling easier management (browsing, searching, querying) of the Web. In this article, we detail the process of this system and give an experimental analysis of its results.Received: 16 December 2002, Accepted: 16 April 2003, Published online: 17 September 2003
Keywords:World Wide Web  Link analysis  Similarity measure  Document clustering  Link management  Semantics
本文献已被 SpringerLink 等数据库收录!
设为首页 | 免责声明 | 关于勤云 | 加入收藏

Copyright©北京勤云科技发展有限公司  京ICP备09084417号