首页 | 本学科首页   官方微博 | 高级检索  
     

汉维主题网页自动获取技术的研究
引用本文:梁建飞,吐尔根·依布拉音,田生伟,赛依旦·阿不力米提.汉维主题网页自动获取技术的研究[J].计算机应用与软件,2012(1):42-45.
作者姓名:梁建飞  吐尔根·依布拉音  田生伟  赛依旦·阿不力米提
作者单位:新疆大学信息科学与工程学院
基金项目:国家社科基金资助项目(10BTQ045);国家自然科学基金资助项目(60963017);新疆自冶区高校科研计划重点项目(XJEDU2009I05)
摘    要:为了获得大量用于机器翻译研究的汉维(维吾尔)文语料,提出一种从网页中自动获取主题信息的方法。考虑到有主题网页中主题信息分布相对集中、文本密度较高,并且这类网页中大量的噪音信息是由链接引入的,提出的算法首先将链接分为噪音链接和非噪音链接,并在源码中删除噪音链接的锚文本和非噪音链接的HTML标签,然后利用容器标签将源码划分为若干部分并删除文本长度和文本密度均小于各自阈值的源码块。针对汉维网页做了实验,实验结果表明,算法在设置合适的阈值的情况下良好率达到90%以上。

关 键 词:有主题网页  主题信息  噪音信息

RESEARCH ON CHINESE-UIGHUR THEME WEBPAGE AUTOMATIC ACQUISITION TECHNOLOGY
Affiliation:Liang Jianfei Turgun Ibrahim Tian Shengwei Sayida Ablimit(School of Information Science and Engineering Technology,Xinjiang University,Urumqi 830046,Xinjiang,China)
Abstract:In order to obtain a lot of Chinese-Uighur corpus for machine translation research,the paper proposes a method to automatically acquire topic information from webpages.Considering the comparatively more concentrative topic information,higher text density and a lot of noise information brought about by links in theme web pages,the algorithm proposed in the paper first of all classifies links into noise links and non-noise links,then edit the source to remove anchor texts from noise links and HTML tags from non-noise links;next according to container tags split the source code into parts,then delete those source blocks whose text lengths and densities are all less than their respective threshold values.Experiment results on Chinese-Uighur webpages illustrate that,when configured with appropriate threshold values the algorithm′s execution result reaches 90% fine rate.
Keywords:Theme web page Topic information Noise information
本文献已被 CNKI 等数据库收录!
设为首页 | 免责声明 | 关于勤云 | 加入收藏

Copyright©北京勤云科技发展有限公司  京ICP备09084417号