基于后缀数组的分词技术① Word Segment Based on Suffix Array期刊界 All Journals 搜尽天下杂志传播学术成果专业期刊搜索期刊信息化学术搜索

基于后缀数组的分词技术①

引用本文：	任雪利,代余彪.基于后缀数组的分词技术①[J].计算机系统应用,2010,19(8):229-230.

作者姓名：	任雪利代余彪

作者单位：	曲靖师范学院计算机科学与工程学院,云南,曲靖,655011

基金项目：	曲靖师范学院基金(2008QN007);云南省教育厅研究课题(09C0188)

摘要：	中文分词技术是机器翻译、分类、搜索引擎以及信息检索的基础,但是,互联网上不断出现的新词严重影响了分词的性能,为了提高新词的识别率,建立待分词内容的后缀数组,然后计算其公共前缀共同出现的次数,采用阈值对其进行过滤筛选出候选词语,实验结果表明,该方法在新词识别方面有一定的优势。
关键词：	后缀数组分词公共前缀长度
收稿时间：	2009/12/4 0:00:00
修稿时间：	2010/1/18 0:00:00
Word Segment Based on Suffix Array

REN Xue-Li and DAI Yu-Biao.Word Segment Based on Suffix Array[J].Computer Systems& Applications,2010,19(8):229-230.

Authors:	REN Xue-Li and DAI Yu-Biao

Affiliation:	(Department of Computer Science and Engineering,Qujing Normal University,Qujing 655011,China)

Abstract:	Chinese word segmentation technology is the basis of machine translation, classification, search engines, as well as information retrieval. But the Internet emerging new words have seriously affected the performance of word segmentation. To improve the recognition rate of new words, suffix array is used in this paper, and the number of length of common prefix is calculated. The candidates on their words are filtered out by the threshold. Experimental results show that the new word recognition method has advantages.

Keywords:	suffix array word segment LCP
本文献已被维普万方数据等数据库收录！
	点击此处可从《计算机系统应用》浏览原始摘要信息
	点击此处可从《计算机系统应用》下载全文

设为首页 | 免责声明 | 关于勤云 | 加入收藏