首页 | 本学科首页   官方微博 | 高级检索  
     

中文文本体裁的自动分类机制
引用本文:方鸷飞,林鸿飞,杨志豪,赵晶.中文文本体裁的自动分类机制[J].中文信息学报,2006,20(2):26-34.
作者姓名:方鸷飞  林鸿飞  杨志豪  赵晶
作者单位:大连理工大学计算机系
基金项目:国家自然科学基金资助项目(60373095)
摘    要:文本按体裁自动分类属于按文本的形式分类的范畴,所以它与按内容自动分类问题有许多的不同之处,本文提出了一种关于中文文本体裁自动分类的新机制。在体裁分类过程中首要的问题是分类特征的选取,体裁分类特征项分为两种方式加以描述,一是集合形式,如基于分类词典和语料统计的政论性词汇和情感词汇等,二是规则形式,如公文标识信息和条文句等。基于根据特征之间的关联性和差异性,采用样本分布决策的方法抽取相应的特征项。最后利用支撑向量机算法进行自动分类。该机制已经在五类体裁的语料上得到实现,并获得了较好的效果。

关 键 词:计算机应用  中文信息处理  体裁分类  特征项选取  样本分布决策  支撑向量机  
文章编号:1003-0077(2006)02-0024-09
收稿时间:2005-03-09
修稿时间:2005-11-17

Automatic Classification of Chinese Text Genre
FANG Zhi-fei,LIN Hong-fei,YANG Zhi-hao,ZHAO Jing.Automatic Classification of Chinese Text Genre[J].Journal of Chinese Information Processing,2006,20(2):26-34.
Authors:FANG Zhi-fei  LIN Hong-fei  YANG Zhi-hao  ZHAO Jing
Affiliation:Department of Computer , Dalian University of Technology
Abstract:Genre is defined as a category on the basis of external criteria,so its classification is different from the classification based on content.A new mechanism for automatic classification of Chinese text genre is presented,and its main idea is as follows.Features for genre classification,as an essential factor in the mechanism,are described in two ways: one is in word-set,such as affective words and political words derived from some related dictionaries and corpus statistics;another one is in rule format,such as document identifiers and items.In terms of the correlativeness and variance of features,an approach of parametric distribution is applied to evaluate various features of the genres and extract the features for genre classification.Support Vector Machine is then used as the learning algorithm to build the classifier.The experiment on automatic classification of Chinese text genres,running on a text corpus consisting of five genres,shows that it can improve the precision of classification.
Keywords:computer application  Chinese information processing  text genre classification  feature selection  parametric distribution  support vector machine
本文献已被 CNKI 维普 万方数据 等数据库收录!
点击此处可从《中文信息学报》浏览原始摘要信息
点击此处可从《中文信息学报》下载全文
设为首页 | 免责声明 | 关于勤云 | 加入收藏

Copyright©北京勤云科技发展有限公司  京ICP备09084417号