首页 | 本学科首页   官方微博 | 高级检索  
     

文本分类中基于概率主题模型的噪声处理方法
引用本文:林洋港,陈恩红.文本分类中基于概率主题模型的噪声处理方法[J].计算机工程与科学,2010,32(7):89-92.
作者姓名:林洋港  陈恩红
作者单位:中国科学技术大学计算机科学与技术学院,安徽,合肥,230027
基金项目:国家自然科学基金资助项目,国家863计划资助项目 
摘    要:训练集中文本质量的好坏直接决定着文本分类的结果。实际应用中训练集的构建不可避免地会产生噪声样本,从而影响文本分类方法的实际应用效果。为此,针对文本分类中的噪声问题,本文提出一种基于概率主题模型的噪声处理方法,首先对训练集中的每个样本计算其类别熵,根据类别熵对噪声样本进行过滤;然后利用主题模型进行数据平滑,进一步减弱噪声样本的影响。这种方法不但能够减弱噪声样本对分类结果的影响,同时还保持了训练集的原有规模。在真实数据上的实验表明,该方法对噪声样本的分布具有较好的鲁棒性,在噪声比例较大的情况下仍能保持较好的分类结果。

关 键 词:噪声数据  文本分类  概率主题模型  类别熵
收稿时间:2009-01-13
修稿时间:2009-05-18

A Probabilistic Topic Model Based Noise Processing Method for Text Classification
LIN Yang-gang,CHEN En-hong.A Probabilistic Topic Model Based Noise Processing Method for Text Classification[J].Computer Engineering & Science,2010,32(7):89-92.
Authors:LIN Yang-gang  CHEN En-hong
Affiliation:(School of Computer Science and Technology,University of Science and Technology of China,Hefei 230027,China)
Abstract:The performance of text classification depends directly on the quality of training corpus.In practical applications,noise samples are unavoidable in the training corpus and thus influence the effect of the text classification approach.To this end,a novel probabilistic topic model based noise processing method is proposed for text classification.In our method,the  noise samples are filtered according to the class entropy.Then the data is smoothed using the generative process of the topic model to further weaken the influence of noise samples,meanwhile the original size of the training corpus is kept.The experimental results of the real world data show that the method proposed is robust to the distribution of noise samples,and has a relative good performance on the data sets with a high noise ratio.
Keywords:noisy data  text classification  probabilistic topic model  class entropy
本文献已被 万方数据 等数据库收录!
点击此处可从《计算机工程与科学》浏览原始摘要信息
点击此处可从《计算机工程与科学》下载全文
设为首页 | 免责声明 | 关于勤云 | 加入收藏

Copyright©北京勤云科技发展有限公司  京ICP备09084417号