基于差异-相似矩阵的文本降维方法 Dimensionality reduction for text document using difference-similitude matrix期刊界 All Journals 搜尽天下杂志传播学术成果专业期刊搜索期刊信息化学术搜索

基于差异-相似矩阵的文本降维方法

引用本文：	黄晓春,晏蒲柳,夏德麟,陈健.基于差异-相似矩阵的文本降维方法[J].计算机应用,2005,25(8):1821-1823.

作者姓名：	黄晓春晏蒲柳夏德麟陈健

作者单位：	武汉大学电子信息学院

基金项目：	国家自然科学基金资助项目(90204008)

摘要：	由于文本文档数量多、词量大,形成的文档空间维度高,很多自动文本分类算法不能直接有效地发挥作用。基于差异-相似矩阵(DSM)的方法在很大程度上降低了文档空间的维度。已经分好类的文集经过预处理后被表示成特征项-文档矩阵,再转化为差异-相似矩阵,其中同类文档采用相似项描述,而异类文档则采用差异项描述。通过对差异-相似矩阵的处理,最终得到维度较低的文本特征集,并同时生成分类规则。实验说明,对于大规模文集,DSM方法能在保持良好的分类质量的同时,获得较高的属性降维率和样本降维率。
关键词：	文本分类维度消减差异&mdash 相似矩阵
文章编号：	1001-9081(2005)08-1821-03
Dimensionality reduction for text document using difference-similitude matrix

HUANG Xiao-chun,Yan Pu-liu,XIA De-lin,CHEN Jian.Dimensionality reduction for text document using difference-similitude matrix[J].journal of Computer Applications,2005,25(8):1821-1823.

Authors:	HUANG Xiao-chun Yan Pu-liu XIA De-lin CHEN Jian

Affiliation:	School of Electronic Information, Wuhan University, Wuhan Hubei 430079,China

Abstract:	Due to the huge amount of text documents and their vocabulary, document spaces are commonly of high dimensionality, and many automatical text categorization algorithms can not get their best performences directly. Difference-similitude Matrix-based (DSM) method reduces dimensionality to a great extend. Pre-classified collection is represented as a item-document matrix after preprocessing, then transmitted into a DSM, in which documents in the same classes are depicted with similitude while documents in different classes with difference. The method generates an item set of low dimensionality and a set of classification rules after dealing with the DSM. Results of experiments suggest that DSM-based method could achieve high attribute reduction degree and sample reduction degree with good classification quality.

Keywords:	text categorization dimensionality reduction DSM(Difference-Similitude Matrix)
本文献已被 CNKI 维普万方数据等数据库收录！
	点击此处可从《计算机应用》浏览原始摘要信息
	点击此处可从《计算机应用》下载全文

设为首页 | 免责声明 | 关于勤云 | 加入收藏