首页 | 本学科首页   官方微博 | 高级检索  
     

考虑样本不平衡的模型无关的基因选择方法
引用本文:李建中,杨昆,高宏,骆吉洲,郭政.考虑样本不平衡的模型无关的基因选择方法[J].软件学报,2006,17(7):1485-1493.
作者姓名:李建中  杨昆  高宏  骆吉洲  郭政
作者单位:1. 哈尔滨工业大学,计算机科学与工程系,黑龙江,哈尔滨,150001
2. 哈尔滨医科大学
基金项目:国家高技术研究发展计划(863计划)
摘    要:在基因表达数据分析中,鉴别基因是后续研究中非常重要的信息基因.有很多研究致力于从基因表达数据中选出信息基因这一挑战性工作,并提出了一些基因选择方法.然而,这些方法(特别是非参数选择方法)都没有考虑不同样本类别中样本大小的不平衡性问题.考虑样本不平衡性和基因选择方法的稳定性,给出一个全新的与数据分布模型无关的基因选择方法.在类内变化小和类间差别大的策略下,选择敏感的度量函数提高方法的鉴别能力,同时,利用类内变化和类间差别的一致性来增加方法的稳定性和适用性.这一方法不但可以应用于两个类别的情况,也可以应用于多个类别的情况.最后,使用两组真实的基因表达数据对所提出的方法进行了验证.实验结果表明,这一方法比其他方法具有更高的有效性和稳健性.

关 键 词:基因选择  基因表达  分类  微阵列
收稿时间:2005-04-19
修稿时间:2005-12-13

Model-Free Gene Selection Method by Considering Unbalanced Samples
LI Jian-Zhong,YANG Kun,GAO Hong,LUO Ji-Zhou and GUO Zheng.Model-Free Gene Selection Method by Considering Unbalanced Samples[J].Journal of Software,2006,17(7):1485-1493.
Authors:LI Jian-Zhong  YANG Kun  GAO Hong  LUO Ji-Zhou and GUO Zheng
Affiliation:Department of Computer Science and Engineering, Harbin Institute of Technology, Harbin 150001, China
Abstract:In gene expression data analysis, discriminator genes are importantly informative genes for further research. Recently, a great deal of research has focused on the challenging task of identifying these informative genes from microarray data. However, the sizes of sample classes in microarray data are often unbalanced. The unbalance of samples has not been explicitly and correctly considered by the existing gene selection methods, especially nonparametric methods. Considering the unbalance of samples and the stability of the approach for identifying informative genes, a novel and model-free gene selection method is proposed in this paper. With considering within-class difference and between-class variation, as well as the homogeneities of the within-class difference and between-class variations, scoring functions of genes are constructed to select discriminator genes. This method is not only applicable in two-category case but also applicable in multi-category case. The experimental results on two publicly available microarray datasets, leukemia data and small round blue cell tumor data, show that the proposed method is very efficient and robust to select discriminator genes.
Keywords:gene selection  gene expression  classification  microarray
本文献已被 CNKI 维普 万方数据 等数据库收录!
点击此处可从《软件学报》浏览原始摘要信息
点击此处可从《软件学报》下载全文
设为首页 | 免责声明 | 关于勤云 | 加入收藏

Copyright©北京勤云科技发展有限公司  京ICP备09084417号