首页 | 本学科首页   官方微博 | 高级检索  
     

基于聚类和自动编码机的缺失数据填充算法
引用本文:卜范玉,陈志奎,张清辰. 基于聚类和自动编码机的缺失数据填充算法[J]. 计算机工程与应用, 2015, 51(18): 13-17
作者姓名:卜范玉  陈志奎  张清辰
作者单位:1.大连理工大学 软件学院,辽宁 大连 1166202.内蒙古财经大学 职业学院,呼和浩特 010010
摘    要:当前的不完整数据处理算法填充缺失值时,精度低下。针对这个问题,提出一种基于CFS聚类和改进的自动编码模型的不完整数据填充算法。利用CFS聚类算法对不完整数据集进行聚类,对降噪自动编码模型进行改进,根据聚类结果,利用改进的自动编码模型对缺失数据进行填充。为了使得CFS聚类算法能够对不完整数据集进行聚类,提出一种部分距离策略,用于度量不完整数据对象之间的距离。实验结果表明提出的算法能够有效填充缺失数据。

关 键 词:不完整数据  快速密度聚类算法(CFS)  自动编码机  部分距离策略  

Missing value imputation algorithm based on clustering and auto- encoder
BU Fanyu,CHEN Zhikui,ZHANG Qingchen. Missing value imputation algorithm based on clustering and auto- encoder[J]. Computer Engineering and Applications, 2015, 51(18): 13-17
Authors:BU Fanyu  CHEN Zhikui  ZHANG Qingchen
Affiliation:1.School of Software Technology, Dalian University of Technology, Dalian, Liaoning 116620, China2.College of Vocation, Inner Mongolia University of Finance and Economics, Huhhot 010010, China
Abstract:Existing algorithms are of low efficiency and effectiveness in imputing missing data. Aiming at this problem, the paper proposes a missing value imputation algorithm based on the CFS clustering and improved auto-encoder model. To cluster the incomplete data set, it improves the CFS clustering algorithm by introducing the partial distance strategy that is used to measure the distance between two objects with missing values. It uses the improved CFS algorithm to cluster the data set. The improved auto-encoder is used to estimate the missing values according to the clustering result. Experiments demonstrate that this proposed algorithm can impute the missing values effectively.
Keywords:incomplete data  Clustering by Fast Search and find of density peaks(CFS)  auto-encoder  partial distance strategy  
本文献已被 万方数据 等数据库收录!
点击此处可从《计算机工程与应用》浏览原始摘要信息
点击此处可从《计算机工程与应用》下载免费的PDF全文
设为首页 | 免责声明 | 关于勤云 | 加入收藏

Copyright©北京勤云科技发展有限公司  京ICP备09084417号