首页 | 本学科首页   官方微博 | 高级检索  
     

基于差异度的不均衡电信客户数据分类方法
引用本文:王林,郭娜娜.基于差异度的不均衡电信客户数据分类方法[J].计算机应用,2017,37(4):1032-1037.
作者姓名:王林  郭娜娜
作者单位:西安理工大学 自动化与信息工程学院, 西安 710048
基金项目:国家自然科学基金资助项目(61405157)。
摘    要:针对传统分类技术对不均衡电信客户数据集中流失客户识别能力不足的问题,提出一种基于差异度的改进型不均衡数据分类(IDBC)算法。该算法在基于差异度分类(DBC)算法的基础上改进了原型选择策略。在原型选择阶段,利用改进型的样本子集优化方法从整体数据集中选择最具参考价值的原型集,从而避免了随机选择所带来的不确定性;在分类阶段,分别利用训练集和原型集、测试集和原型集样本之间的差异性构建相应的特征空间,进而采用传统的分类预测算法对映射到相应特征空间内的差异度数据集进行学习。最后选用了UCI数据库中的电信客户数据集和另外6个普通的不均衡数据集对该算法进行验证,相对于传统基于特征的不均衡数据分类算法,DBC算法对稀有类的识别率平均提高了8.3%,IDBC算法对稀有类的识别率平均提高了11.3%。实验结果表明,所提IDBC算法不受类别分布的影响,而且对不均衡数据集中稀有类的识别能力优于已有的先进分类技术。

关 键 词:客户流失预测  不均衡数据分类  样本子集优化  原型选择  差异度转化  
收稿时间:2016-09-05
修稿时间:2016-12-26

Imbalanced telecom customer data classification method based on dissimilarity
WANG Lin,GUO Nana.Imbalanced telecom customer data classification method based on dissimilarity[J].journal of Computer Applications,2017,37(4):1032-1037.
Authors:WANG Lin  GUO Nana
Affiliation:College of Automation and Information Engineering, Xi'an University of Technology, Xi'an Shaanxi 710048, China
Abstract:It is difficult for conventional classification technology to discriminate churn customers in the context of imbalanced telecom customer dataset, therefore, an Improved Dissimilarity-Based imbalanced data Classification (IDBC) algorithm was proposed by introducing an improved prototype selection strategy to Dissimilarity-Based Classification (DBC) algorithm. In prototype selection stage, the improved sample subset optimization method was adopted to select the most valuable prototype set from the whole dataset, thus avoiding the uncertainties caused by the random selection; in classification stage, new feature space was constructed via dissimilarity between samples from train set and prototype set, and samples from test set and prototype set, and then dissimilarity-based datasets mapped into corresponding feature space were learnt with conventional classification algorithms. Finally, the telecom customer dataset and other six ordinary imbalanced datasets from UCI database were selected to test the performance of IDBC. Compared with the traditional imbalanced data classification algorithm based on features, the recognition rate of DBC algorithm for rare class was improved by 8.3% on average, and the recognition rate of IDBC algorithm for raw class was increased by 11.3%. The experimental results show that the IDBC algorithm is not affected by the category distribution, and the discriminative ability of IDBC algorithm outperforms existing state-of-the-art approaches.
Keywords:customer churn prediction  imbalanced data classification  Sample Subset Optimization (SSO)  prototype selection  dissimilarity transformation  
点击此处可从《计算机应用》浏览原始摘要信息
点击此处可从《计算机应用》下载全文
设为首页 | 免责声明 | 关于勤云 | 加入收藏

Copyright©北京勤云科技发展有限公司  京ICP备09084417号