首页 | 本学科首页   官方微博 | 高级检索  
     

类别严重不均衡应用的在线数据流学习算法
引用本文:赵强利,蒋艳凰.类别严重不均衡应用的在线数据流学习算法[J].计算机科学,2017,44(6):255-259.
作者姓名:赵强利  蒋艳凰
作者单位:湖南商学院计算机与信息工程学院 长沙410205;国防科技大学高性能计算国家重点实验室 长沙410073,国防科技大学高性能计算国家重点实验室 长沙410073
基金项目:本文受国家自然科学基金(61272141,5,61472136),国防科技大学高性能计算国家重点实验室基金(201513-02)资助
摘    要:集成式数据流挖掘是对存在概念漂移的数据流进行学习的重要方法。对于类别分布严重不均衡的应用,集成式数据流挖掘中数据块的学习方式导致样本数多的类别的分类精度高,样本数少的类别的分类精度低的问题,现有算法无法满足此类应用的需求。针对上述问题,对基于回忆机制的集成式数据流学习算法MAE(Memorizing based Adaptive Ensemble)进行改进,提出面向类别严重不均衡应用的在线数据流学习算法UMAE(Unbalanced data Lear-ning based on MAE)。UMAE算法为每个类别设置了一个样本滑动窗口,对于新到达的数据块,其样本依据自身的类别分别进入相应的滑动窗口,最后利用各类别滑动窗口内的样本构建用于在线学习的数据块。与5种典型的数据流挖掘算法的比较结果表明,UMAE算法在满足实时性的同时,不仅整体分类精度高,而且对于样本数很少的小类别的分类精度有大幅度提高;对于异常检测等类别分布严重不均衡的应用,UMAE算法的实用性明显优于其他算法。

关 键 词:在线学习  数据流挖掘  回忆与遗忘机制  不均衡数据学习
收稿时间:2016/5/18 0:00:00
修稿时间:2016/8/28 0:00:00

Online Data Stream Mining for Seriously Unbalanced Applications
ZHAO Qiang-li and JIANG Yan-huang.Online Data Stream Mining for Seriously Unbalanced Applications[J].Computer Science,2017,44(6):255-259.
Authors:ZHAO Qiang-li and JIANG Yan-huang
Affiliation:School of Computer and Information Engineering,Hunan University of Commerce,Changsha 410205,China;State Key Laboratory of High Performance Computing,National University of Defense Technology,Changsha 410073,China and State Key Laboratory of High Performance Computing,National University of Defense Technology,Changsha 410073,China
Abstract:Using ensemble of classifiers on sequential blocks of training instances is a popular strategy for data stream mining with concept drifts.Yet for the seriously unbalanced applications where the number of examples for each class in the data blocks is totally different,traditional data block creation will result in low accuracy for the small classes with much less number of instances.This paper provided an updating algorithm UMAE (Unbalanced data learning based on MAE) for seriously unbalanced applications based on MAE (Memorizing based Adaptive Ensemble).UMAE sets an equal-sized sliding window for each class.When each data block comes,each example in the data block comes into the corresponding sliding window based on its classes.During the learning process,a new data block will be created by using the instances in the current sliding windows.This new data block is adopted to generate a new classifier.Compared with five traditional data stream mining approaches,the results show that UMAE achieves high accuracy for seriously unba-lanced applications,especially for the small classes with much less number of instances in the applications.
Keywords:Online learning  Data stream mining  Recalling and forgetting mechanisms  Unbalanced data learning
点击此处可从《计算机科学》下载全文
设为首页 | 免责声明 | 关于勤云 | 加入收藏

Copyright©北京勤云科技发展有限公司  京ICP备09084417号