首页 | 本学科首页   官方微博 | 高级检索  
     

基于半监督学习的恶意URL检测方法
引用本文:麻瓯勃,刘雪娇,唐旭栋,周宇轩,胡亦承.基于半监督学习的恶意URL检测方法[J].计算机系统应用,2020,29(11):11-20.
作者姓名:麻瓯勃  刘雪娇  唐旭栋  周宇轩  胡亦承
作者单位:杭州师范大学 杭州国际服务工程学院,杭州 311121
基金项目:浙江省自然科学基金(LY19F020021); 浙江省大学生科技创新活动计划(新苗人才计划) (2019R426035)
摘    要:检测恶意URL对防御网络攻击有着重要意义. 针对有监督学习需要大量有标签样本这一问题, 本文采用半监督学习方式训练恶意URL检测模型, 减少了为数据打标签带来的成本开销. 在传统半监督学习协同训练(co-training)的基础上进行了算法改进, 利用专家知识与Doc2Vec两种方法预处理的数据训练两个分类器, 筛选两个分类器预测结果相同且置信度高的数据打上伪标签(pseudo-labeled)后用于分类器继续学习. 实验结果表明, 本文方法只用0.67%的有标签数据即可训练出检测精确度(precision)分别达到99.42%和95.23%的两个不同类型分类器, 与有监督学习性能相近, 比自训练与协同训练表现更优异.

关 键 词:恶意URL检测  半监督学习  协同训练改进算法  Doc2Vec  分类器训练
收稿时间:2019/11/18 0:00:00
修稿时间:2019/12/11 0:00:00

Malicious URL Detection Based on Semi-Supervised Learning
MA Ou-Bo,LIU Xue-Jiao,TANG Xu-Dong,ZHOU Yu-Xuan,HU Yi-Cheng.Malicious URL Detection Based on Semi-Supervised Learning[J].Computer Systems& Applications,2020,29(11):11-20.
Authors:MA Ou-Bo  LIU Xue-Jiao  TANG Xu-Dong  ZHOU Yu-Xuan  HU Yi-Cheng
Affiliation:Hangzhou Institute of Service Engineering, Hangzhou Normal University, Hangzhou 311121, China
Abstract:Detecting malicious URL is important for defending against cyber attacks. In view of the problem that supervised learning requires a large number of labeled samples, this study uses a semi-supervised learning method to train malicious URL detection models, which reduces the cost overhead of labeling data. We propose an improved algorithm based on the traditional co-training. Two kinds of classifiers are trained by using expert knowledge and Doc2Vec pre-processed data, and the data with the same prediction result and the high confidence of the two classifiers are screened and used for classifiers learning after being pseudo-labeled. The experimental results show that the proposed method can train two different types of classifiers with detection precision of 99.42% and 95.23% with only 0.67% of labeled data, which is similar to supervised learning performance and performs better than self-training and co-training.
Keywords:malicious URL detection  semi-supervised learning  co-training improvement algorithm  Doc2Vec  classifier training
本文献已被 万方数据 等数据库收录!
点击此处可从《计算机系统应用》浏览原始摘要信息
点击此处可从《计算机系统应用》下载全文
设为首页 | 免责声明 | 关于勤云 | 加入收藏

Copyright©北京勤云科技发展有限公司  京ICP备09084417号