首页 | 本学科首页   官方微博 | 高级检索  
     

基于Hadoop平台的XML文档重复数据检测
引用本文:李振兴,刘波.基于Hadoop平台的XML文档重复数据检测[J].计算机系统应用,2013,22(11):195-199.
作者姓名:李振兴  刘波
作者单位:暨南大学 信息科学技术学院学院, 广州 510632;暨南大学 信息科学技术学院学院, 广州 510632
摘    要:XML数据越来越广泛地被用于信息交换与集成中,其数据质量问题引起了人们的关注.解决由数据质量引发的问题,实体识别技术非常关键.为了克服现有方法的不足,在海量XML数据上进行高效的重复对象检测,以实体识别技术为基础提出了基于Hadoop平台的XML文档重复检测算法,它将所有标签节点统称为属性,用实体来描述属性,通过属性的比较,快速地找到在某些属性上相同的所有实体对象,并利用Hadoop应用框架处理海量数据的优势实现并行处理.经过试验验证该方法良好的扩展性,伸缩性和高效性.

关 键 词:XML  数据质量  重复检测  Hadoop  分布式
收稿时间:2013/4/22 0:00:00
修稿时间:2013/5/28 0:00:00

XML Data Duplicate Detection Based on Hadoop Platform
LI Zhen-Xing and LIU Bo.XML Data Duplicate Detection Based on Hadoop Platform[J].Computer Systems& Applications,2013,22(11):195-199.
Authors:LI Zhen-Xing and LIU Bo
Affiliation:College of Information Science and Technology, Jinan University, Guangzhou 510632, China;College of Information Science and Technology, Jinan University, Guangzhou 510632, China
Abstract:As being more and more widely used for data exchange and integration, the XML data quality issues cause more concern. In order to overcome the problems caused by data quality, Entity Resolution(ER) is critical. To overcome the drawbacks of current methods's deficiency and perform entity resolution efficiently and effectively on massive XML data set, under the basis of Entity Resolution, an XML data duplicate detection based on hadoop platform algorithm is presented in this paper. The method uses entities to describe their atrributes. By the comparing of the attributes,we can find all the objects that have the same attributes quickly. Meanwhile, taking the advantage of the Hadoop platform which can process massive data parallel. From the experiments, the method has excellent performance in scalability, flexibility and efficiency.
Keywords:XML  data quality  duplicate detection  Hadoop  distribute
本文献已被 维普 等数据库收录!
点击此处可从《计算机系统应用》浏览原始摘要信息
点击此处可从《计算机系统应用》下载全文
设为首页 | 免责声明 | 关于勤云 | 加入收藏

Copyright©北京勤云科技发展有限公司  京ICP备09084417号