首页 | 本学科首页   官方微博 | 高级检索  
     

基于函数依赖与条件约束的数据修复方法
引用本文:金澈清,刘辉平,周傲英.基于函数依赖与条件约束的数据修复方法[J].软件学报,2016,27(7):1671-1684.
作者姓名:金澈清  刘辉平  周傲英
作者单位:华东师范大学 计算机科学与软件工程学院 数据科学与工程研究院, 上海 200062,华东师范大学 计算机科学与软件工程学院 数据科学与工程研究院, 上海 200062,华东师范大学 计算机科学与软件工程学院 数据科学与工程研究院, 上海 200062
基金项目:国家重点基础研究发展规划(973)( 2012CB316203);国家自然科学基金(61370101); 上海市教委科研创新重点项目(14ZZ045).
摘    要:随着经济与信息技术的发展,在许多应用中均产生大量数据.然而,受硬件设备、人工操作、多源数据集成等诸多因素的影响,在这些应用之中往往存在较为严重的数据质量问题,特别是不一致性问题,从而无法有效管理数据.因此,首要的任务就是开发新型数据清洗技术来提升数据质量,以支持后续的数据管理与分析.现有工作主要研究基于函数依赖的数据修复技术,即以函数依赖来描述数据一致性约束,通过变更数据库中部分元组的属性值(而非增加/删除元组)来使得整个数据库遵循函数依赖集合.从一致性约束描述的角度来看,函数依赖并非是唯一的表达方式,还存在其他表达方式,例如硬约束、数量约束、等值约束、非等值约束等.然而,随着一致性约束种类的增加,其处理难度也远比仅有函数依赖的场景要困难.本文考虑以函数依赖与其他一致性约束共同表述数据库的一致性约束,并在此基础上设计数据修复算法,从而提升数据质量.实验结果表明,本文所提方法的执行效率较高.

关 键 词:数据质量  数据修复  函数依赖  条件约束  等价类
收稿时间:2015/10/9 0:00:00
修稿时间:2016/1/12 0:00:00

Functional Dependency and Conditional Constraint Based Data Repair
JIN Che-Qing,LIU Hui-Ping and ZHOU Ao-Ying.Functional Dependency and Conditional Constraint Based Data Repair[J].Journal of Software,2016,27(7):1671-1684.
Authors:JIN Che-Qing  LIU Hui-Ping and ZHOU Ao-Ying
Affiliation:Institute for Data Science and Engineering, School of Computer Science and Software Engineering, East China Normal University, Shanghai 200062, China,Institute for Data Science and Engineering, School of Computer Science and Software Engineering, East China Normal University, Shanghai 200062, China and Institute for Data Science and Engineering, School of Computer Science and Software Engineering, East China Normal University, Shanghai 200062, China
Abstract:Along with the development of economy and information technology, a large amount of data is produced in many applications. However, due to the influence of some factors, such as hardware equipments, manual operations, and multi-source data integration, serious data quality issues arise, including data inconsistency, which makes it more challenging to manage data effectively. Hence, it is urgent to develop new data cleaning technology to improve data quality to support further data management and analysis. Currently, the existing work mainly focuses on the situation where functional dependencies are used to describe data inconsistency. Once some violations are detected, some tuples must be changed to suit for the functional dependency set via update, neither insert nor delete. Besides functional dependency, there also exist other kinds of constraints, such as the hard constraint, quantity constraint, equivalent constraint, non-equivalent constraint, etc. However, it becomes more difficult when more kinds of inconsistent conditions are involved. In this paper, we consider the general scenario where functional dependencies and other constraints co-exist. We design the corresponding data repair algorithm to improve the data quality effectively. Experimental results show that the proposed method behaves effectively and efficiently.
Keywords:data quality  data repair  functional dependency  conditional constraint  equivalence class
点击此处可从《软件学报》浏览原始摘要信息
点击此处可从《软件学报》下载全文
设为首页 | 免责声明 | 关于勤云 | 加入收藏

Copyright©北京勤云科技发展有限公司  京ICP备09084417号