首页 | 本学科首页   官方微博 | 高级检索  
     


Use of graph theory measures to identify errors in record linkage
Authors:Sean M. Randall  James H. Boyd  Anna M. Ferrante  Jacqueline K. Bauer  James B. Semmens
Affiliation:Centre for Data Linkage, Curtin University, Kent Street, Bentley, WA 6102, Australia
Abstract:Ensuring high linkage quality is important in many record linkage applications. Current methods for ensuring quality are manual and resource intensive. This paper seeks to determine the effectiveness of graph theory techniques in identifying record linkage errors. A range of graph theory techniques was applied to two linked datasets, with known truth sets. The ability of graph theory techniques to identify groups containing errors was compared to a widely used threshold setting technique. This methodology shows promise; however, further investigations into graph theory techniques are required. The development of more efficient and effective methods of improving linkage quality will result in higher quality datasets that can be delivered to researchers in shorter timeframes.
Keywords:Record linkage   Graph theory   Data quality
本文献已被 ScienceDirect 等数据库收录!
设为首页 | 免责声明 | 关于勤云 | 加入收藏

Copyright©北京勤云科技发展有限公司  京ICP备09084417号