首页 | 本学科首页   官方微博 | 高级检索  
     

基于混合跳链条件随机场的异构Web记录集成方法
引用本文:黄健斌,姬红兵,孙鹤立.基于混合跳链条件随机场的异构Web记录集成方法[J].软件学报,2008,19(8):2149-2158.
作者姓名:黄健斌  姬红兵  孙鹤立
作者单位:西安电子科技大学,电子工程学院,陕西,西安,710071
基金项目:Supported by the National Natural Science Foundation of China under Grant No.60202004 (国家自然科学基金); the Doctoral Innovation Foundation of Xidian University of China under Grant No.05013 (西安电子科技大学博士创新基金)
摘    要:提出了一种混合跳链条件随机场序列统计学习模型,以实现异构Web记录与关系数据库的模式匹配.该模型可以在由手工标注样本和关系数据库记录组成的联合样本集上进行训练,减少了对繁琐手工标注样本的依赖.此外,通过在线性链条件随机场模型上增加对跳边的支持,使得该模型能够有效地处理状态变量间的长距离依赖.在多个领域的真实数据集上的实验结果表明,所提出的方法能够显著提高异构Web记录语义模式匹配的性能.

关 键 词:混合跳链条件随机场  Web数据集成  模式匹配
收稿时间:2006/10/14 0:00:00
修稿时间:3/8/2007 12:00:00 AM

Integration of Heterogeneous Web Records Using Mixed Skip-Chain Conditional Random Fields
HUANG Jian-Bin,JI Hong-Bing and SUN He-Li.Integration of Heterogeneous Web Records Using Mixed Skip-Chain Conditional Random Fields[J].Journal of Software,2008,19(8):2149-2158.
Authors:HUANG Jian-Bin  JI Hong-Bing and SUN He-Li
Abstract:An improved sequence labeling model named Mixed Skip-Chain Conditional Random Field is presented to solve the problem of schema matching between semi-structured Web records and relational database. The proposed model can be trained on mixed samples set which consists of labeled samples and unlabeled relational database records to reduce the dependence on manually labeled training data.Moreover,it provides a novel way to incorporate the long-distance dependencies between different state variants.Experimental results using a large number of real-world data collected from diverse domains show that the proposed method can improve the performance of schema matching significantly.
Keywords:mixed skip-chain conditional random fields  Web data integration  schema matching
本文献已被 CNKI 维普 万方数据 等数据库收录!
点击此处可从《软件学报》浏览原始摘要信息
点击此处可从《软件学报》下载全文
设为首页 | 免责声明 | 关于勤云 | 加入收藏

Copyright©北京勤云科技发展有限公司  京ICP备09084417号