首页 | 本学科首页   官方微博 | 高级检索  
     


Ensuring annotation consistency and accuracy for Vietnamese treebank
Authors:Quy T Nguyen  Yusuke Miyao  Ha T T Le  Nhung T H Nguyen
Affiliation:1.SOKENDAI (The Graduate University for Advanced Studies),Kanagawa,Japan;2.National Institute of Informatics,Tokyo,Japan;3.University of Social Sciences and Humanities,Ho Chi Minh City,Vietnam;4.University of Science,Ho Chi Minh City,Vietnam
Abstract:Treebanks are important resources for researchers in natural language processing. They provide training and testing materials so that different algorithms can be compared. However, it is not a trivial task to construct high-quality treebanks. We have not yet had a proper treebank for such a low-resource language as Vietnamese, which has probably lowered the performance of Vietnamese language processing. We have been building a consistent and accurate Vietnamese treebank to alleviate such situations. Our treebank is annotated with three layers: word segmentation, part-of-speech tagging, and bracketing. We developed detailed annotation guidelines for each layer by presenting Vietnamese linguistic issues as well as methods of addressing them. Here, we also describe approaches to controlling annotation quality while ensuring a reasonable annotation speed. We specifically designed an appropriate annotation process and an effective process to train annotators. In addition, we implemented several support tools to improve annotation speed and to control the consistency of the treebank. The results from experiments revealed that both inter-annotator agreement and accuracy were higher than 90%, which indicated that the treebank is reliable.
Keywords:
本文献已被 SpringerLink 等数据库收录!
设为首页 | 免责声明 | 关于勤云 | 加入收藏

Copyright©北京勤云科技发展有限公司  京ICP备09084417号