首页 | 本学科首页   官方微博 | 高级检索  
     


Automated classification of content components in technical communication
Authors:Jan Oevermann  Wolfgang Ziegler
Affiliation:1. Faculty 03: Mathematics/Computer Science, University of Bremen, Bremen, Germany;2. Faculty of Information Management and Media, Karlsruhe University of Applied Sciences, Karlsruhe, Germany
Abstract:Automated classification is usually not adjusted to specialized domains due to a lack of suitable data collections and insufficient characterization of the domain‐specific content and its effect on the classification process. This work describes an approach for the automated multiclass classification of content components used in technical communication based on a vector space model. We show that differences in the form and substance of content components require an adaption of document‐based classification methods and validate our assumptions with multiple real‐world data sets in 2 languages. As a result, we propose general adaptions on feature selection and token weighting, as well as new ideas for the measurement of classifier confidence and the semantic weighting of XML‐based training data. We introduce several potential applications of our method and provide prototypical implementation. Our contribution beyond the state of the art is a dedicated procedure model for the automated classification of content components in technical communication, which outperforms current document‐centered or domain‐agnostic approaches.
Keywords:content management  machine learning  technical communication  text classification
设为首页 | 免责声明 | 关于勤云 | 加入收藏

Copyright©北京勤云科技发展有限公司  京ICP备09084417号