首页 | 本学科首页   官方微博 | 高级检索  
     


SEL: A unified algorithm for salient entity linking
Authors:Salvatore Trani  Claudio Lucchese  Raffaele Perego  David E. Losada  Diego Ceccarelli  Salvatore Orlando
Affiliation:1. National Research Council of Italy (CNR), Institute of Information Science and Technologies (ISTI), Pisa, Italy;2. Centro Singular de Investigación en Tecnoloxías da Información (CiTIUS), Universidade de Santiago de Compostela, Santiago de Compostela, Spain;3. Bloomberg LP, London, UK;4. Dipartimento di Scienze Ambientali, Informatica e Statistica (DAIS), Università Ca' Foscari Venezia, Venice, Italy
Abstract:The entity linking task consists in automatically identifying and linking the entities mentioned in a text to their uniform resource identifiers in a given knowledge base. This task is very challenging due to its natural language ambiguity. However, not all the entities mentioned in the document have the same utility in understanding the topics being discussed. Thus, the related problem of identifying the most relevant entities present in the document, also known as salient entities (SE), is attracting increasing interest. In this paper, we propose salient entity linking, a novel supervised 2‐step algorithm comprehensively addressing both entity linking and saliency detection. The first step is aimed at identifying a set of candidate entities that are likely to be mentioned in the document. The second step, besides detecting linked entities, also scores them according to their saliency. Experiments conducted on 2 different data sets show that the proposed algorithm outperforms state‐of‐the‐art competitors and is able to detect SE with high accuracy. Furthermore, we used salient entity linking for extractive text summarization. We found that entity saliency can be incorporated into text summarizers to extract salient sentences from text. The resulting summarizers outperform well‐known summarization systems, proving the importance of using the SE information.
Keywords:entity linking  machine learning  salient entities  text summarization
设为首页 | 免责声明 | 关于勤云 | 加入收藏

Copyright©北京勤云科技发展有限公司  京ICP备09084417号