27 papers · ranked by Valyu relevance
Andre Lamurias, Pedro Ruas, Francisco M. Couto
Background Biomedical literature concerns a wide range of concepts, requiring controlled vocabularies to maintain a consistent terminology across different research groups. However, as new concepts are introduced, biomedical literature is prone to ambiguity, specifically in fields that are advancing more rapidly, for…
Dahlia Shehata
Despite the advantages of their low-resource settings, traditional sparse retrievers depend on exact matching approaches between high-dimensional bag-of-words (BoW) representations of both the queries and the collection. As a result, retrieval performance is restricted by semantic discrepancies and vocabulary gaps. On…
Álvaro García-Barragán, Ahmad Sakor, Maria-Esther Vidal, Ernestina Menasalvas + 3 more
'Ernestina Menasalvas' 'Juan Cristobal Sanchez Gonzalez' 'Mariano Provencio' 'Víctor Robles'] Abstract Accurate recognition and linking of oncologic entities in clinical notes is essential for extracting insights across cancer research, patient care, clinical decision-making, and treatment optimization. We present the…
Johannes M. van Hulst, Faegheh Hasibi, Koen Dercksen, Krisztian Balog + 1 more
'Krisztian Balog' 'Arjen P. de Vries'] Entity linking is a standard component in modern retrieval system that is often performed by third-party toolkits. Despite the plethora of open source options, it is difficult to find a single system that has a modular architecture where certain components may be replaced, does…
Nicholas Botzer, Yifan Ding, Tim Weninger
We introduce and make publicly available an entity linking dataset from Reddit that contains 17,316 linked entities, each annotated by three human annotators and then grouped into Gold, Silver, and Bronze to indicate inter-annotator agreement. We analyze the different errors and disagreements made by annotators and…
Jin G Zheng, Daniel Howsmon, Boliang Zhang, Juergen Hahn + 3 more
'Deborah McGuinness' 'James Hendler' 'Heng Ji'] Background The Entity Linking (EL) task links entity mentions from an unstructured document to entities in a knowledge base. Although this problem is well-studied in news and social media, this problem has not received much attention in the life science domain. One…
Majid Asgari-Bidhendi, Farzane Fakhrian, Behrouz Minaei-Bidgoli
In recent years, social media data has exponentially increased, which can be enumerated as one of the largest data repositories in the world. A large portion of this social media data is natural language text. However, the natural language is highly ambiguous due to exposure to the frequent occurrences of entities…
Tuan Lai, Heng Ji, ChengXiang Zhai
Entity linking (EL) is the task of linking entity mentions in a document to referent entities in a knowledge base (KB). Many previous studies focus on Wikipedia-derived KBs. There is little work on EL over Wikidata, even though it is the most extensive crowdsourced KB. The scale of Wikidata can open up many new…
Vaibhav Kasturia, Marcel Gohsen, Matthias Hagen
Web search queries can be ambiguous: is source of the nile meant to find information on the actual river or on a board game of that name? We tackle this problem by deriving entity-based query interpretations: given some query, the task is to derive all reasonable ways of linking suitable parts of the query to…
Munira Syed, Daheng Wang, Meng Jiang, Oliver Conway + 3 more
'Sriram Subramanian' 'Nitesh V. Chawla'] To improve consumer engagement and satisfaction, online news services employ strategies for personalizing and recommending articles to their users based on their interests. In addition to news agencies’ own digital platforms, they also leverage social media to reach out to a…
Jiexian Liu, Chen Zhang
In this paper, a multimodal knowledge mapping approach is used to digitize enterprise carbon assets, and a corresponding neural network model is designed for use in the practical process. Rich textual entity labels associated with images are obtained using an entity annotation system. A topology-based data fusion…
Yajie Luo, Yihong Wu, Muzhi Li, Fengran Mo + 5 more
Some Question Answering (QA) systems rely on knowledge bases (KBs) to provide accurate answers. Entity Linking (EL) plays a critical role in linking natural language mentions to KB entries. However, most existing EL methods are designed for long contexts and do not perform well on short, ambiguous user questions in QA…
Chuanqi Tan, Furu Wei, Pengjie Ren, Weifeng Lv + 1 more
We present a simple yet effective approach for linking entities in queries. The key idea is to search sentences similar to a query from Wikipedia articles and directly use the human-annotated entities in the similar sentences as candidate entities for the query. Then, we employ a rich set of features, such as…
Richard A A Jonker, Tiago Almeida, Rui Antunes, João R Almeida + 1 more
'Sérgio Matos'] Title: Abstract The identification of medical concepts from clinical narratives has a large interest in the biomedical scientific community due to its importance in treatment improvements or drug development research. Biomedical named entity recognition (NER) in clinical texts is crucial for automated…
Ghadeer Mobasher, Lukrécia Mertová, Sucheta Ghosh, Olga Krebs + 2 more
Chemical named entity recognition (NER) is a significant step for many downstream applications like entity linking for the chemical text-mining pipeline. However, the identification of chemical entities in a biomedical text is a challenging task due to the diverse morphology of chemical entities and the different types…
Paul Anderson, Damon Lin, Jean Davidson, Theresa Migler + 10 more
Link prediction and entity resolution play pivotal roles in uncovering hidden relationships within networks and ensuring data quality in the era of heterogeneous data integration. This paper explores the utilization of large language models to enhance link prediction, particularly through knowledge graphs derived from…
Noam H. Rotenberg, Robert Leaman, Rezarta Islamaj, Helena Kuivaniemi + 10 more
The variety of cell phenotypes identified by single-cell technologies is rapidly expanding, yet this knowledge is dispersed across the scientific literature and incompletely represented in structured resources. We present the CellLink corpus, a manually annotated collection of over 22,000 mentions of human and mouse…
John A Bachman, Benjamin M Gyori, Peter K Sorger
For automated reading of scientific publications to extract useful information about molecular mechanisms it is critical that genes, proteins and other entities be correctly associated with uniform identifiers, a process known as named entity linking or “grounding.” Correct grounding is essential for resolving…
Mohamad Yaser Jaradeh, Kuldeep Singh, Markus Stocker, Andreas Both + 1 more
'Sören Auer'] In the last decade, a large number of knowledge graph (KG) completion approaches were proposed. Albeit effective, these efforts are disjoint, and their collective strengths and weaknesses in effective KG completion have not been studied in the literature. We extend Plumber, a framework that brings…
Vitor D.T Andrade, Pedro Ruas, Francisco M. Couto
Biomedical literature is the main mean of communication for researchers to share their findings. Since biomedical literature is composed of a large collection of text expressed in natural language, the usage of text mining tools to extract information from those texts automatically is of utmost importance. The problem…
Roderic D. M. Page
Taxonomic names remain fundamental to linking biodiversity data, but information on these names resides in separate silos. Despite often making their contents available in RDF, records in these taxonomic databases are rarely linked to identifiers in external databases, such as DOIs for publications, or ORCIDs for…
Roderic D. M. Page
A major gap in the biodiversity knowledge graph is a connection between taxonomic names and the taxonomic literature. While both names and publications often have persistent identifiers (PIDs), such as Life Science Identifiers (LSIDs) or Digital Object Identifiers (DOIs), LSIDs for names are rarely linked to DOIs for…
Authors not listed
The nanosafety domain has seen significant advancements in data generation and sharing, yet challenges remain in ensuring data interoperability and reuse. This article focuses on developing a semantic interoperability framework for nanosafety data to maximize the FAIRness (Findability, Accessibility, Interoperability…
Authors not listed
As the volume and diversity of bioactivity data in ChEMBL continues to grow, ensuring that assay metadata is standardized, interoperable, and machine-readable is critical for effective use in cheminformatics and ML applications. In this work, we present recent efforts to enhance the quality and granularity of bioassay…
Rachana Niranjan Murthy, Sai Teja Potu, Akhil Thomas, Lokesh Mishra + 2 more
Retrieving structured materials information from unstructured textual data is essential for data mining and automatically developing comprehensive ontologies. Information extraction is a complex task composed of multiple subtasks and thus often relies on systems of task-specialized language models. A foundation…
Authors not listed
Sharing knowledge on chemicals in the digital age has been the playground of databases such as the Chemical Abstract Services and PubChem. Wikipedia complements this field by providing context to chemicals aimed at a broad audience, but is not easily read by machines. Wikidata was started as a database service to…
Philip Strömert, Johannes Hunold, Stuart Chalk, Leah McEwen + 1 more
This whitepaper aims to provide guidance to improve standardization and quality of ontologies development and curation concerning term definitions. We outline an approach on how definitions of The IUPAC Compendium of Chemical Terminology (colloquially known as the "Gold Book") could be applied as a definition source…