Search · four archives
Search · four archives
26 papers · ranked by Valyu relevance
Cassandra Königs, Marcel Friedrichs, Theresa Dietrich
Heterogeneous biomedical pharmacological databases are important for multiple fields in bioinformatics. Hetionet is a freely available database combining diverse entities and relationships from 29 public resources. Therefore, it is used as the basis for this project. 19 additional pharmacological medical and biological…
Leonardo Guerreiro Azevedo, R.F. Souza, Elton Soares, Raphael Thiago + 3 more
'Julio Cesar Cardoso Tesolin' 'A. C. C. de A. Oliveira' 'Márcio Ferreira Moreno'] Modern applications commonly need to manage dataset types composed of heterogeneous data and schemas, making it difficult to access them in an integrated way. A single data store to manage heterogeneous data using a common data model is…
Alejandro Roldán, Tomás Golomb Durán, Antoni Josep Far, Maria Capa + 2 more
The era of Big Data has revolutionised biodiversity research, yet the potential of this information is frequently constrained by data heterogeneity, incompatible schemas, and the fragmentation of resources. Whilst standards such as Darwin Core have improved interoperability, significant barriers persist in harmonising…
Priya Deshpande, Alexander Rasin, Roselyne Tchoua, Jacob Furst + 4 more
'Daniela Raicu' 'Michiel Schinkel' 'Hari Trivedi' 'Sameer Antani'] Data integration is a well-motivated problem in the clinical data science domain. Availability of patient data, reference clinical cases, and datasets for research have the potential to advance the healthcare industry. However, the unstructured (text…
Flore N’kam Suguem, Sébastien Déjean, Philippe Saint Pierre, Nicolas Savy
One of the challenges encountered when merging heterogeneous observational clinical datasets is the recoding of categorical target variables that may have been measured differently across data sources. Standard machine learning-based approaches, such as Multiple Imputation by Chained Equations and the k-Nearest…
Chuangtao Ma, Arijit Khan
Large language models (LLMs) and AI agents have demonstrated strong potential for data integration in zero-shot and few-shot settings. However, they continue to face significant accuracy and cost challenges in enterprise environments due to a persistent knowledge gap. This paper envisions trustworthy, scalable, and…
Makbule Gulcin Ozsoy
Large language models have significantly improved natural language interfaces to databases by translating user questions into executable queries. In particular, Text2Cypher focuses on generating Cypher queries for graph databases, enabling users to access graph data without query language expertise. Most existing…
Rihan Hai, Yan Kang, Christos Koutras, Andra Ionescu + 1 more
'Asterios Katsifodimos'] Abstract—Machine learning (ML) training data is often scattered across disparate collections of datasets, called data silos. This fragmentation poses a major challenge for data-intensive ML applications: integrating and transforming data residing in different sources demand a lot of manual work…
Christina Khnaisser, Luc Lavoie, Benoit Fraikin, Adrien Barton + 3 more
'Samuel Dussault' 'Anita Burgun' 'Jean-François Ethier'] Background A large volume of heavily fragmented data is generated daily in different healthcare contexts and is stored using various structures with different semantics. This fragmentation and heterogeneity make secondary use of data a challenge. Data integration…
Nicolas Le Guillarme, Wilfried Thuiller
With the rapid accumulation of biodiversity data, data integration has emerged as a hot topic in soil ecology. Data integration has indeed the potential to advance our knowledge of global patterns in soil biodiversity by facilitating large-scale meta-analytical studies of soil ecosystems. However, ecologists are still…
Frédéric Burdet, Pierre-Marie Allard, Louis-Felix Nothias, Olivier Kirchhoffer + 16 more
Plants have a complex chemo-diversity and represent a reservoir of potential new therapeutic agents. Within a Swiss research project, six scientific research groups from different disciplines are collaborating to investigate a collection of more than 17’000 unique dried plant extracts. It aims to find new bioactive…
Keti Korini, Christian Bizer
The goal of schema integration is, given a set of input schemata or tables, to derive a global, unified schema that is able to represent the concepts, attributes, and relationships of all input tables in a coherent fashion. This paper presents SINT-Flow, a schema integration framework composed of five LLM-based…
Ting Wang
Multi-source heterogeneous knowledge graph fusion faces significant challenges due to schema heterogeneity, entity conflicts, and relationship inconsistencies across different knowledge sources. This paper proposes CausalFusion, a novel adaptive fusion algorithm that leverages causal discovery principles to guide the…
Stefano Silvestri, Giuseppe Tricomi, Salvatore Rosario Bassolillo, Riccardo De Benedictis + 4 more
This paper describes a novel architecture that aims to create a template for the implementation of an IT platform, supporting the deployment and integration of the different digital twin subsystems that compose a complex urban intelligence system. In more detail, the proposed Smart City IT architecture has the…
Muhammad Fahad
This paper presents a semantic system named OntMed for an ontology-based data integration of heterogeneous data sources to achieve interoperability between heterogeneous data sources. Our system is based on the quality criteria (consistency, completeness and conciseness) for building the reliable analysis contexts to…
Ralf Bill, Jörg Blankenbach, Martin Breunig, Jan-Henrik Haunert + 10 more
'Christian Heipke' 'Stefan Herle' 'Hans-Gerd Maas' 'Helmut Mayer' 'Liqui Meng' 'Franz Rottensteiner' 'Jochen Schiewe' 'Monika Sester' 'Uwe Sörgel' 'Martin Werner'] Geospatial information science (GI science) is concerned with the development and application of geodetic and information science methods for modeling…
Gianluca Cima, Marco Console, Maurizio Lenzerini, Antonella Poggi
It is well-known that Artificial Intelligence (AI), and in particular Machine Learning (ML), is not effective without good data preparation, as also pointed out by the recent wave of data-centric AI. Data preparation is the process of gathering, transforming and cleaning raw data prior to processing and analysis. Since…
Dave Bunten, Jenna Tomkinson, Erik Serrano, Michael J. Lippincott + 4 more
High-content imaging (HCI) involves the automated acquisition and quantitative analysis of cell phenotypes from microscopy images. These studies often rely on screening, which can involve thousands of chemical or genetic perturbations that produce terabytes of microscopy data. To extract meaningful biological insights…
Authors not listed
Artificial intelligence (AI) is poised to transform heterogeneous catalysis, ushering in a new paradigm for catalytic materials discovery. By uncovering intricate patterns in high-dimensional data, AI has been reshaping our pursuit of sustainable catalytic processes across the energy, environmental, and chemical…
Nishu Nehra, Rohit Swami, Dharani Dadi, Ritika Mishra + 5 more
Getting clinical data from different sources to “talk” to each other within the OMOP Common Data Model (CDM) is arguably the most tedious part of multi-center research. While this integration is essential, the transformation process is frequently a manual grind, requiring a rare overlap of deep clinical knowledge and…
Zhengtong Yan, Yuan, Gongsheng, Qing-Yi Guo + 1 more
Modern enterprises are increasingly driven by the DATA+AI paradigm, in which Database Management Systems (DBMSs) and Large Language Models (LLMs) have become two foundational infrastructures powering a wide range of industrial and business applications, such as enterprise analytics, intelligent customer service, and…
Ragunathan Mariappan, Aishwarya Jayagopal, Ho Zong Sien, Vaibhav Rajan
In many biomedical studies, there arises the need to integrate data from multiple directly or indirectly related sources. Collective matrix factorization (CMF) and its variants are models designed to collectively learn from arbitrary collections of matrices. The latent factors learnt are rich integrative…
Martin Starman, Fabian Kirchner, Martin Held, Catriona Eschke + 4 more
Electronic Lab Notebooks (ELNs) have become indispensable tools for modern research laboratories, facilitating data management, collaboration, and documentation of scientific experiments. However, the proliferation of diverse ELN platforms poses challenges for researchers who need to seamlessly exchange data between…
Authors not listed
The nanosafety domain has seen significant advancements in data generation and sharing, yet challenges remain in ensuring data interoperability and reuse. This article focuses on developing a semantic interoperability framework for nanosafety data to maximize the FAIRness (Findability, Accessibility, Interoperability…
Michael Statt, Brian Rohr, Dan Guevarra, Ja'Nya Breeden + 2 more
Materials knowledge is inherently hierarchical. While high-level descriptors such as composition and structure are valuable for contextualizing materials data, the data must ultimately be considered in the context of its low-level acquisition details. Graph databases offer an opportunity to represent hierarchical…
Arnaud Gaudry, Marco Pagni, Florence Mehl, Sébastien Moretti + 10 more
Modern natural products (NPs) research relies on untargeted liquid chromatography coupled with mass spectrometry metabolomics. Together with cutting-edge processing and computational annotation strategies, such approaches can yield extensive spectral and structural information. However, current processing workflows…