Search · four archives
Search · four archives
26 papers · ranked by Valyu relevance
Vadim Y. Bichutskiy, Richard Colman, Rainer K. Brachmann, Richard H. Lathrop
'Richard H. Lathrop'] Complex problems in life science research give rise to multidisciplinary collaboration, and hence, to the need for heterogeneous database integration. The tumor suppressor p53 is mutated in close to 50% of human cancers, and a small drug-like molecule with the ability to restore native function to…
Giuseppe Fusco, Lerina Aversano, Sebastian Ventura
Integrating data from multiple heterogeneous data sources entails dealing with data distributed among heterogeneous information sources, which can be structured, semi-structured or unstructured, and providing the user with a unified view of these data. Thus, in general, gathering information is challenging, and one of…
Vijay Gadepally, Jeremy Kepner
Data forms a key component of any enterprise. The need for high quality and easy access to data is further amplified by organizations wishing to leverage machine learning or artificial intelligence for their operations. To this end, many organizations are building resources for managing heterogenous data, providing…
Raul Castro Fernandez, Samuel Madden
Data-driven analysis is important in virtually every modern organization. Yet, most data is underutilized because it remains locked in silos inside of organizations; large organizations have thousands of databases, and billions of files that are not integrated together in a single, queryable repository. Despite 40+…
Leonardo Guerreiro Azevedo, R.F. Souza, Elton Soares, Raphael Thiago + 3 more
'Julio Cesar Cardoso Tesolin' 'A. C. C. de A. Oliveira' 'Márcio Ferreira Moreno'] Modern applications commonly need to manage dataset types composed of heterogeneous data and schemas, making it difficult to access them in an integrated way. A single data store to manage heterogeneous data using a common data model is…
Wang Yan, Le Jiajin, Zhang Yun
The main challenges that marine heterogeneous data integration faces are the problem of accurate schema mapping between heterogeneous data sources. In order to improve the schema mapping efficiency and get more accurate learning results, this paper proposes a heterogeneous data schema mapping method basing on…
Alejandro Roldán, Tomás Golomb Durán, Antoni Josep Far, Maria Capa + 2 more
The era of Big Data has revolutionised biodiversity research, yet the potential of this information is frequently constrained by data heterogeneity, incompatible schemas, and the fragmentation of resources. Whilst standards such as Darwin Core have improved interoperability, significant barriers persist in harmonising…
Nadia Anwar, Ela Hunt
Background This paper summarises the lessons and experiences gained from a case study of the application of semantic web technologies to the integration of data from the bacterial species Francisella tularensis novicida (Fn). Fn data sources are disparate and heterogeneous, as multiple laboratories across the world…
Teng Lin
—The integration of heterogeneous databases into a unified querying framework remains a critical challenge, particularly in resource-constrained environments. This paper presents a novel Small Language Model (SLM)-driven system that synergizes advancements in lightweight Retrieval-Augmented Generation (RAG) and…
Rahman Ali, Muhammad Hameed Siddiqi, Muhammad Idris, Taqdir Ali + 5 more
'Shujaat Hussain' 'Eui-Nam Huh' 'Byeong Ho Kang' 'Sungyoung Lee' 'Vittorio M.N. Passaro'] A wide array of biomedical data are generated and made available to healthcare experts. However, due to the diverse nature of data, it is difficult to predict outcomes from it. It is therefore necessary to combine these diverse…
Flore N’kam Suguem, Sébastien Déjean, Philippe Saint Pierre, Nicolas Savy
One of the challenges encountered when merging heterogeneous observational clinical datasets is the recoding of categorical target variables that may have been measured differently across data sources. Standard machine learning-based approaches, such as Multiple Imputation by Chained Equations and the k-Nearest…
Amineh Amini, Hadi Saboohi, Nasser Nemat bakhsh
Data integration is one of the main problems in distributed data sources. An approach is to provide an integrated mediated schema for various data sources. This research work aims at developing a framework for defining an integrated schema and querying on it. The basic idea is to employ recent standard languages and…
Jason T. Liu
In information ecosystems, semantic heterogeneity is known as the root issue for the difficulties of data integration, and the Relational Model is not designed for addressing such challenges, i.e. to re-use data that is modeled from other sources in the local data model design. Although researchers have proposed many…
Nicolas Le Guillarme, Wilfried Thuiller
With the rapid accumulation of biodiversity data, data integration has emerged as a hot topic in soil ecology. Data integration has indeed the potential to advance our knowledge of global patterns in soil biodiversity by facilitating large-scale meta-analytical studies of soil ecosystems. However, ecologists are still…
Frédéric Burdet, Pierre-Marie Allard, Louis-Felix Nothias, Olivier Kirchhoffer + 16 more
Plants have a complex chemo-diversity and represent a reservoir of potential new therapeutic agents. Within a Swiss research project, six scientific research groups from different disciplines are collaborating to investigate a collection of more than 17’000 unique dried plant extracts. It aims to find new bioactive…
Rafat Hammad, Malek Barhoush, Bilal H. Abed-alguni
Healthcare information systems can reduce the expenses of treatment, foresee episodes of pestilences, help stay away from preventable illnesses, and improve personal life satisfaction. As of late, considerable volumes of heterogeneous and differing medicinal services data are being produced from different sources…
Yan Chak Li, Linhua Wang, Jeffrey N. Law, T. M. Murali + 1 more
Integrating multimodal data represents an effective approach to predicting biomedical characteristics, such as protein functions and disease outcomes. However, existing data integration approaches do not sufficiently address the heterogeneous semantics of multimodal data. In particular, early and intermediate…
Muhammad Fahad
This paper presents a semantic system named OntMed for an ontology-based data integration of heterogeneous data sources to achieve interoperability between heterogeneous data sources. Our system is based on the quality criteria (consistency, completeness and conciseness) for building the reliable analysis contexts to…
Ralf Bill, Jörg Blankenbach, Martin Breunig, Jan-Henrik Haunert + 10 more
'Christian Heipke' 'Stefan Herle' 'Hans-Gerd Maas' 'Helmut Mayer' 'Liqui Meng' 'Franz Rottensteiner' 'Jochen Schiewe' 'Monika Sester' 'Uwe Sörgel' 'Martin Werner'] Geospatial information science (GI science) is concerned with the development and application of geodetic and information science methods for modeling…
Authors not listed
Artificial intelligence (AI) is poised to transform heterogeneous catalysis, ushering in a new paradigm for catalytic materials discovery. By uncovering intricate patterns in high-dimensional data, AI has been reshaping our pursuit of sustainable catalytic processes across the energy, environmental, and chemical…
Giovanni Delussu, Luca Lianas, Francesca Frexia, Gianluigi Zanetti
This work presents a scalable data access layer, called PyEHR, intended for building data management systems for secondary use of structured heterogeneous biomedical and clinical data. PyEHR adopts openEHR formalisms to guarantee the decoupling of data descriptions from implementation details and exploits structures…
Martin Starman, Fabian Kirchner, Martin Held, Catriona Eschke + 4 more
Electronic Lab Notebooks (ELNs) have become indispensable tools for modern research laboratories, facilitating data management, collaboration, and documentation of scientific experiments. However, the proliferation of diverse ELN platforms poses challenges for researchers who need to seamlessly exchange data between…
Authors not listed
The nanosafety domain has seen significant advancements in data generation and sharing, yet challenges remain in ensuring data interoperability and reuse. This article focuses on developing a semantic interoperability framework for nanosafety data to maximize the FAIRness (Findability, Accessibility, Interoperability…
Michael Statt, Brian Rohr, Dan Guevarra, Ja'Nya Breeden + 2 more
Materials knowledge is inherently hierarchical. While high-level descriptors such as composition and structure are valuable for contextualizing materials data, the data must ultimately be considered in the context of its low-level acquisition details. Graph databases offer an opportunity to represent hierarchical…
Michael Statt, Kristopher Brown, Santosh Suram, Linda Hung + 3 more
In this work, we present DBgen, a Python library that provides a framework for defining extract-transform-load (ETL) pipelines to create and populate SQL databases. DBgen is most useful when the underlying data has complex relationships, requires multi-step analysis, is large-scale, and the type of data being collected…
Authors not listed
The exponential growth of chemical literature necessitates the development of automated tools for extracting and curating molecular information from unstructured scientific publications into open-access chemical databases. Current optical chemical structure recognition (OCSR) and named entity recognition solutions…