25 papers · ranked by Valyu relevance
Gunther Eysenbach, Ieuan Clay, Christian Angelopoulos, Anne Lord Bailey + 9 more
Data integration, the processes by which data are aggregated, combined, and made available for use, has been key to the development and growth of many technological solutions. In health care, we are experiencing a revolution in the use of sensors to collect data on patient behaviors and experiences. Yet, the potential…
Ralf Bill, Jörg Blankenbach, Martin Breunig, Jan-Henrik Haunert + 10 more
'Christian Heipke' 'Stefan Herle' 'Hans-Gerd Maas' 'Helmut Mayer' 'Liqui Meng' 'Franz Rottensteiner' 'Jochen Schiewe' 'Monika Sester' 'Uwe Sörgel' 'Martin Werner'] Geospatial information science (GI science) is concerned with the development and application of geodetic and information science methods for modeling…
Ryu Kyung Kim, Young Min Kim, Won Jin Lee, Jongho Im + 4 more
Data integration including statistical matching, and record linkage is the process of merging information from several heterogeneous datasets . Data integration and analysis using linked data have been steadily studied in epidemiology . Data integration has the advantage of enabling users or researchers to implement…
Aziz Fouché, Loïc Chadoutaud, Olivier Delattre, Andrei Zinovyev
Data integration of single-cell data describes the task of embedding datasets obtained from different sources into a common space, so that cells with similar cell type or state end up close from one another in this representation independently from their dataset of origin. Data integration is a crucial early step in…
Chao Wang, Michael J. O’Connell
In cancer research, different levels of high-dimensional data are often collected for the same subjects. Effective integration of these data by considering the shared and specific information from each data source can help us better understand different types of cancer. In this study we propose a novel autoencoder (AE)…
Sharon Zanti, Emily Berkowitz, Matthew Katz, Amy Hawn Nelson + 3 more
'T.C. Burnett' 'Dennis Culhane' 'Yixi Zhou'] Use of administrative data to inform decision making is now commonplace throughout the public sector, including program and policy evaluation. While reuse of these data can reduce costs, improve methodologies, and shorten timelines, challenges remain. This article informs…
Roman Lukyanenko
Data Management Authors: ['Roman Lukyanenko'] In an era dominated by information technology, the critical discipline of data management remains undervalued compared to the innovations it enables, such as artificial intelligence and social media. The ambiguity surrounding what constitutes data management and its…
G Eysenbach, Christian Gulden, Axel Newe, Benjamin Kinast + 3 more
'Hannes Ulrich' 'Björn Bergh' 'Björn Schreiweis'] Background In patient care, data are historically generated and stored in heterogeneous databases that are domain specific and often noninteroperable or isolated. As the amount of health data increases, the number of isolated data silos is also expected to grow…
Alon Halevy, Jane Dwivedi-Yu
One of the limitations of large language models is that they do not have access to up-to-date, proprietary or personal data. As a result, there are multiple efforts to extend language models with techniques for accessing external data. In that sense, LLMs share the vision of data integration systems whose goal is to…
Adam Coscia, Ashley Suh, Remco Chang, Alex Endert
—Data integration is often performed to consolidate information from multiple disparate data sources during visual data analysis. However, integration operations are usually separate from visual analytics operations such as encode and filter in both interface design and empirical research. We conducted a preliminary…
Muhammad Fahad
This paper presents a semantic system named OntMed for an ontology-based data integration of heterogeneous data sources to achieve interoperability between heterogeneous data sources. Our system is based on the quality criteria (consistency, completeness and conciseness) for building the reliable analysis contexts to…
Anthony Huffman, Feng-Yu Yeh, Junguk Hur, Jie Zheng + 5 more
With the increasing volume of biomedical experimental data, standardizing, sharing, and integrating heterogeneous experimental data across domains has become a major challenge. To address this challenge, we have developed an ontology-supported Study-Experiment-Assay (SEA) common data model (CDM), which includes 10 core…
Nicolas Le Guillarme, Wilfried Thuiller
With the rapid accumulation of biodiversity data, data integration has emerged as a hot topic in soil ecology. Data integration has indeed the potential to advance our knowledge of global patterns in soil biodiversity by facilitating large-scale meta-analytical studies of soil ecosystems. However, ecologists are still…
Ahmed Fawzy, Amjed Tahir, Matthias Galster, Peng Liang
Literature Review Authors: ['Ahmed Fawzy' 'Amjed Tahir' 'Matthias Galster' 'Peng Liang'] Method: We used a Systematic Literature Review (SLR) to collect and analyse relevant studies. We identified 45 studies related to data management in agile software development. We then manually analysed and mapped data from these…
Alejandro Roldán, Tomás Golomb Durán, Antoni Josep Far, Maria Capa + 2 more
The era of Big Data has revolutionised biodiversity research, yet the potential of this information is frequently constrained by data heterogeneity, incompatible schemas, and the fragmentation of resources. Whilst standards such as Darwin Core have improved interoperability, significant barriers persist in harmonising…
Authors not listed
The nanosafety domain has seen significant advancements in data generation and sharing, yet challenges remain in ensuring data interoperability and reuse. This article focuses on developing a semantic interoperability framework for nanosafety data to maximize the FAIRness (Findability, Accessibility, Interoperability…
Frédéric Burdet, Pierre-Marie Allard, Louis-Felix Nothias, Olivier Kirchhoffer + 16 more
Plants have a complex chemo-diversity and represent a reservoir of potential new therapeutic agents. Within a Swiss research project, six scientific research groups from different disciplines are collaborating to investigate a collection of more than 17’000 unique dried plant extracts. It aims to find new bioactive…
Angelo Martella, Cristian Martella, Antonella Longo
> Abstract. The emerging paradigm of data economy can constitute an unmissable and attractive opportunity for companies that aim to consider their data as valuable assets. To fully leverage this opportunity, data owners need to have specific and precise guarantees regarding the protection of data they share from…
Sergi Nadal, Petar Jovanovic, Besim Bilalli, Oscar Romero
The ability to cross data from multiple sources represents a competitive advantage for organizations. Yet, the governance of the data lifecycle, from the data sources into valuable insights, is largely performed in an ad-hoc or manual manner. This is specifically concerning in scenarios where tens or hundreds of…
Flore N’kam Suguem, Sébastien Déjean, Philippe Saint Pierre, Nicolas Savy
One of the challenges encountered when merging heterogeneous observational clinical datasets is the recoding of categorical target variables that may have been measured differently across data sources. Standard machine learning-based approaches, such as Multiple Imputation by Chained Equations and the k-Nearest…
Authors not listed
Mass spectrometry (MS) is a cornerstone technology in modern molecular biology, powering diverse applications across proteomics, metabolomics, lipidomics, glycomics, and beyond. As the field continues to evolve, rapid advancements in instrumentation, acquisition strategies, machine learning, and scalable computing have…
Authors not listed
Artificial intelligence (AI) is poised to transform heterogeneous catalysis, ushering in a new paradigm for catalytic materials discovery. By uncovering intricate patterns in high-dimensional data, AI has been reshaping our pursuit of sustainable catalytic processes across the energy, environmental, and chemical…
Authors not listed
High-quality data preprocessing is essential for untargeted metabolomics experiments, where increasing dataset scale and complexity demand adaptable, robust, and reproducible software solutions. Modern preprocessing tools must evolve to integrate seamlessly with downstream analysis platforms, ensuring efficient and…
Authors not listed
The exponential growth of chemical literature necessitates the development of automated tools for extracting and curating molecular information from unstructured scientific publications into open-access chemical databases. Current optical chemical structure recognition (OCSR) and named entity recognition solutions…
Charlotte Neidiger, Tarek Saier, Kai Kühn, Victor Larignon + 12 more
In this work, a concept for an open chemistry knowledge base was developed to integrate chemical research results into a collaboratively usable platform. To achieve this, we enhanced Semantic MediaWiki (SMW) to support the collection and structured summary of chemical data contained in publications. We implemented…