26 papers · ranked by Valyu relevance
David Gomez-Cabrero, Imad Abugessaisa, Dieter Maier, Andrew Teschendorff + 6 more
'Andrew Teschendorff' 'Matthias Merkenschlager' 'Andreas Gisel' 'Esteban Ballestar' 'Erik Bongcam-Rudloff' 'Ana Conesa' 'Jesper Tegnér'] To integrate heterogeneous and large omics data constitutes not only a conceptual challenge but a practical hurdle in the daily analysis of omics data. With the rise of novel omics…
Vasileios Lapatas, Michalis Stefanidakis, Rafael C. Jimenez, Allegra Via + 1 more
Data sharing, integration and annotation are essential to ensure the reproducibility of the analysis and interpretation of the experimental findings. Often these activities are perceived as a role that bioinformaticians and computer scientists have to take with no or little input from the experimental biologist. On the…
Amineh Amini, Hadi Saboohi, Nasser Nemat bakhsh
Data integration is one of the main problems in distributed data sources. An approach is to provide an integrated mediated schema for various data sources. This research work aims at developing a framework for defining an integrated schema and querying on it. The basic idea is to employ recent standard languages and…
Nadia Anwar, Ela Hunt
Background This paper summarises the lessons and experiences gained from a case study of the application of semantic web technologies to the integration of data from the bacterial species Francisella tularensis novicida (Fn). Fn data sources are disparate and heterogeneous, as multiple laboratories across the world…
Jason T. Liu
In information ecosystems, semantic heterogeneity is known as the root issue for the difficulties of data integration, and the Relational Model is not designed for addressing such challenges, i.e. to re-use data that is modeled from other sources in the local data model design. Although researchers have proposed many…
Martin Telefont
Large-scale data gathering efforts of the past, like the Human Genome Project, have shown that their most valuable contribution is data, which allows researchers to link their own experimental findings to them. This process, data integration, will be of central importance in how the scientific community to will be able…
G Eysenbach, Christian Gulden, Axel Newe, Benjamin Kinast + 3 more
'Hannes Ulrich' 'Björn Bergh' 'Björn Schreiweis'] Background In patient care, data are historically generated and stored in heterogeneous databases that are domain specific and often noninteroperable or isolated. As the amount of health data increases, the number of isolated data silos is also expected to grow…
El Kindi Rezig, Michael Cafarella, Vijay Gadepally
AI application developers typically begin with a dataset of interest and a vision of the end analytic or insight they wish to gain from the data at hand. Although these are two very important components of an AI workflow, one often spends the first few weeks (sometimes months) in the phase we refer to as data…
Gunther Eysenbach, Ieuan Clay, Christian Angelopoulos, Anne Lord Bailey + 9 more
Data integration, the processes by which data are aggregated, combined, and made available for use, has been key to the development and growth of many technological solutions. In health care, we are experiencing a revolution in the use of sensors to collect data on patient behaviors and experiences. Yet, the potential…
Fritz Laux, Malcolm Crowe
—Schema and data integration have been a challenge for more than 40 years. While data warehouse technologies are quite a success story, there is still a lack of information integration methods, especially if the data sources are based on different data models or do not have a schema. Enterprise Information Integration…
Anne E. Thessen, Paul Bogdan, David J. Patterson, Theresa M. Casey + 3 more
'César Hinojo-Hinojo' 'Orlando de Lange' 'Melissa A. Haendel'] Decades of reductionist approaches in biology have achieved spectacular progress, but the proliferation of subdisciplines, each with its own technical and social practices regarding data, impedes the growth of the multidisciplinary and interdisciplinary…
Andra Waagmeester, Gregory Stupp, Sebastian Burgstaller-Muehlbacher, Benjamin M. Good + 20 more
Wikidata is a community-maintained knowledge base that epitomizes the FAIR principles of Findability, Accessibility, Interoperability, and Reusability. Here, we describe the breadth and depth of biomedical knowledge contained within Wikidata, assembled from primary knowledge repositories on genomics, proteomics…
Julian Eberius, Katrin Braunschweig, Maik Thiele, Wolfgang Lehner
Open data platforms such as data.gov or opendata.socrata. com provide a huge amount of valuable information. Their free-for-all nature, the lack of publishing standards and the multitude of domains and authors represented on these platforms lead to new integration and standardization problems. At the same time…
Muhammad Fahad
This paper presents a semantic system named OntMed for an ontology-based data integration of heterogeneous data sources to achieve interoperability between heterogeneous data sources. Our system is based on the quality criteria (consistency, completeness and conciseness) for building the reliable analysis contexts to…
Ana Claudia Sima, Tarcisio Mendes de Farias, Erich Zbinden, Maria Anisimova + 5 more
Data integration promises to be one of the main catalysts in enabling new insights to be drawn from the wealth of biological data available publicly. However, the heterogeneity of the different data sources, both at the syntactic and the semantic level, still poses significant challenges for achieving interoperability…
Nicolas Le Guillarme, Wilfried Thuiller
With the rapid accumulation of biodiversity data, data integration has emerged as a hot topic in soil ecology. Data integration has indeed the potential to advance our knowledge of global patterns in soil biodiversity by facilitating large-scale meta-analytical studies of soil ecosystems. However, ecologists are still…
Maryam Alizadeh, Maliheh Heydarpour Shahrezaei, Farajollah Tahernezhad-Javazm
'Farajollah Tahernezhad-Javazm'] Abstract—An ontology makes a special vocabulary which describes the domain of interest and the meaning of the term on that vocabulary. Based on the precision of the specification, the concept of the ontology contains several data and conceptual models. The notion of ontology has emerged…
Authors not listed
The nanosafety domain has seen significant advancements in data generation and sharing, yet challenges remain in ensuring data interoperability and reuse. This article focuses on developing a semantic interoperability framework for nanosafety data to maximize the FAIRness (Findability, Accessibility, Interoperability…
Frédéric Burdet, Pierre-Marie Allard, Louis-Felix Nothias, Olivier Kirchhoffer + 16 more
Plants have a complex chemo-diversity and represent a reservoir of potential new therapeutic agents. Within a Swiss research project, six scientific research groups from different disciplines are collaborating to investigate a collection of more than 17’000 unique dried plant extracts. It aims to find new bioactive…
Flore N’kam Suguem, Sébastien Déjean, Philippe Saint Pierre, Nicolas Savy
One of the challenges encountered when merging heterogeneous observational clinical datasets is the recoding of categorical target variables that may have been measured differently across data sources. Standard machine learning-based approaches, such as Multiple Imputation by Chained Equations and the k-Nearest…
Authors not listed
Mass spectrometry (MS) is a cornerstone technology in modern molecular biology, powering diverse applications across proteomics, metabolomics, lipidomics, glycomics, and beyond. As the field continues to evolve, rapid advancements in instrumentation, acquisition strategies, machine learning, and scalable computing have…
Michael Statt, Brian Rohr, Dan Guevarra, Ja'Nya Breeden + 2 more
Materials knowledge is inherently hierarchical. While high-level descriptors such as composition and structure are valuable for contextualizing materials data, the data must ultimately be considered in the context of its low-level acquisition details. Graph databases offer an opportunity to represent hierarchical…
Authors not listed
High-quality data preprocessing is essential for untargeted metabolomics experiments, where increasing dataset scale and complexity demand adaptable, robust, and reproducible software solutions. Modern preprocessing tools must evolve to integrate seamlessly with downstream analysis platforms, ensuring efficient and…
Authors not listed
The exponential growth of chemical literature necessitates the development of automated tools for extracting and curating molecular information from unstructured scientific publications into open-access chemical databases. Current optical chemical structure recognition (OCSR) and named entity recognition solutions…
Bohdan B. Khomtchouk, Kasra A. Vand, Thor Wahlestedt, Kelly Khomtchouk + 2 more
We propose a search engine and file retrieval system for all bioinformatics databases worldwide. PubData searches biomedical data in a user-friendly fashion similar to how PubMed searches biomedical literature. PubData is built on novel network programming, natural language processing, and artificial intelligence…
Charlotte Neidiger, Tarek Saier, Kai Kühn, Victor Larignon + 12 more
In this work, a concept for an open chemistry knowledge base was developed to integrate chemical research results into a collaboratively usable platform. To achieve this, we enhanced Semantic MediaWiki (SMW) to support the collection and structured summary of chemical data contained in publications. We implemented…