13 papers · ranked by Valyu relevance
Aziz Fouché, Loïc Chadoutaud, Olivier Delattre, Andrei Zinovyev
Data integration of single-cell data describes the task of embedding datasets obtained from different sources into a common space, so that cells with similar cell type or state end up close from one another in this representation independently from their dataset of origin. Data integration is a crucial early step in…
Chao Wang, Michael J. O’Connell
In cancer research, different levels of high-dimensional data are often collected for the same subjects. Effective integration of these data by considering the shared and specific information from each data source can help us better understand different types of cancer. In this study we propose a novel autoencoder (AE)…
Jingru Zhang, Erjia Cui, Hongzhe Li, Haochang Shou
Wearable devices and digital phenotyping are increasingly used in observational and interventional studies to assess physical activity. However, integrating and comparing data across studies and cohorts remains challenging due to variability in device types, acquisition protocols, and preprocessing methods. A key…
Eva Luz Tejada-Gutierrez, Jordi Mateo-Fornés, Francesc Solsona, Rui Alves
DATABASE URL: https://forestforward.udl.cat Mitigating the effects of environmental exploitation on forests requires robust data analysis tools to inform sustainable management strategies and enhance ecosystem resilience. Access to extensive, integrated plant biodiversity data, spanning decades, is essential for this…
Yonatan Itai, Nimrod Rappoport, Ron Shamir
Integrative analysis of multi-omic datasets has proven to be extremely valuable in cancer research and precision medicine. However, obtaining multimodal data from the same samples is often difficult. Integrating multiple datasets of different omics remains a challenge, with only a few available algorithms developed to…
Anthony Huffman, Feng-Yu Yeh, Junguk Hur, Jie Zheng + 5 more
With the increasing volume of biomedical experimental data, standardizing, sharing, and integrating heterogeneous experimental data across domains has become a major challenge. To address this challenge, we have developed an ontology-supported Study-Experiment-Assay (SEA) common data model (CDM), which includes 10 core…
Anthony Huffman, Feng-Yu Yeh, Junguk Hur, Jie Zheng + 5 more
With the increasing volume of biomedical experimental data, standardizing, sharing, and integrating heterogeneous experimental data across domains has become a major challenge. To address this challenge, we have developed an ontology-supported Study-Experiment-Assay (SEA) common data model (CDM), which includes 10 core…
Nicolas Le Guillarme, Wilfried Thuiller
With the rapid accumulation of biodiversity data, data integration has emerged as a hot topic in soil ecology. Data integration has indeed the potential to advance our knowledge of global patterns in soil biodiversity by facilitating large-scale meta-analytical studies of soil ecosystems. However, ecologists are still…
Connor Bernard, Gabriel Silva Santos, Jacques Deere, Roberto Rodriguez-Caro + 5 more
The ecological sciences have joined the big data revolution. However, despite exponential growth in data availability, broader interoperability amongst datasets is still needed to unlock the potential of open access. The interface of demography and functional traits is well-positioned to benefit from said…
Theodoros Visvikis, Wouter-Michiel Vierdag, Luca Marconato, Ron M.A. Heeren + 1 more
Mass Spectrometry Imaging (MSI) is a powerful technique for mapping molecular distributions, and its integration with other imaging modalities is crucial for comprehensive understanding of molecular systems. Fragmented data formats and the limitations of existing standards like imzML, challenge spatial biology centric…
Alejandro Roldán, Tomás Golomb Durán, Antoni Josep Far, Maria Capa + 2 more
The era of Big Data has revolutionised biodiversity research, yet the potential of this information is frequently constrained by data heterogeneity, incompatible schemas, and the fragmentation of resources. Whilst standards such as Darwin Core have improved interoperability, significant barriers persist in harmonising…
Flore N’kam Suguem, Sébastien Déjean, Philippe Saint Pierre, Nicolas Savy
One of the challenges encountered when merging heterogeneous observational clinical datasets is the recoding of categorical target variables that may have been measured differently across data sources. Standard machine learning-based approaches, such as Multiple Imputation by Chained Equations and the k-Nearest…
Ginger Tsueng, Marco A. Alvarado Cano, José Bento, Candice Czech + 14 more
Biomedical datasets are increasing in size, stored in many repositories, and face challenges in FAIRness (findability, accessibility, interoperability, reusability). As a Consortium of infectious disease researchers from 15 Centers, we aim to adopt open science practices to promote transparency, encourage…