23 papers · ranked by Valyu relevance
Henk van der Pol, Tina Kringelbach, Maria Martin Agudo, Gabriel Bratseeth Stav + 19 more
When merging the data from different DLCTs, there are three critical decision triggering time-points: 1. The first decision involves choosing the cohorts to merge from the individual trials. Due to the high complexity of particularly genomic biomarkers and cancer type definitions, this is not limited to the…
Abraham D. Smith, Paul Bendich, John Harer
Collections of measures on compact metric spaces form a model category ("data complexes"), whose morphisms are marginalization integrals. The fibrant objects in this category represent collections of measures in which there is a measure on a product space that marginalizes to any measures on pairs of its factors. The…
Jingjing Yu, Xiao-Feng Li, Elizabeth Lewis, Stephen Blenkinsop + 1 more
'Hayley J. Fowler'] There is an urgent need for high-quality and high-spatial-resolution hourly precipitation products around the globe, including the UK. Although hourly precipitation products exist for the UK, these either contain large errors, or are insufficient in spatial resolution. An efficient way to solve this…
Aaron S. Brewster, Daniel W. Paley, Asmit Bhowmick, David W. Mittan-Moreau + 5 more
The cctbx.xfel suite of processing programs and tools allows fast, visual analysis of serial diffraction images from synchrotrons and XFELs. Built on DIALS and cctbx, cctbx.xfel is designed for real-time and post-experiment processing with a fully featured graphical user interface. Users can quickly identify hitrates…
Xiaobo Sun, Jingjing Gao, Peng Jin, Celeste Eng + 7 more
In this report, we describe three cluster-based schemas running on the Apache Hadoop (MapReduce), HBase, and Spark platforms for performing sorted merging of variants identified from WGS. We show that all three schemas are scalable on both input data size and computing resources, suggesting large-scale “-omics” data…
Flore N’kam Suguem, Sébastien Déjean, Philippe Saint Pierre, Nicolas Savy
One of the challenges encountered when merging heterogeneous observational clinical datasets is the recoding of categorical target variables that may have been measured differently across data sources. Standard machine learning-based approaches, such as Multiple Imputation by Chained Equations and the k-Nearest…
Yuzhao Yang, Jérôme Darmont, Franck Ravat, Olivier Teste
Using data warehouses to analyse multidimensional data is a significant task in company decision-making. The need for analyzing data stored in different data warehouses generates the requirement of merging them into one integrated data warehouse. The data warehouse merging process is composed of two steps: matching…
José Roberto Lomelí-Huerta, Juan Pablo Rivera-Caicedo, Miguel De-la-Torre, Brenda Acevedo-Juárez + 3 more
This paper proposes an approach to fill in missing data from satellite images using data-intensive computing platforms. The proposed approach merges satellite imagery from diverse sources to reduce the impact of the holes in images that result from acquisition conditions: occlusion, the satellite trajectory, sunlight…
Nalin Ranjan, Zechao Shang, Sanjay Krishnan, Aaron J. Elmore
We propose MindPalace, a prototype of a versioned database for efficient collaborative data management. MindPalace supports offline collaboration, where users work independently without real-time correspondence. The core of MindPalace is a critical step of offline collaboration: reconciling divergent branches made by…
Yimei Li, Matt Hall, Brian T. Fisher, Alix E. Seif + 10 more
National Cancer Institute (NCI)-funded cooperative group clinical trials have improved cure rates for children with cancer and have set standards of care for the treatment of adult malignancies . However, such clinical trials data have important limitations, particularly the lack of resource utilization and cost data.…
Agneta Egenvall, Ane Nødtvedt, Lars Roepstorff, Brenda Bonnett
In a world of limited resources, using existing databases in research is a potentially cost-effective way to increase knowledge, given that correct and meaningful results are gained. Nordic examples of the use of secondary small animal and equine databases include studies based on data from tumour registries, breeding…
Xueyuan Ren, Frank Li, Yang Wang
—This paper explores a new opportunity to improve the performance of transaction processing at the application side by merging structurely similar statements or transactions. Concretely, we re-write transactions to 1) merge similar statements using specific SQL semantics; 2) eliminate redundant reads; and 3) merge…
Ksenia Khelik, Alexander Johan Nederbragt, Geir Kjetil Sandve, Torbjørn Rognes
In spite of the major breakthroughs in the second-generation sequencing technologies and the developments of a plethora of assemblers over the last ten years, the resulting genome assemblies may still be fragmented and contain errors. It is typical in genome projects with second-generation reads involved to run…
Noah Brown, Charles Danis, Vazira Ahmedjanova, Jennifer L. Guler
Structural variants (SVs) are abundant across all life, and have major impacts on the genome and transcriptome. However, it is difficult to appreciate the individual significance of SVs when they are heterogeneously distributed across a genomic neighborhood. Further, low-input sequencing technologies or sequencing of…
Javier Cabau-Laporta, Alex M. Ascensión, Mikel Arrospide-Elgarresta, Daniela Gerovska + 1 more
High-throughput cell-data technologies such as single-cell RNA-Seq create a demand for algorithms for automatic cell classification and characterization. There exist several classification ontologies of cells with complementary information. However, one needs to merge them in order to combine synergistically their…
Federico Ottomano, Giovanni De Felice, Vladimir Gusev, Taylor Sparks
Recent Machine Learning (ML) developments have opened new perspectives on accelerating the discovery of new materials. However, in the field of materials informatics, the performance of ML estimators is heavily limited by the nature of the available training datasets, which are often severely restricted and unbalanced.…
Nora yahia Ibrahim, Sahar A. Mokhtar, Hany Harb
—This Ontologies are widely used as a means for solving the information heterogeneity problems on the web because of their capability to provide explicit meaning to the information. They become an efficient tool for knowledge representation in a structured manner. There is always more than one ontology for the same…
Oļegs Verhodubs
Ontology merging is important, but not always effective. The main reason, why ontology merging is not effective, is that ontology merging is performed without considering goals. Goals define the way, in which ontologies to be merged more effectively. The paper illustrates ontology merging by means of rules, which are…
Authors not listed
In recent years, the development of large language models (LLMs) has revolutionized various fields of natural science, yet their application in molecular data processing remains constrained due to the reliance on single-modality inputs and outputs. To bridge the gap between experimenters and computational tools, we…
Matteo P. Ferla, Rubén Sánchez-García, Rachael E. Skyner, Stefan Gahbauer + 4 more
Current strategies centred on either merging or linking initial hits from fragment-based drug design (FBDD) crystallographic screens ignore 3D structural information. We show that an algorithmic approach (Fragmenstein) that ‘stitches’ the ligand atoms from this structural information together can provide more accurate…
Michael Statt, Brian Rohr, Dan Guevarra, Ja'Nya Breeden + 2 more
Materials knowledge is inherently hierarchical. While high-level descriptors such as composition and structure are valuable for contextualizing materials data, the data must ultimately be considered in the context of its low-level acquisition details. Graph databases offer an opportunity to represent hierarchical…
Authors not listed
Fragment-based drug discovery (FBDD) is a widely used strategy in early-stage drug development, but accurately predicting the binding affinities of fragments and their elaborated analogs poses unique computational challenges. These difficulties arise from weak binding affinities, diverse chemical scaffolds, and limited…
Michael Statt, Kristopher Brown, Santosh Suram, Linda Hung + 3 more
In this work, we present DBgen, a Python library that provides a framework for defining extract-transform-load (ETL) pipelines to create and populate SQL databases. DBgen is most useful when the underlying data has complex relationships, requires multi-step analysis, is large-scale, and the type of data being collected…