24 papers · ranked by Valyu relevance
Peter Sestoft
Research relies on ever larger amounts of data from experiments, automated production equipment, questionnaries, times series such as weather records, and so on. A major task in science is to combine, process and analyse such data to obtain evidence of patterns and correlations. Most research data are on digital form…
Dexian Yang, Jiong Yu, Zhenzhen He, Ping Li + 1 more
This study explores the analysis and modeling of energy consumption in the context of database workloads, aiming to develop an eco-friendly database management system (DBMS). It leverages vibration energy harvesting systems with self-sustaining wireless vibration sensors (WVSs) in combination with the least square…
Aleem Akhtar
—Databases are considered to be integral part of modern information systems. Almost every web or mobile application uses some kind of database. Database management systems are considered to be a crucial element from both business and technological standpoint. This paper divides different types of database management…
Paul L. Soto
Data collection and analysis are central to scientific research, including in applied and basic behavior analysis. A substantial amount of attention has been given to how to rigorously collect and analyze data. Less attention has been paid to storing and maintaining research data, which becomes a critical step in the…
Agapi Rissaki, Ilias Fountalis, Wolfgang Gatterbauer, Benny Kimelfeld
In recent years, there has been significant progress in the development of deep learning models over relational databases, including architectures based on heterogeneous graph neural networks (hetero-GNNs) and heterogeneous graph transformers. In effect, such architectures state how the database records and links…
Dimitri Yatsenko, Edgar Y. Walker, Andreas S. Tolias
The relational data model offers unrivaled rigor and precision in defining data structure and querying complex data. Yet the use of relational databases in scientific data pipelines is limited due to their perceived unwieldiness. We propose a simplified and conceptually refined relational data model named DataJoint.…
Jakub Galgonek, Jiří Vondrášek
Current biological and chemical research is increasingly dependent on the reusability of previously acquired data, which typically come from various sources. Consequently, there is a growing need for database systems and databases stored in them to be interoperable with each other. One of the possible solutions to…
Jakub Peleška, Gustav Šír
—Transformer models have continuously expanded into all machine learning domains convertible to the underlying sequence-to-sequence representation, including tabular data. However, while ubiquitous, this representation restricts their extension to the more general case of relational databases. In this paper, we…
Dhanamma Jagli, Priyanka Gaikwad, Shubhangi Gunjal, Chaitanya Bilaware
'Chaitanya Bilaware'] Abstract. This paper describes practical observations during the Database system Lab. Oracle 10g DBMS is used in the data base system lab and performed SQL queries based many concepts like Data Definition Language Commands (DDL), Data Modification Language Commands ((DML), Views, Integrity…
Mohamed Hassan
In the world of data management, the rapidly growing area of big data offers both intriguing possibilities and substantial obstacles. Big data, defined by its enormous size, complex structure, and varied sources - which include everything from social media engagement and sensor data to financial transactions and…
Jeremy Kepner, Vijay Gadepally, David Hutchison, Hayden Jananthan + 3 more
'Timothy G. Mattson' 'Siddharth Samsi' 'Albert Reuther'] Abstract—The success of SQL, NoSQL, and NewSQL databases is a reflection of their ability to provide significant functionality and performance benefits for specific domains, such as financial transactions, internet search, and data analysis. The BigDAWG polystore…
Dimitri Yatsenko, Jacob Reimer, Alexander S. Ecker, Edgar Y. Walker + 6 more
The rise of big data in modern research poses serious challenges for data management: Large and intricate datasets from diverse instrumentation must be precisely aligned, annotated, and processed in a variety of ways to extract new insights. While high levels of data integrity are expected, research teams have diverse…
Keith J. Fraga, Yuanpeng J. Huang, Theresa A. Ramelot, G.V.T. Swapna + 4 more
NMR is a valuable experimental tool in the structural biologist’s toolkit to elucidate the structures, functions, and motions of biomolecules. The progress of machine learning, particularly in structural biology, reveals the critical importance of large, diverse, and reliable datasets in developing new methods and…
Authors not listed
Mass spectrometry (MS) generates large datasets that are stored in increasingly optimized and complex file types, demanding technical expertise to extract information rapidly and easily. We wondered whether a simple structured query language (SQL) database could hold raw MS data and allow for easily readable queries…
Mateusz Jundzill, Riccardo Spott, Mara Lohde, Martin Hölzer + 2 more
'Adrian Viehweger' 'Christian Brandt'] Title: Abstract With the rapidly growing amount of biological data, powerful but also flexible data management and visualization systems are of increasingly crucial importance. The COVID-19 pandemic has more than highlighted this need and the challenges scientists are facing.…
Zongyue Qin, Luo Chen, Zhengyang Wang, Haoming Jiang + 1 more
Large language models (LLMs) excel in many natural language processing (NLP) tasks. However, since LLMs can only incorporate new knowledge through training or supervised finetuning processes, they are unsuitable for applications that demand precise, up-to-date, and private information not available in the training…
Michael Statt, Kristopher Brown, Santosh Suram, Linda Hung + 3 more
In this work, we present DBgen, a Python library that provides a framework for defining extract-transform-load (ETL) pipelines to create and populate SQL databases. DBgen is most useful when the underlying data has complex relationships, requires multi-step analysis, is large-scale, and the type of data being collected…
Authors not listed
This work presents the LabIMotion extension for the Chemotion Electronic Lab Notebook (ELN), expanding its capabilities from organic chemistry to support interdisciplinary research and enabling the description of workflows. LabIMotion enhances documentation by introducing customizable components structured across three…
Venkata Chandrasekhar Nainala, Kohulan Rajan, Sri Ram Sagar Kanakam, Nisha Sharma + 3 more
The COCONUT (COlleCtion of Open Natural prodUcTs) database was launched in 2021 as an aggregation of openly available natural product datasets and has been one of the biggest open natural product databases since. Apart from the chemical structures of natural products, COCONUT contains information about names and…
Sajid Mughal, Ismail Moghul, Jing Yu, Tristan Clark + 2 more
Efficient storage and querying of large amounts of genetic and phenotypic data is crucial to contemporary clinical genetic research. This introduces computational challenges for classical relational databases, due to the sparsity and sheer volume of the data. Our Java based solution loads annotated genetic variants and…
Marco Falda
FuzzyVariantExplorer is a web interface and a fuzzy search system for exploring richly annotated genomes in a flexible way. The application allows combining vague constraints in expressive logical queries to retrieve graded sets of genes and their associated variants. Results can be further refined in a visual way by…
Authors not listed
In recent years, the development of large language models (LLMs) has revolutionized various fields of natural science, yet their application in molecular data processing remains constrained due to the reliance on single-modality inputs and outputs. To bridge the gap between experimenters and computational tools, we…
Clark C. Evans, Kyrylo Simonov
A new way to conceptualize computations, Query Combinators, can be used to create a data processing environment shared among the entire medical research team. For a given research context, a domain specific query language can be created that represents data sources, analysis methods, and integrative domain knowlege.…
Flore N’kam Suguem, Sébastien Déjean, Philippe Saint Pierre, Nicolas Savy
One of the challenges encountered when merging heterogeneous observational clinical datasets is the recoding of categorical target variables that may have been measured differently across data sources. Standard machine learning-based approaches, such as Multiple Imputation by Chained Equations and the k-Nearest…