26 papers · ranked by Valyu relevance
Peter Sestoft
Research relies on ever larger amounts of data from experiments, automated production equipment, questionnaries, times series such as weather records, and so on. A major task in science is to combine, process and analyse such data to obtain evidence of patterns and correlations. Most research data are on digital form…
Rebekah Duke, Vinayak Bhat, Chad Risko
As buzzwords like “big data,” “machine learning,” and “high-throughput” expand through chemistry, chemists need to consider more than ever their data storage, data management, and data accessibility, whether in their own laboratories or with the broader community. While it is commonplace for chemists to use…
Mateusz Jundzill, Riccardo Spott, Mara Lohde, Martin Hölzer + 2 more
'Adrian Viehweger' 'Christian Brandt'] Title: Abstract With the rapidly growing amount of biological data, powerful but also flexible data management and visualization systems are of increasingly crucial importance. The COVID-19 pandemic has more than highlighted this need and the challenges scientists are facing.…
Aleem Akhtar
—Databases are considered to be integral part of modern information systems. Almost every web or mobile application uses some kind of database. Database management systems are considered to be a crucial element from both business and technological standpoint. This paper divides different types of database management…
Jasper Kyle Catapang
Databases play an essential role in our society today. Databases are embedded in sectors like corporations, institutions, and government organizations, among others. These databases are used for our video and audio streaming platforms, social gaming, finances, cloud storage, e-commerce, healthcare, economy, etc. It is…
Mohamed Hassan
In the world of data management, the rapidly growing area of big data offers both intriguing possibilities and substantial obstacles. Big data, defined by its enormous size, complex structure, and varied sources - which include everything from social media engagement and sensor data to financial transactions and…
Shahid Ullah, Wajeeha Rahman, Farhan Ullah, Gulzar Ahmad + 2 more
'Muhmmad Ijaz' 'Tianshun Gao'] Background: The achievement of the human genome project provides a basis for the systematic study of the human genome from evolutionary history to disease-specific medicine. With the explosive growth of biological data, a growing number of biological databases are being established to…
Paul L. Soto
Data collection and analysis are central to scientific research, including in applied and basic behavior analysis. A substantial amount of attention has been given to how to rigorously collect and analyze data. Less attention has been paid to storing and maintaining research data, which becomes a critical step in the…
Heidi J. Imker
Online resources enable unfettered access to and analysis of scientific data and are considered crucial for the advancement of modern science. Despite the clear power of online data resources, including web-available databases, proliferation can be problematic due to challenges in sustainability and long-term…
Dimitri Yatsenko, Jacob Reimer, Alexander S. Ecker, Edgar Y. Walker + 6 more
The rise of big data in modern research poses serious challenges for data management: Large and intricate datasets from diverse instrumentation must be precisely aligned, annotated, and processed in a variety of ways to extract new insights. While high levels of data integrity are expected, research teams have diverse…
Rukshan Alexander, Prashanthi Rukshan, Sinnathamby Mahesan
It is a long term desire of the computer users to minimize the communication gap between the computer and a human. On the other hand, almost all ICT applications store information in to databases and retrieve from them. Retrieving information from the database requires knowledge of technical languages such as…
Bohdan B. Khomtchouk, Kasra A. Vand, Thor Wahlestedt, Kelly Khomtchouk + 2 more
We propose a search engine and file retrieval system for all bioinformatics databases worldwide. PubData searches biomedical data in a user-friendly fashion similar to how PubMed searches biomedical literature. PubData is built on novel network programming, natural language processing, and artificial intelligence…
Renzo Angles, Claudio Gutiérrez
A graph database is a database where the data structures for the schema and/or instances are modeled as a (labeled)(directed) graph or generalizations of it, and where querying is expressed by graphoriented operations and type constructors. In this article we present the basic notions of graph databases, give an…
Chaimae Asaad, Karim Bäına, Mounir Ghogho
The demanding requirements of the new Big Data intensive era raised the need for flexible storage systems capable of handling huge volumes of unstructured data and of tackling the challenges that traditional databases were facing. NoSQL Databases, in their heterogeneity, are a powerful and diverse set of databases…
Authors not listed
Mass spectrometry (MS) generates large datasets that are stored in increasingly optimized and complex file types, demanding technical expertise to extract information rapidly and easily. We wondered whether a simple structured query language (SQL) database could hold raw MS data and allow for easily readable queries…
Authors not listed
In recent years, the development of large language models (LLMs) has revolutionized various fields of natural science, yet their application in molecular data processing remains constrained due to the reliance on single-modality inputs and outputs. To bridge the gap between experimenters and computational tools, we…
Tiqing Liu, Linda Hwang, Stephen K Burley, Carmen I Nitsche + 3 more
BindingDB (bindingdb.org) is a public, web-accessible database of experimentally measured binding affinities between small molecules and proteins, which supports diverse applications including medicinal chemistry, biochemical pathway annotation, training of artificial intelligence models, and computational chemistry…
Xiao-Qin Xia, Michael McClelland, Yipeng Wang
Background With advances in high-throughput genomics and proteomics, it is challenging for biologists to deal with large data files and to map their data to annotations in public databases. Results We developed TabSQL, a MySQL-based application tool, for viewing, filtering and querying data files with large numbers of…
Michael Statt, Kristopher Brown, Santosh Suram, Linda Hung + 3 more
In this work, we present DBgen, a Python library that provides a framework for defining extract-transform-load (ETL) pipelines to create and populate SQL databases. DBgen is most useful when the underlying data has complex relationships, requires multi-step analysis, is large-scale, and the type of data being collected…
Prasetya Ajie Utama, Nathaniel Weir, Fuat Basık, Carsten Binnig + 5 more
'Uğur Çetintemel' 'Benjamin Hättasch' 'Amir Ilkhechi' 'Shekar Ramaswamy' 'Arif Usta'] The ability to extract insights from new data sets is critical for decision making. Visual interactive tools play an important role in data exploration since they provide non-technical users with an effective way to visually compose…
Magali Ruffier, Andreas Kähäri, Monika Komorowska, Stephen Keenan + 10 more
The Ensembl software resources are a stable infrastructure to store, access and manipulate genome assemblies and their functional annotations. The Ensembl “Core” database and Application Programming Interface (API) was our first major piece of software infrastructure and remains at the centre of all of our genome…
Christopher Southan
This article assesses a key aspect of data sharing that has the potential to accelerate the progress and impact of medicinal chemistry. To achieve this the community needs to increase the outward flow of experimental results locked-up in millions of published PDFs into structured open databases that explicitly capture…
Alexander M. Waldrop, John B. Cheadle, Kira Bradford, Alexander Preiss + 14 more
As the number of public data resources continues to proliferate, identifying relevant datasets across heterogenous repositories is becoming critical to answering scientific questions. To help researchers navigate this data landscape, we developed Dug: a semantic search tool for biomedical datasets utilizing…
Michael Statt, Brian Rohr, Dan Guevarra, Ja'Nya Breeden + 2 more
Materials knowledge is inherently hierarchical. While high-level descriptors such as composition and structure are valuable for contextualizing materials data, the data must ultimately be considered in the context of its low-level acquisition details. Graph databases offer an opportunity to represent hierarchical…
Alejandro Gómez-García, Ann-Kathrin Prinz, Daniel A. Acuña Jiménez, William J. Zamora + 14 more
Compound databases of natural products play a crucial role in drug discovery and development projects and have implications in other areas, such as food chemical research, ecology and metabolomics. Recently, we put together the first version of the Latin American Natural Product database (LANaPDB) as a collective…
Gerardo Lagunes-García, Alejandro Rodríguez-González, Lucía Prieto-Santamaría, Eduardo P. García del Valle + 2 more
Within the global endeavour of improving population health, one major challenge is the increasingly high cost associated with drug development. Drug repositioning, i.e. finding new uses for existing drugs, is a promising alternative; yet, its effectiveness has hitherto been hindered by our limited knowledge about…