26 papers · ranked by Valyu relevance
Shahid Ullah, Wajeeha Rahman, Farhan Ullah, Gulzar Ahmad + 2 more
'Muhmmad Ijaz' 'Tianshun Gao'] Background: The achievement of the human genome project provides a basis for the systematic study of the human genome from evolutionary history to disease-specific medicine. With the explosive growth of biological data, a growing number of biological databases are being established to…
Rebekah Duke, Vinayak Bhat, Chad Risko
As buzzwords like “big data,” “machine learning,” and “high-throughput” expand through chemistry, chemists need to consider more than ever their data storage, data management, and data accessibility, whether in their own laboratories or with the broader community. While it is commonplace for chemists to use…
Mateusz Jundzill, Riccardo Spott, Mara Lohde, Martin Hölzer + 2 more
'Adrian Viehweger' 'Christian Brandt'] Title: Abstract With the rapidly growing amount of biological data, powerful but also flexible data management and visualization systems are of increasingly crucial importance. The COVID-19 pandemic has more than highlighted this need and the challenges scientists are facing.…
Aleem Akhtar
—Databases are considered to be integral part of modern information systems. Almost every web or mobile application uses some kind of database. Database management systems are considered to be a crucial element from both business and technological standpoint. This paper divides different types of database management…
Mohamed Hassan
In the world of data management, the rapidly growing area of big data offers both intriguing possibilities and substantial obstacles. Big data, defined by its enormous size, complex structure, and varied sources - which include everything from social media engagement and sensor data to financial transactions and…
Baldeep Singh, Randall Martyr, Thomas Medland, Jamie Astin + 2 more
'Gordon Hunter' 'Jean-Christophe Nebel'] About fifty years ago, the world’s first fully automated system for trading securities was introduced by Instinet in the US. Since then the world of trading has been revolutionised by the introduction of electronic markets and automatic order execution. Nowadays, financial…
Paul L. Soto
Data collection and analysis are central to scientific research, including in applied and basic behavior analysis. A substantial amount of attention has been given to how to rigorously collect and analyze data. Less attention has been paid to storing and maintaining research data, which becomes a critical step in the…
Dhouha Grissa, Alexander Junge, Tudor I. Oprea, Lars Juhl Jensen
The scientific knowledge about which genes are involved in which diseases grows rapidly, which makes it difficult to keep up with new publications and genetics datasets. The DISEASES database aims to provide a comprehensive overview by systematically integrating and assigning confidence scores to evidence for…
Authors not listed
Mass spectrometry (MS) generates large datasets that are stored in increasingly optimized and complex file types, demanding technical expertise to extract information rapidly and easily. We wondered whether a simple structured query language (SQL) database could hold raw MS data and allow for easily readable queries…
Heming Zhang, Shunning Liang, Tim Xu, Wenyu Li + 15 more
Artificial intelligence (AI) is revolutionizing scientific discovery because of its super capability, following the neural scaling laws, to integrate and analyze large-scale datasets to mine knowledge. Foundation models, large language models (LLMs) and large vision models (LVMs), are among the most important…
Michael J. Sullivan, Zhibo Chen, Elvis Pranskevichus, Robert J. Simmons + 3 more
For applications that store structured data in relational databases, there is an impedance mismatch between the flat representations encouraged by relational data models and the deeply nested information that applications expect to receive. In this work, we present the graph-relational database model, which provides a…
Zhengtong Yan, Yuan, Gongsheng, Qing-Yi Guo + 1 more
Modern enterprises are increasingly driven by the DATA+AI paradigm, in which Database Management Systems (DBMSs) and Large Language Models (LLMs) have become two foundational infrastructures powering a wide range of industrial and business applications, such as enterprise analytics, intelligent customer service, and…
Authors not listed
In recent years, the development of large language models (LLMs) has revolutionized various fields of natural science, yet their application in molecular data processing remains constrained due to the reliance on single-modality inputs and outputs. To bridge the gap between experimenters and computational tools, we…
João Pedro de Magalhães, Zoya Abidi, Gabriel Arantes dos Santos, Roberto A. Avelar + 11 more
Ageing is a complex and multifactorial process. For two decades, the Human Ageing Genomic Resources (HAGR) have aided researchers in the study of various aspects of ageing and its manipulation. Here we present the key features and recent enhancements of these resources, focusing on its six main databases. One database…
Tiqing Liu, Linda Hwang, Stephen K Burley, Carmen I Nitsche + 3 more
BindingDB (bindingdb.org) is a public, web-accessible database of experimentally measured binding affinities between small molecules and proteins, which supports diverse applications including medicinal chemistry, biochemical pathway annotation, training of artificial intelligence models, and computational chemistry…
Eduardo Nascimento, Caio Viktor S. Avila, Yenier Torres Izquierdo, Grettel Monteagudo García + 4 more
'Grettel Monteagudo García' 'Lucas Vinícius Moreira de Andrade' 'Michelle S. P. Facina' 'Melissa Lemos' 'Marco A. Casanova'] Abstract: Text-to-SQL prompt strategies based on Large Language Models (LLMs) achieve remarkable performance on well-known benchmarks. However, when applied to real-world databases, their…
Carla Cicero, Michelle S. Koo, Emily Braker, John Abbott + 15 more
Museum collections house millions of objects and associated data records that document biological and cultural diversity. In recent decades, digitization efforts have greatly increased accessibility to these data, thereby revolutionizing interdisciplinary studies in evolutionary biology, biogeography, epidemiology…
Elizabeth Wenk, Payal Bal, David Coleman, Rachael Gallagher + 2 more
Trait databases have proliferated over the past decades, facilitating research on the ecology, evolution, and conservation of taxa across the Tree of Life. Typically, teams of independent researchers build these databases, and each must develop their own workflow and output structure. This divests research hours from…
José J. Naveja-Romero, Fernanda I. Saldívar-González, Diana L. Prado-Romero, Angel J. Ruiz-Moreno + 3 more
The manuscript discusses recent advances on computer-aided drug discovery (CADD) with focus on data-dependent drug discovery. Herein, we do not intend to review the many CADD methodologies comprehensively. Instead, the review discusses progress on selected concepts, methodologies, resources, and applications that are…
Rania Mkhinini Gahar, Olfa Arfaoui, Minyar Sassi Hidri
—To succeed in a Big Data strategy, you have to arm yourself with a wide range of data skills and best practices. This strategy can result in an impressive asset that can streamline operational costs, reduce time to market, and enable the creation of new products. However, several Big Data challenges may take place in…
Alejandro Roldán, Tomás Golomb Durán, Antoni Josep Far, Maria Capa + 2 more
The era of Big Data has revolutionised biodiversity research, yet the potential of this information is frequently constrained by data heterogeneity, incompatible schemas, and the fragmentation of resources. Whilst standards such as Darwin Core have improved interoperability, significant barriers persist in harmonising…
Sergio Lifschitz, Edward H. Haeusler, Marcos Catanho, Antonio B. de Miranda + 6 more
'Antonio B. de Miranda' 'Elvismary Molina de Armas' 'Alexandre Heine' 'Sergio G. M. P. Moreira' 'Cristian Tristão' 'Pietro Pinoli' 'Anna Bernasconi'] DNA sequencers output a large set of very long biological data strings that we should persist in databases rather than basic text file systems. Many different data models…
Flore N’kam Suguem, Sébastien Déjean, Philippe Saint Pierre, Nicolas Savy
One of the challenges encountered when merging heterogeneous observational clinical datasets is the recoding of categorical target variables that may have been measured differently across data sources. Standard machine learning-based approaches, such as Multiple Imputation by Chained Equations and the k-Nearest…
Michael Statt, Brian Rohr, Dan Guevarra, Ja'Nya Breeden + 2 more
Materials knowledge is inherently hierarchical. While high-level descriptors such as composition and structure are valuable for contextualizing materials data, the data must ultimately be considered in the context of its low-level acquisition details. Graph databases offer an opportunity to represent hierarchical…
Alejandro Gómez-García, Ann-Kathrin Prinz, Daniel A. Acuña Jiménez, William J. Zamora + 14 more
Compound databases of natural products play a crucial role in drug discovery and development projects and have implications in other areas, such as food chemical research, ecology and metabolomics. Recently, we put together the first version of the Latin American Natural Product database (LANaPDB) as a collective…
Changhao Zhu, Junzhe Li, Ziyue Zhong, Cong Yue + 1 more
The success of blockchain technology in cryptocurrencies reveals its potential in the data management field. Recently, there is a trend in the database community to integrate blockchains and traditional databases to obtain security, efficiency, and privacy from the two distinctive but related systems. In this survey…