24 papers · ranked by Valyu relevance
Rebekah Duke, Vinayak Bhat, Chad Risko
As buzzwords like “big data,” “machine learning,” and “high-throughput” expand through chemistry, chemists need to consider more than ever their data storage, data management, and data accessibility, whether in their own laboratories or with the broader community. While it is commonplace for chemists to use…
Qin Wang, Youhuan Li, Yansong Feng, Si Chen + 11 more
'Zhichao Shi' 'Yuequn Dou' 'chuchu Gao' 'Zebin Huang' 'Zihui Si' 'Yixuan Chen' 'Zhaohai Sun' 'Ke Tang' 'Wenqiang Jin'] The relational database design would output a schema based on user's requirements, which defines table structures and their interrelated relations. Translating requirements into accurate schema…
Shi Heng Zhang, Zhengjie Miao, Jiannan Wang
Logical database design has traditionally optimized database schemas, including tables, columns, keys, constraints, and views, for correctness, integrity, and human-written application queries. LLM-based Text-to-SQL changes the consumer: the schema is now often read as text by a language model, so design choices that…
Chuhao Zhou, Shuren Guo, Dong Xiang, Huatang Cao + 4 more
Casting process design is crucial in manufacturing; however, traditional design workflows are time-consuming and seriously reliant on the experience and expertise of designers. To overcome these challenges, database technology has emerged as a promising solution to optimize the design process and enhance efficiency.…
Chen Chen, Yuanyuan Liu, Lei Wang, Jingyi Sai + 8 more
With the rapid accumulation of diverse omics datasets, achieving efficient management and integrative analysis of plant multi-omics data remains a major challenge. Conventional solutions rely on constructing web-based databases, which often demand substantial programming expertise and long-term financial support. To…
Keith J. Fraga, Yuanpeng J. Huang, Theresa A. Ramelot, G.V.T. Swapna + 4 more
NMR is a valuable experimental tool in the structural biologist’s toolkit to elucidate the structures, functions, and motions of biomolecules. The progress of machine learning, particularly in structural biology, reveals the critical importance of large, diverse, and reliable datasets in developing new methods and…
Sherri Weitl-Harms
This paper describes a service learning project used in an upper-level/graduate-level database systems course. Students complete a small database project for a real client. The final product must match the client specification and needs, and include the database design and the final working database system with…
Baldeep Singh, Randall Martyr, Thomas Medland, Jamie Astin + 2 more
'Gordon Hunter' 'Jean-Christophe Nebel'] About fifty years ago, the world’s first fully automated system for trading securities was introduced by Instinet in the US. Since then the world of trading has been revolutionised by the introduction of electronic markets and automatic order execution. Nowadays, financial…
Malcolm Crowe, Fritz Laux
– This paper reviews suggestions for changes to database technology coming from the work of many researchers, particularly those working with evolving big data. We discuss new approaches to remote data access and standards that better provide for durability and auditability in settings including business and scientific…
Authors not listed
Mass spectrometry (MS) generates large datasets that are stored in increasingly optimized and complex file types, demanding technical expertise to extract information rapidly and easily. We wondered whether a simple structured query language (SQL) database could hold raw MS data and allow for easily readable queries…
Daniel Walke, Daniel Micheel, Kay Schallert, Thilo Muth + 3 more
'David Broneske' 'Gunter Saake' 'Robert Heyer'] Title: Abstract The increasing amount and complexity of clinical data require an appropriate way of storing and analyzing those data. Traditional approaches use a tabular structure (relational databases) for storing data and thereby complicate storing and retrieving…
Mateusz Jundzill, Riccardo Spott, Mara Lohde, Martin Hölzer + 2 more
'Adrian Viehweger' 'Christian Brandt'] Title: Abstract With the rapidly growing amount of biological data, powerful but also flexible data management and visualization systems are of increasingly crucial importance. The COVID-19 pandemic has more than highlighted this need and the challenges scientists are facing.…
Aristide Grange
database itself Authors: ['Aristide Grange'] SQL adventure builder (SQLab) is an open-source framework for creating SQL games that are embedded within the very database they query. Students' answers are evaluated using query fingerprinting, a novel technique that allows for better feedback than traditional SQL online…
Marta Chronowska, Michael J. Stam, Derek N. Woolfson, Luigi F. Di Constanzo + 1 more
The field of protein design has changed dramatically over the last 40 years, with a range of methods developing from rational design to more recent data-driven approaches. While considerable insight could be gained from analysing designed proteins, there is no single resource that brings together all the relevant data…
Shunfan Zheng, Dongsheng Shi, Yue Li, Xin Yi + 2 more
Large Language Models (LLMs) are transforming database interaction paradigms, evolving from simple query translators to autonomous database administrators (DBAs). However, current evaluation benchmarks remain disproportionately fixated on Text-to-SQL tasks, neglecting the holistic Database Lifecycle-from initial schema…
Authors not listed
In recent years, the development of large language models (LLMs) has revolutionized various fields of natural science, yet their application in molecular data processing remains constrained due to the reliance on single-modality inputs and outputs. To bridge the gap between experimenters and computational tools, we…
Flore N’kam Suguem, Sébastien Déjean, Philippe Saint Pierre, Nicolas Savy
One of the challenges encountered when merging heterogeneous observational clinical datasets is the recoding of categorical target variables that may have been measured differently across data sources. Standard machine learning-based approaches, such as Multiple Imputation by Chained Equations and the k-Nearest…
Michael Statt, Brian Rohr, Dan Guevarra, Ja'Nya Breeden + 2 more
Materials knowledge is inherently hierarchical. While high-level descriptors such as composition and structure are valuable for contextualizing materials data, the data must ultimately be considered in the context of its low-level acquisition details. Graph databases offer an opportunity to represent hierarchical…
Sumesh Kumar, Joseph Zambreno, Ashfaq Khokhar, Shoaib Akram + 1 more
Improving the speed and efficiency of database search algorithms that deduce peptides from mass spectrometry (MS) data has been an active area of research for more than three decades. The significance of the need for faster database search methods has rapidly increased due to the growing interest in studying non-model…
Tobias B. Alter, Pascal A. Pieters, Colton J. Lloyd, Adam M. Feist + 3 more
Whole-cell biocatalysis facilitates the production of a wide range of industrially and pharmaceutically relevant molecules from sustainable feedstocks such as plastic wastes, carbon dioxide, lignocellulose, or plant-based sugar sources. The identification and use of efficient enzymes in the applied biocatalyst is key…
Authors not listed
In the past decade, many approaches have been suggested to execute ML workloads on a DBMS. However, most of them have looked at in-DBMS ML from a training perspective, whereas ML inference has been largely overlooked. We think that this is an important gap to fill for two main reasons: (1) in the near future, every…
José J. Naveja-Romero, Fernanda I. Saldívar-González, Diana L. Prado-Romero, Angel J. Ruiz-Moreno + 3 more
The manuscript discusses recent advances on computer-aided drug discovery (CADD) with focus on data-dependent drug discovery. Herein, we do not intend to review the many CADD methodologies comprehensively. Instead, the review discusses progress on selected concepts, methodologies, resources, and applications that are…
Tiqing Liu, Linda Hwang, Stephen K Burley, Carmen I Nitsche + 3 more
BindingDB (bindingdb.org) is a public, web-accessible database of experimentally measured binding affinities between small molecules and proteins, which supports diverse applications including medicinal chemistry, biochemical pathway annotation, training of artificial intelligence models, and computational chemistry…
Diana L. Prado-Romero, Fernanda I. Saldívar-González, Iván López-Mata, Pedro A. Laurel-García + 2 more
Designing and developing inhibitors against the epigenetic target DNA methyltransferase (DNMT) is an attractive strategy in epigenetic drug discovery. DNMT1 is one of the epigenetic enzymes with significant clinical relevance. Structure-based de novo design is a drug discovery strategy used in combination with…