26 papers · ranked by Valyu relevance
Heidi J. Imker
For decades, life science researchers have had cost-free, unrestricted access to data through online databases. However, the sustainability of even well-established resources was already tenuous, and abrupt changes in science funding in the United States seems poised to exacerbate these challenges. This study employed…
Authors not listed
Mass spectrometry (MS) generates large datasets that are stored in increasingly optimized and complex file types, demanding technical expertise to extract information rapidly and easily. We wondered whether a simple structured query language (SQL) database could hold raw MS data and allow for easily readable queries…
Diana Martínez-Minguet, René Noel, Alberto García S., Mireia Costa + 1 more
Research into the genetics of autism spectrum disorder (ASD) seeks to unravel its complex genetic background by identifying genes associated with the condition at varying levels of confidence. While these findings hold significant potential for clinical applications, the dispersed nature of scientific evidence presents…
Saikrishna Sudarshan, Tanay Kulkarni, Manasi Patwardhan, Lovekesh Vig + 2 more
—We address the task of routing natural language queries in multi-database enterprise environments. We construct realistic benchmarks by extending existing NL-to-SQL datasets. Our study shows that routing becomes increasingly challenging with larger, domain-overlapping DB repositories and ambiguous queries, motivating…
Mingrui Liu, Zelin Ye, Haiyu Liu, Pengzhen Ma + 6 more
This review provides a representative overview of public biomedical databases and their use in biomedical research. These resources are categorized into four major types according to their dominant data content: public health databases, clinical databases, comprehensive cohort databases, and omics databases. For each…
Kim Boesen, Lars G Hemkens, Perrine Janiaud, Julian Hirt
The seven COVID-19 databases were initiated shortly after the SARS-CoV-2 outbreak in early 2020. By January 2024, the last of these databases, the Cochrane COVID-19 Study Register had ceased updates. All databases were publicly funded or supported by non-profit institutions and were openly accessible to the public…
Shunfan Zheng, Dongsheng Shi, Yue Li, Xin Yi + 2 more
Large Language Models (LLMs) are transforming database interaction paradigms, evolving from simple query translators to autonomous database administrators (DBAs). However, current evaluation benchmarks remain disproportionately fixated on Text-to-SQL tasks, neglecting the holistic Database Lifecycle-from initial schema…
Chen Chen, Yuanyuan Liu, Lei Wang, Jingyi Sai + 8 more
With the rapid accumulation of diverse omics datasets, achieving efficient management and integrative analysis of plant multi-omics data remains a major challenge. Conventional solutions rely on constructing web-based databases, which often demand substantial programming expertise and long-term financial support. To…
Allan Garland, Peter Dodek, Kednapa Thavorn, Rita Wissa + 3 more
Introduction While Canada is rich in databases useful to support healthcare research, they are widely distributed, often poorly documented, and it is challenging to identify relevant databases, apply for access, and eventually use, link or harmonise the data. Even if the databases needed to address specific questions…
Authors not listed
Chemical reaction databases have become core scientific infrastructure. Most prominent datasets focus on or- ganic reactions, or only include reactants and product rather than full reaction pathways, leaving organometallic chemistry particularly underserved despite its centrality to homogeneous catalysis. This gap…
Alejandro Roldán, Tomás Golomb Durán, Antoni Josep Far, Maria Capa + 2 more
The era of Big Data has revolutionised biodiversity research, yet the potential of this information is frequently constrained by data heterogeneity, incompatible schemas, and the fragmentation of resources. Whilst standards such as Darwin Core have improved interoperability, significant barriers persist in harmonising…
Yifei Jiang, Xiaozhuan Jia, Zhenzhong Yang
Medicinal plants have long served as an important asset in the treatment of diseases. Recent developments in computer science have enabled the rise of specialized databases cataloging medicinal plant knowledge. However, a systematic comparison of available region-specific medicinal plant databases is lacking. This…
Marina Martínez de Pinillos González, Ana Álvarez Fernández, Beatriz Delgado Esteban, Miguel Delgado + 13 more
Human dental morphology is diverse and varies both within and between populations worldwide. Common variants include different numbers of cusps and roots, as well as different configurations in the fissures, ridges, and grooves on tooth crowns. Because teeth preserve well in taphonomic contexts and retain strong…
Miklós Bán
The increasing reliance of biodiversity research on large-scale databases has brought significant progress in data accessibility but also new challenges in data comparability, reliability, and interpretability. While global platforms, such as GBIF and iNaturalist standardize and disseminate vast quantities of…
Abhishek Halder, Manvendra Singh, Rohit Kesarwani, Bernadette Mathew + 11 more
Biomedical research increasingly relies on expert-curated databases to connect diseases, genes, variants, phenotypes, pathways and therapeutics. Incomplete or irreproducible retrieval can distort interpretation, mechanistic inference, variant assessment or therapeutic prioritization despite an unchanged evidence base.…
Denis Mayr Lima Martins, Gottfried Vossen
Self-Organizing Maps (SOMs) have long been used as exploratory tools for high-dimensional data: they organize objects into a two-dimensional topology that reveals clusters, gradients, sparse regions, dense regions, and boundaries. Yet, in modern data systems, SOMs are typically trained and visualized outside the DBMS…
Flore N’kam Suguem, Sébastien Déjean, Philippe Saint Pierre, Nicolas Savy
One of the challenges encountered when merging heterogeneous observational clinical datasets is the recoding of categorical target variables that may have been measured differently across data sources. Standard machine learning-based approaches, such as Multiple Imputation by Chained Equations and the k-Nearest…
Yizhu Jiao, Sha Li, Sizhe Zhou, Heng Ji + 1 more
The task of information extraction (IE) is to extract structured knowledge from text. However, it is often not straightforward to utilize IE output due to the mismatch between the IE ontology and the downstream application needs. We propose a new formulation of IE TEXT2DB that emphasizes the integration of IE output…
Hassan Soliman, Vivek Gupta, Dan Roth, Iryna Gurevych
Realistic text-to-SQL workflows often require joining multiple tables. As a result, accurately retrieving the relevant set of tables becomes a key bottleneck for end-to-end performance. We study an open-book setting where queries must be answered over large, heterogeneous table collections pooled from many sources…
Hongshen Gou, Feng Tian, Long Wang, Nan Deng + 1 more
The rapid advancement of artificial intelligence has elevated data to a cornerstone of modern software systems. As data projects become increasingly complex and dynamic, version control for data has become essential rather than merely convenient. Existing version control systems designed for source code are inadequate…
Authors not listed
The materials-science literature is the richest reservoir of domain knowledge, yet converting its unstructured text—especially narrative passages and complex tables—into machine-readable data for analysis and ML model training remains challenging. To address this, we present KnowMat, an agentic, multi-stage pipeline…
Kazi F. Akhter, Bharath Ajendla, Manar D. Samad
Relational databases (RDBs) are the primary data infrastructure in many enterprises, yet recent deep learning methods designed for RDBs have been evaluated under inconsistent experimental protocols, making fair comparison difficult. We present one of the first systematic benchmarking studies of recently released deep…
Authors not listed
Machine learning is increasingly used to predict reaction properties such as barrier heights, reaction energies, rates, or yields, as well as the underlying molecular geometries, including transition state structures. While such predictions have the potential to provide mechanistic insight for high-impact applications…
Authors not listed
Molecular mechanisms governing initiation steps of the assembly of thousands of endogenous multi-protein complexes (EMCs) remain incompletely understood. Here, multiple lines of observations are reported reflecting the biological functions-aligned initiation sequence of hybrid assembly pathways (HAPs) of EMCs. HAPs…
Authors not listed
The discovery of chemically novel or structurally anomalous metal-organic frameworks (MOFs) is essential for expanding reticular design space and enhancing dataset reliability. We present CHEM-AD (Chemically Unusual Metal–organic Frameworks via Autoencoder-based Detection), a label-free, CPU-efficient pipeline that…
Authors not listed
DNA-encoded libraries (DELs) have emerged as a powerful platform for screening ultra-large chemical spaces by leveraging DNA barcodes to tag and track individual small molecules. Recent work has shown that machine learning can enhance DEL based hit discovery by denoising sequencing artifacts and improving binder…