15 papers · ranked by Valyu relevance
Authors not listed
Mass spectrometry (MS) generates large datasets that are stored in increasingly optimized and complex file types, demanding technical expertise to extract information rapidly and easily. We wondered whether a simple structured query language (SQL) database could hold raw MS data and allow for easily readable queries…
Authors not listed
In recent years, the development of large language models (LLMs) has revolutionized various fields of natural science, yet their application in molecular data processing remains constrained due to the reliance on single-modality inputs and outputs. To bridge the gap between experimenters and computational tools, we…
Michael Statt, Kristopher Brown, Santosh Suram, Linda Hung + 3 more
In this work, we present DBgen, a Python library that provides a framework for defining extract-transform-load (ETL) pipelines to create and populate SQL databases. DBgen is most useful when the underlying data has complex relationships, requires multi-step analysis, is large-scale, and the type of data being collected…
Tiqing Liu, Linda Hwang, Stephen K Burley, Carmen I Nitsche + 3 more
BindingDB (bindingdb.org) is a public, web-accessible database of experimentally measured binding affinities between small molecules and proteins, which supports diverse applications including medicinal chemistry, biochemical pathway annotation, training of artificial intelligence models, and computational chemistry…
Authors not listed
Natural products are outstanding resources of bioactive compounds with potential applications not only in drug discovery but also in the cosmetic industry and natural pesticides. Costa Rica is among the most biologically diverse countries in terms of the number of known species per unit of area, even above…
José J. Naveja-Romero, Fernanda I. Saldívar-González, Diana L. Prado-Romero, Angel J. Ruiz-Moreno + 3 more
The manuscript discusses recent advances on computer-aided drug discovery (CADD) with focus on data-dependent drug discovery. Herein, we do not intend to review the many CADD methodologies comprehensively. Instead, the review discusses progress on selected concepts, methodologies, resources, and applications that are…
Authors not listed
The field of computational chemistry is increasingly leveraging machine learning (ML) potentials to predict molecular properties with high accuracy and efficiency, providing a viable alternative to traditional quantum mechanical (QM) methods, which are often computationally intensive. Central to the success of ML…
Venkata Chandrasekhar Nainala, Kohulan Rajan, Sri Ram Sagar Kanakam, Nisha Sharma + 3 more
The COCONUT (COlleCtion of Open Natural prodUcTs) database was launched in 2021 as an aggregation of openly available natural product datasets and has been one of the biggest open natural product databases since. Apart from the chemical structures of natural products, COCONUT contains information about names and…
Authors not listed
Sharing knowledge on chemicals in the digital age has been the playground of databases such as the Chemical Abstract Services and PubChem. Wikipedia complements this field by providing context to chemicals aimed at a broad audience, but is not easily read by machines. Wikidata was started as a database service to…
Michael Statt, Brian Rohr, Dan Guevarra, Ja'Nya Breeden + 2 more
Materials knowledge is inherently hierarchical. While high-level descriptors such as composition and structure are valuable for contextualizing materials data, the data must ultimately be considered in the context of its low-level acquisition details. Graph databases offer an opportunity to represent hierarchical…
Evan Walter Clark Spotte-Smith, Orion Archer Cohen, Samuel Blau, Jason Munro + 8 more
Advanced chemical research is increasingly reliant on large computed datasets of molecules and reactions to discover new functional molecules, understand chemicaltrends,train machine learning models, and more. To be of greatest use to the scientific community, such datasets should follow FAIR principles (i.e. be…
Christopher Southan
This article assesses a key aspect of data sharing that has the potential to accelerate the progress and impact of medicinal chemistry. To achieve this the community needs to increase the outward flow of experimental results locked-up in millions of published PDFs into structured open databases that explicitly capture…
Charlotte Neidiger, Tarek Saier, Kai Kühn, Victor Larignon + 12 more
In this work, a concept for an open chemistry knowledge base was developed to integrate chemical research results into a collaboratively usable platform. To achieve this, we enhanced Semantic MediaWiki (SMW) to support the collection and structured summary of chemical data contained in publications. We implemented…
Alejandro Gómez-García, Ann-Kathrin Prinz, Daniel A. Acuña Jiménez, William J. Zamora + 14 more
Compound databases of natural products play a crucial role in drug discovery and development projects and have implications in other areas, such as food chemical research, ecology and metabolomics. Recently, we put together the first version of the Latin American Natural Product database (LANaPDB) as a collective…
HANIYEH ABDOLLAHZADEH, Tonya Peeples, Mohammad Shahcheraghi
DNA-based nanomaterials have shown great potential in numerous applications, thanks to their unique properties including DNA's various molecular interactions, programmability, and versatility with biological modules. Meanwhile, the DNA origami platforms have shown promise in the creation of drug carriers. This…