15 papers · ranked by Valyu relevance
Authors not listed
High-quality data preprocessing is essential for untargeted metabolomics experiments, where increasing dataset scale and complexity demand adaptable, robust, and reproducible software solutions. Modern preprocessing tools must evolve to integrate seamlessly with downstream analysis platforms, ensuring efficient and…
Authors not listed
The precision of thermodynamic modeling for ionic liquid (IL)–solute systems is fundamentally reliant on the quality of experimental data. However, prevalent databases such as ILThermo frequently exhibit conflicting measurements for the same systems under identical temperature and pressure conditions. These disparities…
Helle W. van den Maagdenberg, Martin Šícho, David Alencar Araripe, Sohvi Luukkonen + 9 more
Building reliable and robust quantitative structure-property relationship (QSPR) models is a challenging task. First, the experimental data needs to be obtained, analyzed and curated. Second, the number of available methods is continuously growing and evaluating different algorithms and methodologies can be arduous.…
Authors not listed
The integration of artificial intelligence technologies into pharmaceutical research is crucial for gaining an early understanding of molecular properties, thereby facilitating successful drug design. Constructing a machine learning (ML) model however, requires knowledge spanning from data preprocessing and feature…
Authors not listed
The discovery of chemically novel or structurally anomalous metal-organic frameworks (MOFs) is essential for expanding reticular design space and enhancing dataset reliability. We present CHEM-AD (Chemically Unusual Metal–organic Frameworks via Autoencoder-based Detection), a label-free, CPU-efficient pipeline that…
Authors not listed
Methanol synthesis from syngas (CO/CO₂/H₂) is vital for sustainable chemical production; however, traditional kinetic models hinder rapid reactor optimisation [1]. We present a reproducible machine-learning pipeline to predict methanol yield in a double-pass plug-flow reactor, utilising a synthetic dataset (n = 5,000)…
Authors not listed
In recent years, the development of large language models (LLMs) has revolutionized various fields of natural science, yet their application in molecular data processing remains constrained due to the reliance on single-modality inputs and outputs. To bridge the gap between experimenters and computational tools, we…
Authors not listed
This study presents a validation and refinement of the “yellow cards” error detection workflow that can be applied to any property connected to molecular structure. In our implementation the workflow employed 5 predictive models with each assigning a “yellow card” to 5% of the entries with worst prediction accuracy.…
Tobias Rehfeldt, Ralf Gabriels, Robbin Bouwmeester, Siegfried Gessulat + 6 more
Dataset acquisition and curation are often the hardest and most time-consuming parts of a machine learning endeavor. This is especially true for proteomics-based LC-IM-MS datasets, due to the high-throughput data structure with high levels of noise and complexity between raw and machine learning-ready formats. While…
Authors not listed
DNA-encoded libraries (DELs) are a powerful way to find chemical starting points against challenging biological targets, by rapidly generating billion-scale structure-activity datasets. However, DEL experiment design and interpretation, especially the optimal use of machine learning (ML) to analyse the vast amount of…
Feng Feng, Zhenru Chen, Jianyuan Ni, Yuanxun Zhang + 3 more
Drinking water is essential to public health and socioeconomic growth. Therefore, assessing and ensuring drinking water supply is a critical task in modern society. Conventional approaches to analyzing and controlling drinking water quality are labor-intensive and costly with a low throughput. Machine learning (ML) is…
Mahnoor Zulfiqar, Michael R. Crusoe, Birgitta König-Ries, Christoph Steinbeck + 2 more
Scientific workflows facilitate the automation of data analysis tasks by integrating various software and tools executed in a particular order. To enable transparency and reusability in workflows, it is essential to implement the FAIR principles. Here, we describe our experiences implementing the FAIR principles for…
Authors not listed
Kinetic modeling is essential for predicting changes in food quality during processing and storage. This study evaluates the application of physics-informed neural networks (PINN) for food kinetic modeling, integrating kinetic insights into neural network frameworks. Based on three case studies, namely seed drying…
Michael Statt, Kristopher Brown, Santosh Suram, Linda Hung + 3 more
In this work, we present DBgen, a Python library that provides a framework for defining extract-transform-load (ETL) pipelines to create and populate SQL databases. DBgen is most useful when the underlying data has complex relationships, requires multi-step analysis, is large-scale, and the type of data being collected…
Aaron Liu, Myeongyeon Lee, Rahul Venkatesh, Jessica Bonsu + 4 more
Polymer-based semiconductors and organic electronics encapsulate a significant research thrust for informatics-driven materials development. However, device measurements are described by a complex array of design and parameter choices, many of which are sparsely reported. For example, the mobility of a polymer-based…