15 papers · ranked by Valyu relevance
Naoual El aboudi, Laila Benhlima
The growing amount of data in healthcare industry has made inevitable the adoption of big data techniques in order to improve the quality of healthcare delivery. Despite the integration of big data processing approaches and platforms in existing data management architectures for healthcare systems, these architectures…
Lorraine Buis, Richard Matovu, Lei Guo, Bengie L Ortiz + 8 more
Background Wearable sensors are increasingly being explored in health care, including in cancer care, for their potential in continuously monitoring patients. Despite their growing adoption, significant challenges remain in the quality and consistency of data collected from wearable sensors. Moreover, preprocessing…
Ioannis K. Gallos, Dimitrios Tryfonopoulos, Gidi Shani, Angelos Amditis + 4 more
Early detection of colorectal cancer is crucial for improving outcomes and reducing mortality. While there is strong evidence of effectiveness, currently adopted screening methods present several shortcomings which negatively impact the detection of early stage carcinogenesis, including low uptake due to patient…
Davide Chicco, Luca Oneto, Erica Tavazzi, Francis Ouellette
Applying computational statistics or machine learning methods to data is a key component of many scientific studies, in any field, but alone might not be sufficient to generate robust and reliable outcomes and results. Before applying any discovery method, preprocessing steps are necessary to prepare the data to the…
Joseph C. Mellor, Michael A. Stone, John Keane
Principles and Potential Authors: ['Joseph C. Mellor' 'Michael A. Stone' 'John Keane'] The ubiquity and cheapness of miniature low-power sensors, digital processing, and large amounts of storage contained in small packages has heralded the ability to acquire large amounts of data about systems during their course of…
Authors not listed
The precision of thermodynamic modeling for ionic liquid (IL)–solute systems is fundamentally reliant on the quality of experimental data. However, prevalent databases such as ILThermo frequently exhibit conflicting measurements for the same systems under identical temperature and pressure conditions. These disparities…
S. M. Kamruzzaman, A. M. Jehad Sarkar
Classification is one of the data mining problems receiving enormous attention in the database community. Although artificial neural networks (ANNs) have been successfully applied in a wide range of machine learning applications, they are however often regarded as black boxes, i.e., their predictions cannot be…
Authors not listed
High-quality data preprocessing is essential for untargeted metabolomics experiments, where increasing dataset scale and complexity demand adaptable, robust, and reproducible software solutions. Modern preprocessing tools must evolve to integrate seamlessly with downstream analysis platforms, ensuring efficient and…
Amaryllis Mavragani, Rohan Alexander, Hyo Jung Kim, Manping Guo + 11 more
'Yiming Wang' 'Qiaoning Yang' 'Rui Li' 'Yang Zhao' 'Chenfei Li' 'Mingbo Zhu' 'Yao Cui' 'Xin Jiang' 'Song Sheng' 'Qingna Li' 'Rui Gao'] With the rapid development of science, technology, and engineering, large amounts of data have been generated in many fields in the past 20 years. In the process of medical research…
Authors not listed
In recent years, the development of large language models (LLMs) has revolutionized various fields of natural science, yet their application in molecular data processing remains constrained due to the reliance on single-modality inputs and outputs. To bridge the gap between experimenters and computational tools, we…
Authors not listed
This study presents a validation and refinement of the “yellow cards” error detection workflow that can be applied to any property connected to molecular structure. In our implementation the workflow employed 5 predictive models with each assigning a “yellow card” to 5% of the entries with worst prediction accuracy.…
Feng Feng, Zhenru Chen, Jianyuan Ni, Yuanxun Zhang + 3 more
Drinking water is essential to public health and socioeconomic growth. Therefore, assessing and ensuring drinking water supply is a critical task in modern society. Conventional approaches to analyzing and controlling drinking water quality are labor-intensive and costly with a low throughput. Machine learning (ML) is…
Tobias Rehfeldt, Ralf Gabriels, Robbin Bouwmeester, Siegfried Gessulat + 6 more
Dataset acquisition and curation are often the hardest and most time-consuming parts of a machine learning endeavor. This is especially true for proteomics-based LC-IM-MS datasets, due to the high-throughput data structure with high levels of noise and complexity between raw and machine learning-ready formats. While…
Authors not listed
Kinetic modeling is essential for predicting changes in food quality during processing and storage. This study evaluates the application of physics-informed neural networks (PINN) for food kinetic modeling, integrating kinetic insights into neural network frameworks. Based on three case studies, namely seed drying…
Sarah E. Lindley, Yiyang Lu, Diwakar Shukla
Guide to Machine Learning for Small Molecule Design Authors: ['Sarah\nE. Lindley' 'Yiyang Lu' 'Diwakar Shukla'] Initially part of the field of artificial intelligence, machine learning (ML) has become a booming research area since branching out into its own field in the 1990s. After three decades of refinement, ML…