25 papers · ranked by Valyu relevance
Alhassan Mumuni, Fuseini Mumuni
—Modern approach to artificial intelligence (AI) aims to design algorithms that learn directly from data. This approach has achieved impressive results and has contributed significantly to the progress of AI, particularly in the sphere of supervised deep learning. It has also simplified the design of machine learning…
Canchen Li
—Data mining is about obtaining new knowledge from existing datasets. However, the data in the existing datasets can be scattered, noisy, and even incomplete. Although lots of effort is spent on developing or fine-tuning data mining models to make them more robust to the noise of the input data, their qualities still…
Naoual El aboudi, Laila Benhlima
The growing amount of data in healthcare industry has made inevitable the adoption of big data techniques in order to improve the quality of healthcare delivery. Despite the integration of big data processing approaches and platforms in existing data management architectures for healthcare systems, these architectures…
Lorraine Buis, Richard Matovu, Lei Guo, Bengie L Ortiz + 8 more
Background Wearable sensors are increasingly being explored in health care, including in cancer care, for their potential in continuously monitoring patients. Despite their growing adoption, significant challenges remain in the quality and consistency of data collected from wearable sensors. Moreover, preprocessing…
Lydia R Lucchesi, Petra Kuhnert, Jenny Davis, Lexing Xie
Data preprocessing is a crucial stage in the data analysis pipeline, with both technical and social aspects to consider. Yet, the attention it receives is often lacking in research practice and dissemination. We present the Smallset Timeline, a visualisation to help reflect on and communicate data preprocessing…
Oscar Esteban, Christopher J. Markiewicz, Ross W. Blair, Craig A. Moodie + 12 more
Preprocessing of functional MRI (fMRI) involves numerous steps to clean and standardize data before statistical analysis. Generally, researchers create ad hoc preprocessing workflows for each new dataset, building upon a large inventory of tools available for each step. The complexity of these workflows has snowballed…
Valerie Restat
Data preparation, especially data cleaning, is very important to ensure data quality and to improve the output of automated decision systems. Since there is no single tool that covers all steps required, a combination of tools – namely a data preparation pipeline – is required. Such process comes with a number of…
Ioannis K. Gallos, Dimitrios Tryfonopoulos, Gidi Shani, Angelos Amditis + 4 more
Early detection of colorectal cancer is crucial for improving outcomes and reducing mortality. While there is strong evidence of effectiveness, currently adopted screening methods present several shortcomings which negatively impact the detection of early stage carcinogenesis, including low uptake due to patient…
Davide Chicco, Luca Oneto, Erica Tavazzi, Francis Ouellette
Applying computational statistics or machine learning methods to data is a key component of many scientific studies, in any field, but alone might not be sufficient to generate robust and reliable outcomes and results. Before applying any discovery method, preprocessing steps are necessary to prepare the data to the…
Joseph C. Mellor, Michael A. Stone, John Keane
Principles and Potential Authors: ['Joseph C. Mellor' 'Michael A. Stone' 'John Keane'] The ubiquity and cheapness of miniature low-power sensors, digital processing, and large amounts of storage contained in small packages has heralded the ability to acquire large amounts of data about systems during their course of…
Yousef Koka, David Selby, Gerrit Großmann, Sebastián Vollmer
using reinforcement learning Authors: ['Yousef Koka' 'David Selby' 'Gerrit Großmann' 'Sebastián Vollmer'] Data preprocessing is a critical yet frequently neglected aspect of machine learning, often paid little attention despite its potentially significant impact on model performance. While automated machine learning…
Authors not listed
The precision of thermodynamic modeling for ionic liquid (IL)–solute systems is fundamentally reliant on the quality of experimental data. However, prevalent databases such as ILThermo frequently exhibit conflicting measurements for the same systems under identical temperature and pressure conditions. These disparities…
S. M. Kamruzzaman, A. M. Jehad Sarkar
Classification is one of the data mining problems receiving enormous attention in the database community. Although artificial neural networks (ANNs) have been successfully applied in a wide range of machine learning applications, they are however often regarded as black boxes, i.e., their predictions cannot be…
Authors not listed
High-quality data preprocessing is essential for untargeted metabolomics experiments, where increasing dataset scale and complexity demand adaptable, robust, and reproducible software solutions. Modern preprocessing tools must evolve to integrate seamlessly with downstream analysis platforms, ensuring efficient and…
Greg Finak, Bryan T. Mayer, William Fulp, Paul Obrecht + 4 more
A central tenet of reproducible research is that scientific results are published along with the underlying data and software code necessary to reproduce and verify the findings. A host of tools and software have been released that facilitate such work-flows and scientific journals have increasingly demanded that code…
Amaryllis Mavragani, Rohan Alexander, Hyo Jung Kim, Manping Guo + 11 more
'Yiming Wang' 'Qiaoning Yang' 'Rui Li' 'Yang Zhao' 'Chenfei Li' 'Mingbo Zhu' 'Yao Cui' 'Xin Jiang' 'Song Sheng' 'Qingna Li' 'Rui Gao'] With the rapid development of science, technology, and engineering, large amounts of data have been generated in many fields in the past 20 years. In the process of medical research…
M. Zanin, D. Papo, P. A. Sousa, E. Menasalvas + 3 more
The increasing power of computer technology does not dispense with the need to extract meaningful in-formation out of data sets of ever growing size, and indeed typically exacerbates the complexity of this task. To tackle this general problem, two methods have emerged, at chronologically different times, that are now…
Martin A. Lindquist, Stephan Geuter, Tor D. Wager, Brian S. Caffo
The preprocessing pipelines typically used in both task and restingstate fMRI (rs-fMRI) analysis are modular in nature: They are composed of a number of separate filtering/regression steps, including removal of head motion covariates and band-pass filtering, performed sequentially and in a flexible order. In this paper…
Gökmen Altay, Jose Zapardiel-Gonzalo, Bjoern Peters
Gene network inference (GNI) methods have the potential to reveal functional relationships between different genes and their products. Most GNI algorithms have been developed for microarray gene expression datasets and their application to RNA-seq data is relatively recent. As the characteristics of RNA-seq data are…
Authors not listed
In recent years, the development of large language models (LLMs) has revolutionized various fields of natural science, yet their application in molecular data processing remains constrained due to the reliance on single-modality inputs and outputs. To bridge the gap between experimenters and computational tools, we…
Authors not listed
This study presents a validation and refinement of the “yellow cards” error detection workflow that can be applied to any property connected to molecular structure. In our implementation the workflow employed 5 predictive models with each assigning a “yellow card” to 5% of the entries with worst prediction accuracy.…
Feng Feng, Zhenru Chen, Jianyuan Ni, Yuanxun Zhang + 3 more
Drinking water is essential to public health and socioeconomic growth. Therefore, assessing and ensuring drinking water supply is a critical task in modern society. Conventional approaches to analyzing and controlling drinking water quality are labor-intensive and costly with a low throughput. Machine learning (ML) is…
Tobias Rehfeldt, Ralf Gabriels, Robbin Bouwmeester, Siegfried Gessulat + 6 more
Dataset acquisition and curation are often the hardest and most time-consuming parts of a machine learning endeavor. This is especially true for proteomics-based LC-IM-MS datasets, due to the high-throughput data structure with high levels of noise and complexity between raw and machine learning-ready formats. While…
Authors not listed
Kinetic modeling is essential for predicting changes in food quality during processing and storage. This study evaluates the application of physics-informed neural networks (PINN) for food kinetic modeling, integrating kinetic insights into neural network frameworks. Based on three case studies, namely seed drying…
Sarah E. Lindley, Yiyang Lu, Diwakar Shukla
Guide to Machine Learning for Small Molecule Design Authors: ['Sarah\nE. Lindley' 'Yiyang Lu' 'Diwakar Shukla'] Initially part of the field of artificial intelligence, machine learning (ML) has become a booming research area since branching out into its own field in the 1990s. After three decades of refinement, ML…