26 papers · ranked by Valyu relevance
Koteeswaran Seerangan, Premalatha Gunasekaran, Nithya Rekha Sivakumar, Resmi Ravi Nair + 4 more
Background/Objectives: Diabetes is one of the most familiar and common diseases among people currently, and is a type of metabolic disease that is caused due to high levels of sugar in the blood for longer periods of time. If the disease is predicted at an earlier stage, the severity and risks associated with diabetes…
Maxwell, David S, Darkoh, Michael + 8 more
- 1 Data Impact and Governance, The University of Texas MD Anderson Cancer Center, Houston, Texas, USA - 2 Department of Genomic Medicine, The University of Texas MD Anderson Cancer Center, Houston, Texas, USA - 3 The Institute for Data Science in Oncology, The University of Texas MD Anderson Cancer Center, Houston…
WENBIAO LI, ANISA HALIMI, JAIDEEP VAIDYA, XIAOQIAN JIANG + 1 more
We present a privacy-preserving framework to verify whether a declared data preprocessing pipeline was correctly applied before training a machine learning model on sensitive data. The verifier has only black-box query access to the model and combines three behavior indicators: shift in prediction accuracy…
Raffaele Emanuele Russo, Martina Fattobene, Silvia Zamponi, Paolo Conti + 7 more
Environmental element monitoring is essential for assessing environmental quality, identifying pollution sources, evaluating ecological risks, and understanding long-term contamination trends. Modern monitoring campaigns routinely generate large volumes of complex data that require advanced analytical strategies. This…
Craig Michoski, Matthew M. Waller, B. Sammuli, Zeyu Li + 14 more
Craig Michoskia,d, ∗ , Matthew Waller a , Brian Sammuli b , Zeyu Li b , Tapan Ganatma Nakkina a , Raffi Nazikian b , Sterling Smith b , David Orozco b , Dongyang Kuang a , Martin Foltin c , Erik Olofsson b , Mike Fredrickson a , Jerry Louis-Jeune a , David R. Hatcha,d, Todd A. Olivera,d, Mitchell Clark b , Steph-Yves…
Zidong Yan, Jiaqi Li, Weican Zhang, Haonan Wen + 4 more
on Best Practices and Pitfalls in Applying Machine Learning to Environmental Research Authors: Zidong Yan, Jiaqi Li, Weican Zhang, Haonan Wen, Hao Yu, Miao Yu, Qian Liu, Guibin Jiang Machine learning (ML) has become a powerful paradigm for extracting structures from complex environmental data and supporting scientific…
Daria D. Tyurina, Sergey V. Stasenko, Konstantin V. Lushnikov, Maria V. Vedunova
This study introduces a novel method for predicting cognitive age using psychophysiological tests. To determine cognitive age, subjects were asked to complete a series of psychological tests measuring various cognitive functions, including reaction time and cognitive conflict, short-term memory, verbal functions, and…
Jose Manuel Valencia-Moreno, Everardo Gutierrez-Lopez, Jose Angel Gonzalez-Fraga, Rodolfo Alan Martinez Rodriguez + 2 more
This article presents a dataset of breast cancer risk factors collected from 1697 Cuban women between 2001 and 2018, as a tool to design and support the development and validation of predictive models in public health for breast cancer risk. A reproducible methodology for quality control and variable enrichment was…
Shahab Ahmad Al Maaytah, Ayman Qahmash, Zeyar Aung
Predicting whether a patient will attend a scheduled medical appointment is essential for reducing inefficiencies in healthcare systems and optimizing resource allocation. This study introduces a local, LLM-assisted pipeline that uses LLaMA 7B solely to automate semantic preprocessing such as column renaming, datatype…
Henrik Meyer, Lars Ahlers, Pedro Querini, Erica Fernandez + 3 more
Machine Learning, Artificial Intelligence, among others, are very promising methodologies and technologies that are emerging for implementing a broad spectrum of analytics within digitalized eco-systems. Analytics containing adequate analytical models are generating a burgeoning interest from Business Intelligence-…
Daniele Scanzi, Dylan A. Taylor, Katie A. McNair, Rohan O. C. King + 2 more
Electroencephalography (EEG) data are inherently contaminated by non-neuronal noise, including eye movements, muscle activity, cardiac signals, electrical interference, and technical issues such as poorly connected electrodes. Preprocessing to remove these artefacts is essential, yet the optimal method remains unclear…
Xingyu Liu, Yijun Zhang, Zi Yin, Zonglei Zhen + 1 more
Macaque MRI bridges non-invasive systems neuroscience with cellular and circuit-level mechanisms, but preprocessing tools remain difficult to integrate and deploy reproducibly. We present Brainana, an automated, BIDS-compatible preprocessing and visualization framework for macaque neuroimaging. Brainana integrates…
Authors not listed
The discovery of chemically novel or structurally anomalous metal-organic frameworks (MOFs) is essential for expanding reticular design space and enhancing dataset reliability. We present CHEM-AD (Chemically Unusual Metal–organic Frameworks via Autoencoder-based Detection), a label-free, CPU-efficient pipeline that…
Loïc Brun, Jonas Rothrock, Erica van de Waal, Ebi Antony George
Although the use of accelerometer-based behavioural classification to quantify animal activity budgets is gaining widespread traction, the interactions between key preprocessing decisions and modern classification algorithms remain poorly understood. Moreover, classification pipelines are commonly assessed using global…
Tobias K. Mildenberger, Federico Maioli, Casper W. Berg
Scientific bottom-trawl surveys provide essential fisheries-independent data for fisheries and ecosystem research. In the Northeast Atlantic, the ICES Database of Trawl Surveys (DATRAS) compiles haul-level information, species- and length-specific catch data, and individual biological observations across multiple…
Matthieu Vilain, Stéphane Aris-Brosou
The ever-growing amount of available biological data leads modern analysis to be performed on large datasets. Unfortunately, bioinformatics tools for preprocessing and analyzing data are not always designed to treat such large amounts of data efficiently. Notably, this is the case when encoding DNA and RNA sequences…
Phanindra Reddy Madduru, Bijo Thomas
This paper proposes a preprocessing framework for optimizing large-scale graph database ingestion through intelligent edge filtering based on value ranking. We combine adapted PageRank algorithms with business-specific metrics and edge type importance to evaluate and rank edges, enabling selective retention of…
Authors not listed
The integration of machine learning methods is transforming many areas of research by, for instance, accelerating molecular dynamics simulations and enabling improved prediction and optimization of chemical reactions. However, despite this progress, the adoption of data-driven approaches in atomic layer deposition…
Marcos L. P. Bueno, Vanschoren, Joaquin
The goal of automated machine learning (AutoML) is to reduce trial and error when doing machine learning (ML). Although AutoML methods for classification are able to deal with data imperfections, such as outliers, multiple scales and missing data, their behavior is less known on dirty categorical datasets. These…
Chunqing (Tony) Liang, Tajveer Grewal, Asees Singh, Amrit Singh
Multimodal biomedical studies increasingly profile multiple molecular and clinical modalities from the same samples, creating new opportunities for disease prediction and biological discovery. However, benchmarking multimodal integration methods remains difficult because studies often use inconsistent preprocessing…
Authors not listed
Cycloaddition of CO2 to epoxides is considered to be a promising atom-efficient process for converting of the CO2 emissions to the valuable products. The present analytical review is dedicated to computational multivariance analysis of the published experimental data for homogeneously catalyzed reactions. In the vast…
Peter G. Hawkins, Eli M. Swanson, Megan Feichtel
The size of individual single cell samples continues to grow with advancing technologies, as do the number of samples included in individual experiments and across organizations. This presents challenges for processing this data at scale, both in terms of computational throughput and the required size of the machines…
Authors not listed
The materials-science literature is the richest reservoir of domain knowledge, yet converting its unstructured text—especially narrative passages and complex tables—into machine-readable data for analysis and ML model training remains challenging. To address this, we present KnowMat, an agentic, multi-stage pipeline…
Aljawarneh, Shadi, Lara, Juan A. + 2 more
The Meteorology is a field where huge amounts of data are generated, mainly collected by sensors at weather stations, where different variables can be measured. Those data have some particularities such as high volume and dimensionality, the frequent existence of missing values in some stations, and the high…
Authors not listed
Mass spectrometry (MS) generates large datasets that are stored in increasingly optimized and complex file types, demanding technical expertise to extract information rapidly and easily. We wondered whether a simple structured query language (SQL) database could hold raw MS data and allow for easily readable queries…
Authors not listed
High-throughput experimentation (HTE) accelerates chemical discovery by shortening the lead times for molecule synthesis. The choice of initial reaction conditions directly influences the outcome and length of a screening campaign. But human involvement in plate design and data analysis remains a significant cost…