Search · four archives
Search · four archives
21 papers · ranked by Valyu relevance
Habeeb Abolaji Babatunde, Owen M. McDougal, Timothy Andersen, Hongbin Pu
'Hongbin Pu'] The preprocessing of infrared spectra can significantly improve predictive accuracy for protein, carbohydrate, lipid, or other nutrition components, yet optimal preprocessing selection is typically empirical, tedious, and dataset specific. This study introduces a Bayesian optimization-based framework…
Alexander Isenko, Ruben Mayer, Jeffrey Jedele, Hans‐Arno Jacobsen
Preprocessing pipelines in deep learning aim to provide sufficient data throughput to keep the training processes busy. Maximizing resource utilization is becoming more challenging as the throughput of training processes increases with hardware innovations (e.g., faster GPUs, TPUs, and inter-connects) and advanced…
Peng Li, Zhiyi Chen, Xu Chu, Kexin Rong
Data preprocessing is a crucial step in the machine learning process that transforms raw data into a more usable format for downstream ML models. However, it can be costly and time-consuming, often requiring the expertise of domain experts. Existing automated machine learning (AutoML) frameworks claim to automate data…
Sai Prakash Challa, Melvin Alexis Lara de Leon, Jiri Koziorek, Ibrahim A. Hameed + 1 more
Machine vision and AI-based defect detection systems are increasingly deployed in manufacturing to support consistent product quality and high production efficiency. However, these automated inspection systems often suffer from sensitivity to imaging variability, dependence on large labeled datasets, and the need for…
Mathieu Dugré, Yohan Chatelain, Tristan Glatard
Magnetic resonance imaging (MRI) preprocessing is a critical step for neuroimaging analysis. However, the computational cost of MRI preprocessing pipelines is a major bottleneck for large cohort studies and some clinical applications. While high-performance computing and, more recently, deep learning have been adopted…
Paulito Palmes, Akihiro Kishimoto, Radu Marinescu, Parikshit Ram + 1 more
'Elizabeth Daly'] The pipeline optimization problem in machine learning requires simultaneous optimization of pipeline structures and parameter adaptation of their elements. Having an elegant way to express these structures can help lessen the complexity in the management and analysis of their performances together…
Sebastian Pineda Arango, Josif Grabocka
Automated Machine Learning (AutoML) is a promising direction for democratizing AI by automatically deploying Machine Learning systems with minimal human expertise. The core technical challenge behind AutoML is optimizing the pipelines of Machine Learning systems (e.g. the choice of preprocessing, augmentations, models…
Gufran Ahmad Ansari, Salliah Shafi, Lamees Alhazzaa, Dechang Chen + 2 more
Background: Lung cancer remains one of the leading causes of cancer-related mortality worldwide, primarily due to late diagnosis. Although machine learning (ML) techniques have been widely applied for lung cancer classification, many studies lack a fully optimized end-to-end pipeline using routine clinical data. This…
Olesya Melnichenko, Venkat S. Malladi
In the field of genomics, bioinformatics pipelines play a crucial role in processing and analyzing vast biological datasets. These pipelines, consisting of interconnected tasks, can be optimized for efficiency and scalability by leveraging cloud platforms such as Microsoft Azure. The choice of compute resources…
Leonardo Rosa Amado, Adriano Vogel, Dalvan Griebler, Gabriel Paludo Licks + 2 more
'Gabriel Paludo Licks' 'Eric Simon' 'Felipe Meneguzzi'] Abstract. Data pipeline frameworks provide abstractions for implementing sequences of data-intensive transformation operators, automating the deployment and execution of such transformations in a cluster. Deploying a data pipeline, however, requires computing…
Jochen Sieg, Christian Wolfgang Feldmann, Jennifer Hemmerich, Conrad Stork + 3 more
The open-source package scikit-learn provides various machine learning algorithms and data processing tools, including the Pipeline class, which allows users to prepend custom data transformation steps to the machine learning model. We introduce the MolPipeline package, which extends this concept to chemoinformatics by…
Elliot Xie, Lingxin Cheng, Yujia Cai, Jack Shireman + 1 more
Performance bottlenecks in widely used genomics and bioinformatics software present a substantial and growing burden as biological datasets continue to increase in size and number. Relieving these bottlenecks relies largely on expert manual optimization and therefore remains difficult to scale. Here we present…
Isabel Mogollon, Michaela Feodoroff, Pedro Neto, Alba Montedeoca + 2 more
Understanding cellular function within 3D multicellular spheroids is essential for advancing cancer research, particularly in studying cell-stromal interactions as potential targets for novel drug therapies. However, accurate single-cell segmentation in 3D cultures is challenging due to dense cell clustering and the…
Jiang Wu, Hongbo Wang, Chunhe Ni, Chenwei Zhang + 1 more
—Data Pipeline plays an indispensable role in tasks such as modeling machine learning and developing data products. With the increasing diversification and complexity of Data sources, as well as the rapid growth of data volumes, building an efficient Data Pipeline has become crucial for improving work efficiency and…
Authors not listed
The integration of artificial intelligence technologies into pharmaceutical research is crucial for gaining an early understanding of molecular properties, thereby facilitating successful drug design. Constructing a machine learning (ML) model however, requires knowledge spanning from data preprocessing and feature…
Authors not listed
Methanol synthesis from syngas (CO/CO₂/H₂) is vital for sustainable chemical production; however, traditional kinetic models hinder rapid reactor optimisation [1]. We present a reproducible machine-learning pipeline to predict methanol yield in a double-pass plug-flow reactor, utilising a synthetic dataset (n = 5,000)…
Heeseung Lee, Daeho Kim, Heyin Lee, Namyoung Gwak + 6 more
- 1. Computational Science Research Center, Korea Institute of Science and Technology, Seoul 02792, Republic of Korea - 2. Department of Materials Science and Engineering, Korea University, 145 Anam-ro, Seoul 02841, Republic of Korea - 3. Department of Chemical and Biological Engineering, Korea University, Seoul 02841…
Authors not listed
Bayesian optimization (BO) has become increasingly important for experimental optimization across scientific domains, yet implementing BO pipelines requires significant programming expertise and familiarity with specialized frameworks. This creates a barrier for domain experts who could benefit from BO but lack the…
Guido Schlögel, Rüdiger Lück, Stefan Kittler, Oliver Spadiut + 3 more
Biotechnological production of recombinant molecules relies heavily on fed-batch processes. However, as the cells’ growth, substrate uptake, and production kinetics are often unclear, the fed-batches are frequently operated under sub-optimal conditions. Process design is based on simple feed profiles (e.g., constant or…
Raúl Miñón, Josu Diaz-de-Arcaya, Ana I. Torre-Bastida, Juan López-de-Armentia + 4 more
'Juan López-de-Armentia' 'Gorka Zarate' 'Lander Bonilla' 'Asier Garcia-Perez' 'Jon Aguirre-Usandizaga'] Machine learning is already integrated in diverse domains enhancing their performance and decision support. For laboratories, this approach is normally sufficient. However, in real environments, these models can not…
Authors not listed
The global drive towards net-zero has accelerated the adoption of carbon fibre reinforced polymers (CFRP) for lightweight structures in various sectors such as aerospace, automotive, energy and biomedical. Mechanical machining of CFRP is often necessary to meet dimensional or assembly-related requirements. However…