12 papers · ranked by Valyu relevance
Daniele Dall’Olio, Nico Curti, Eugenio Fonzi, Claudia Sala + 3 more
'Daniel Remondini' 'Gastone Castellani' 'Enrico Giampieri'] Background Current high-throughput technologies-i.e. whole genome sequencing, RNA-Seq, ChIP-Seq, etc.-generate huge amounts of data and their usage gets more widespread with each passing year. Complex analysis pipelines involving several…
Marco Aldinucci, Cristina Calcagno, Mario Coppo, Ferruccio Damiani + 5 more
'Maurizio Drocco' 'Eva Sciacca' 'Salvatore Spinella' 'Massimo Torquati' 'Angelo Troina'] The paper arguments are on enabling methodologies for the design of a fully parallel, online, interactive tool aiming to support the bioinformatics scientists .In particular, the features of these methodologies, supported by the…
Marcin Cieślik, Cameron Mura
Background Bioinformatic analyses typically proceed as chains of data-processing tasks. A pipeline, or 'workflow', is a well-defined protocol, with a specific structure defined by the topology of data-flow interdependencies, and a particular functionality arising from the data transformations applied at each step. In…
Michael Grauer, Patrick Reynolds, Marion Hoogstoel, Francois Budin + 2 more
'Martin A. Styner' 'Ipek Oguz'] Image processing is an important quantitative technique for neuroscience researchers, but difficult for those who lack experience in the field. In this paper we present a web-based platform that allows an expert to create a brain image processing pipeline, enabling execution of that…
Jochen Sieg, Christian Wolfgang Feldmann, Jennifer Hemmerich, Conrad Stork + 3 more
The open-source package scikit-learn provides various machine learning algorithms and data processing tools, including the Pipeline class, which allows users to prepend custom data transformation steps to the machine learning model. We introduce the MolPipeline package, which extends this concept to chemoinformatics by…
Satoshi Ito, Masaaki Yadome, Tatsuo Nishiki, Shigeru Ishiduki + 3 more
'Hikaru Inoue' 'Rui Yamaguchi' 'Satoru Miyano'] Background Supercomputers have become indispensable infrastructures in science and industries. In particular, most state-of-the-art scientific results utilize massively parallel supercomputers ranked in TOP500. However, their use is still limited in the bioinformatics…
Michael Statt, Kristopher Brown, Santosh Suram, Linda Hung + 3 more
In this work, we present DBgen, a Python library that provides a framework for defining extract-transform-load (ETL) pipelines to create and populate SQL databases. DBgen is most useful when the underlying data has complex relationships, requires multi-step analysis, is large-scale, and the type of data being collected…
Onur Yukselen, Osman Turkyilmaz, Ahmet Rasit Ozturk, Manuel Garber + 1 more
'Alper Kucukural'] Background The emergence of high throughput technologies that produce vast amounts of genomic data, such as next-generation sequencing (NGS) is transforming biological research. The dramatic increase in the volume of data, the variety and continuous change of data processing tools, algorithms and…
Authors not listed
Background: Pharmaceutical batch scheduling in multi-reactor configurations presents complex optimization challenges under operational uncertainty, yet limited research addresses how parallel processing capacity affects heuristic performance and predictive modeling. Objectives: This study investigated scheduling…
Ido Ben-Shalom, Charles Lin, Brian Radak, Woody Sherman + 1 more
Molecular dynamics (MD) simulations of proteins are commonly used to sample from the Boltzmann distribution of conformational states, with wide-ranging applications spanning chemistry, biophysics, and drug discovery. However, MD can be inefficient at equilibrating water occupancy for buried cavities in proteins that…
Authors not listed
This comprehensive review examines the evolution of autonomous materials synthesis laboratories that integrate artificial intelligence with advanced robotics to accelerate discovery. Traditional materials development pipelines typically require 10-20 years, but self-driving laboratories (SDLs) and Materials…
Peter Kraus, Edan Bainglass, Francisco F. Ramirez, Enea Svaluto-Ferro + 7 more
Compliance with good research data management practices means trust in the integrity of the data, and it is achievable by a full control of the data gathering process. In this work, we demonstrate tooling which bridges these two aspects, and illustrate its use in a case study of automated battery cycling. We…