24 papers · ranked by Valyu relevance
Hugo López-Fernández, Osvaldo Graña-Castro, Alba Nogueira-Rodríguez, Miguel Reboiro-Jato + 2 more
'Miguel Reboiro-Jato' 'Daniel Glez-Peña' 'Robert Winkler'] Compi is an application framework to develop end-user, pipeline-based applications with a primary emphasis on: (i) user interface generation, by automatically generating a command-line interface based on the pipeline specific parameter definitions; (ii)…
Jaroslav Budiš, Werner Krampl, Marcel Kucharík, Rastislav Hekel + 12 more
'Adrián Goga' 'Jozef Sitarčík' 'Michal Lichvár' 'Dávid Smol’ak' 'Miroslav Böhmer' 'Andrej Baláž' 'František Ďuriš' 'Juraj Gazdarica' 'Katarína Šoltys' 'Ján Turňa' 'Ján Radvánszky' 'Tomáš Szemes'] Title: Abstract With the rapid growth of massively parallel sequencing technologies, still more laboratories are utilising…
Jaroslav Budiš, Werner Krampl, Marcel Kucharík, Rastislav Hekel + 11 more
'Adrian Goga' 'Michal Lichvár' 'Dávid Smoľak' 'Miroslav Böhmer' 'Andrej Baláž' 'František Ďuriš' 'Juraj Gazdarica' 'Katarína Šoltýs' 'Ján Turňa' 'Ján Radvánszky' 'Tomáš Szemes'] Geneton Ltd., 841 04 Bratislava, Slovakia Slovak Centre of Scientific and Technical Information, 811 04 Bratislava, Slovakia Comenius…
Anthony Federico, Tanya Karagiannis, Kritika Karri, Dileep Kishore + 3 more
The advent of high-throughput sequencing technologies has led to the need for flexible and user-friendly data pre-processing platforms. The Pipeliner framework provides an out-of-the-box solution for processing various types of sequencing data. It combines the Nextflow scripting language and Anaconda package manager to…
Anthony Federico, Tanya Karagiannis, Kritika Karri, Dileep Kishore + 3 more
'Yusuke Koga' 'Joshua D. Campbell' 'Stefano Monti'] The advent of high-throughput sequencing technologies has led to the need for flexible and user-friendly data preprocessing platforms. The Pipeliner framework provides an out-of-the-box solution for processing various types of sequencing data. It combines the Nextflow…
Sahil Seth, Samir Amin, Xingzhi Song, Xizeng Mao + 3 more
Bioinformatics analyses have become increasingly intensive computing processes, with lowering costs and increasing numbers of samples. Each laboratory spends time creating and maintaining a set of pipelines, which may not be robust, scalable, or efficient. Further, the existence of different computing environments…
Mathieu Bourgey, Rola Dali, Robert Eveleigh, Kuang Chung Chen + 19 more
With the decreasing cost of sequencing and the rapid developments in genomics technologies and protocols, the need for validated bioinformatics software that enables efficient large-scale data processing is growing. Here we present GenPipes, a flexible Python-based framework that facilitates the development and…
Jason P. Kurs, Manuele Simi, Fabien Campagne
Computational workflows and pipelines are often created to automate series of processing steps. For instance, workflows enable one to standardize analysis for large projects or core facilities, but are also useful for individual biologists who need to perform repetitive data processing. Some workflow systems, designed…
Pablo Cingolani, Rob Sladek, Mathieu Blanchette
Motivation: The analysis of large biological datasets often requires complex processing pipelines that run for a long time on large computational infrastructures. We designed and implemented a simple script-like programming language with a clean and minimalist syntax to develop and manage pipeline execution and provide…
Bruno Dantas, Calmenelias Fleitas, Alexandre P. Francisco, José Simão + 1 more
'José Simão' 'Cátia Vaz'] Biosciences have been revolutionized by next generation sequencing (NGS) technologies in last years, leading to new perspectives in medical, industrial and environmental applications. And although our motivation comes from biosciences, the following is true for many areas of science: published…
Marcin Cieślik, Cameron Mura
PaPy, which stands for parallel pipelines in Python, is a highly flexible framework that enables the construction of robust, scalable workflows for either generating or processing voluminous datasets. A workflow is created from user-written Python functions (nodes) connected by 'pipes' (edges) into a directed acyclic…
Samuel Lampa, Martin Dahlö, Jonathan Alvarsson, Ola Spjuth
The complex nature of biological data has driven the development of specialized software tools. Scientific workflow management systems simplify the assembly of such tools into pipelines and assist with job automation and aids reproducibility of analyses. Many contemporary workflow tools are specialized and not designed…
Camillo Lugaresi, Jiuqiang Tang, Hadon Nash, Chris McClanahan + 10 more
'Esha Uboweja' 'Michael L. Hays' 'Fan Zhang' 'Chuo-Ling Chang' 'Ming Guang Yong' 'Ju Hyun Lee' 'Wan-Teh Chang' 'Wei Hua' 'Manfred Georg' 'Matthias Grundmann'] Building applications that perceive the world around them is challenging. A developer needs to (a) select and develop corresponding machine learning algorithms…
Jochen Sieg, Christian Wolfgang Feldmann, Jennifer Hemmerich, Conrad Stork + 3 more
The open-source package scikit-learn provides various machine learning algorithms and data processing tools, including the Pipeline class, which allows users to prepend custom data transformation steps to the machine learning model. We introduce the MolPipeline package, which extends this concept to chemoinformatics by…
Michael Grauer, Patrick Reynolds, Marion Hoogstoel, Francois Budin + 2 more
'Martin A. Styner' 'Ipek Oguz'] Image processing is an important quantitative technique for neuroscience researchers, but difficult for those who lack experience in the field. In this paper we present a web-based platform that allows an expert to create a brain image processing pipeline, enabling execution of that…
Hyungro Lee, André Merzky, Li Lynn Tan, Mikhail Titov + 14 more
'Matteo Turilli' 'Dario Alfè' 'Agastya P. Bhati' 'Alex Brace' 'Austin Clyde' 'Peter V. Coveney' 'Heng Ma' 'Arvind Ramanathan' 'Rick Stevens' 'Anda Trifan' 'Hubertus J. J. van Dam' 'Shunzhou Wan' 'Sean Wilkinson' 'Shantenu Jha'] Abstract—COVID-19 has claimed more 106 lives and resulted in over 40 × 106 infections. There…
Nikolay O. Nikitin, Sergey Teryoshkin, Valerii Pokrovskii, Sergey Pakulin + 1 more
'Sergey Pakulin' 'Denis Nasonov'] Resource-intensive computations are a major factor that limits the effectiveness of automated machine learning solutions. In the paper, we propose a modular approach that can be used to increase the quality of evolutionary optimization for modelling pipelines with a graph-based…
Authors not listed
This comprehensive review examines the evolution of autonomous materials synthesis laboratories that integrate artificial intelligence with advanced robotics to accelerate discovery. Traditional materials development pipelines typically require 10-20 years, but self-driving laboratories (SDLs) and Materials…
Jumana Dakka, Matteo Turilli, David W. Wright, Stefan J. Zasada + 4 more
'Vivek Balasubramanian' 'Shunzhou Wan' 'Peter V. Coveney' 'Shantenu Jha'] Background Resistance to chemotherapy and molecularly targeted therapies is a major factor in limiting the effectiveness of cancer treatment. In many cases, resistance can be linked to genetic changes in target proteins, either pre-existing or…
Nikolay O. Nikitin, Pavel Vychuzhanin, Mikhail Sarafanov, Iana S. Polonskaia + 5 more
'Iana S. Polonskaia' 'Ilia Revin' 'Irina V. Barabanova' 'Gleb Maximov' 'Anna V. Kalyuzhnaya' 'Alexander V. Boukhanovsky'] The effectiveness of the machine learning methods for real-world tasks depends on the proper structure of the modeling pipeline. The proposed approach is aimed to automate the design of composite…
Peter Kraus, Edan Bainglass, Francisco F. Ramirez, Enea Svaluto-Ferro + 7 more
Compliance with good research data management practices means trust in the integrity of the data, and it is achievable by a full control of the data gathering process. In this work, we demonstrate tooling which bridges these two aspects, and illustrate its use in a case study of automated battery cycling. We…
Authors not listed
We have developed Aitomia – a platform powered by AI to assist in performing AI-driven atomistic and quantum chemical (QC) simulations. This evolving intelligent assistant platform is equipped with chatbots and AI agents to help experts and guide non-experts in setting up and running atomistic simulations, monitoring…
Althea Hansel-Harris, Andreas Tillack, Diogo Santos-Martins, Matthew Holcomb + 1 more
Virtual screening using molecular docking is now routinely used for the rapid evaluation of very large ligand libraries. As such, it has become an increasingly common approach in early-stage drug discovery. These screenings generate large amounts of data proportional to the size of the compound library used, which must…
Authors not listed
The era of exascale computing presents both exciting opportunities and unique challenges for quantum mechanical simulations. While the transition from petaflops to exascale computing has been marked by a steady increase in computational power, the shift towards heterogeneous architectures, particularly the dominant…