21 papers · ranked by Valyu relevance
Anthony Federico, Tanya Karagiannis, Kritika Karri, Dileep Kishore + 3 more
The advent of high-throughput sequencing technologies has led to the need for flexible and user-friendly data pre-processing platforms. The Pipeliner framework provides an out-of-the-box solution for processing various types of sequencing data. It combines the Nextflow scripting language and Anaconda package manager to…
Oleksandr Semeniuta, Petter Falkman, Mario Luca Bernardi
Many data processing systems are naturally modeled as pipelines, where data flows though a network of computational procedures. This representation is particularly suitable for computer vision algorithms, which in most cases possess complex logic and a big number of parameters to tune. In addition, online vision…
Anthony Federico, Tanya Karagiannis, Kritika Karri, Dileep Kishore + 3 more
'Yusuke Koga' 'Joshua D. Campbell' 'Stefano Monti'] The advent of high-throughput sequencing technologies has led to the need for flexible and user-friendly data preprocessing platforms. The Pipeliner framework provides an out-of-the-box solution for processing various types of sequencing data. It combines the Nextflow…
Pablo Cingolani, Rob Sladek, Mathieu Blanchette
Motivation: The analysis of large biological datasets often requires complex processing pipelines that run for a long time on large computational infrastructures. We designed and implemented a simple script-like programming language with a clean and minimalist syntax to develop and manage pipeline execution and provide…
Jochen Sieg, Christian Wolfgang Feldmann, Jennifer Hemmerich, Conrad Stork + 3 more
The open-source package scikit-learn provides various machine learning algorithms and data processing tools, including the Pipeline class, which allows users to prepend custom data transformation steps to the machine learning model. We introduce the MolPipeline package, which extends this concept to chemoinformatics by…
Michael Grauer, Patrick Reynolds, Marion Hoogstoel, Francois Budin + 2 more
'Martin A. Styner' 'Ipek Oguz'] Image processing is an important quantitative technique for neuroscience researchers, but difficult for those who lack experience in the field. In this paper we present a web-based platform that allows an expert to create a brain image processing pipeline, enabling execution of that…
Philippe Hauchamps, Babak Bayat, Simon Delandre, Mehdi Hamrouni + 4 more
With the increase of the dimensionality in conventional flow cytometry data over the past years, there is a growing need to replace or complement traditional manual analysis (i.e. iterative 2D gating) with automated data analysis pipelines. A crucial part of these pipelines consists of pre-processing and applying…
Anthony Mbata, Somayajulu Sripada, Mingjun Zhong
Currently, a variety of pipeline tools are available for use in data engineering. Data scientists can use these tools to resolve data wrangling issues associated with data and accomplish some data engineering tasks from data ingestion through data preparation to utilization as input for machine learning (ML). Some of…
Michael Statt, Kristopher Brown, Santosh Suram, Linda Hung + 3 more
In this work, we present DBgen, a Python library that provides a framework for defining extract-transform-load (ETL) pipelines to create and populate SQL databases. DBgen is most useful when the underlying data has complex relationships, requires multi-step analysis, is large-scale, and the type of data being collected…
Petar Maymounkov
We propose a new result-oriented semantic for dening data processing workows that manipulate data in dierent semantic forms (les or services) in a unied manner. is approach enables users to dene workows for a vast variety of reproducible data-processing tasks in a simple declarative manner which focuses on…
Samuel Lampa, Martin Dahlö, Jonathan Alvarsson, Ola Spjuth
The complex nature of biological data has driven the development of specialized software tools. Scientific workflow management systems simplify the assembly of such tools into pipelines and assist with job automation and aids reproducibility of analyses. Many contemporary workflow tools are specialized and not designed…
Onur Yukselen, Osman Turkyilmaz, Ahmet Rasit Ozturk, Manuel Garber + 1 more
The emergence of high throughput technologies that produce vast amounts of genomic data, such as next-generation sequencing (NGS) are transforming biological research. The dramatic increase in the volume of data makes analysis the main bottleneck for scientific discovery. The processing of high throughput datasets…
Christopher Bogart, Rajeev Chhajer, Baljit Singh, Tony Fontana + 1 more
'Majd Sakr'] Abstract—As the volume of data available from sensor-enabled devices such as vehicles expands, it is increasingly hard for companies to make informed decisions about the cost of capturing, processing, and storing the data from every device. Business teams may do detailed forecasting of costs associated…
Felix Bänsch, Jonas Schaub, Betül Sevindik, Samuel Behr + 3 more
Developing and implementing computational algorithms for the extraction of specific substructures from molecular graphs (in silico molecule fragmentation) is an iterative process. It involves repeated sequences of implementing a rule set, applying it to relevant structural data, checking the results, and adjusting the…
Mark Burgess, Ewout Prangsma
—Koalja describes a generalized data wiring or 'pipeline' platform, built on top of Kubernetes, for plugin user code. Koalja makes the Kubernetes underlay transparent to users (for a 'serverless' experience), and offers a breadboarding experience for development of data sharing circuitry, to commoditize its gradual…
Sumon Biswas, Mohammad Wardat, Hridesh Rajan
Increasingly larger number of software systems today are including data science components for descriptive, predictive, and prescriptive analytics. The collection of data science stages from acquisition, to cleaning/curation, to modeling, and so on are referred to as data science pipelines. To facilitate research and…
Shadi A. Issa, Romeo Kienzler, Mohamed El-Kalioby, Peter J. Tonellato + 3 more
'Peter J. Tonellato' 'Dennis Wall' 'Rémy Bruggmann' 'Mohamed Abouelhoda'] Cloud computing provides a promising solution to the genomics data deluge problem resulting from the advent of next-generation sequencing (NGS) technology. Based on the concepts of “resources-on-demand” and “pay-as-you-go”, scientists with no or…
Niannian Wang, Jingzheng Zhang, Xiaotian Song, Songling Huang
Deep learning algorithms have achieved encouraging results for pipeline defect segmentation. However, existing defect segmentation methods may encounter challenges in accurately segmenting the complex features of pipeline defects and suffer from low processing speeds. Therefore, in this study, we propose…
Authors not listed
This comprehensive review examines the evolution of autonomous materials synthesis laboratories that integrate artificial intelligence with advanced robotics to accelerate discovery. Traditional materials development pipelines typically require 10-20 years, but self-driving laboratories (SDLs) and Materials…
Aaron Liu, Myeongyeon Lee, Rahul Venkatesh, Jessica Bonsu + 4 more
Polymer-based semiconductors and organic electronics encapsulate a significant research thrust for informatics-driven materials development. However, device measurements are described by a complex array of design and parameter choices, many of which are sparsely reported. For example, the mobility of a polymer-based…
Authors not listed
The exponential growth of chemical literature necessitates the development of automated tools for extracting and curating molecular information from unstructured scientific publications into open-access chemical databases. Current optical chemical structure recognition (OCSR) and named entity recognition solutions…