24 papers · ranked by Valyu relevance
Wilhelm Hasselbring, Maik Wojcieszak, Schahram Dustdar
When we consider the application layer [1] of networked infrastructures, data and control flow are important concerns in distributed systems integration. Modularity is a fundamental principle in software design [2], in particular for distributed system architectures. Modularity emphasizes high cohesion of individual…
Charlotte Capitanchik, Sam Ireland, Alex Harston, Chris Cheshire + 13 more
Ever-increasing volumes of sequencing data offer potential for large-scale meta-analyses to address significant biological questions. However, challenges such as insufficient data processing information, data quality concerns, and issues related to accessibility and curation often present obstacles. Additionally, most…
Catherine Jayapandian, Annan Wei, Priya Ramesh, Bilal Zonjy + 4 more
'Samden D. Lhatoo' 'Kenneth Loparo' 'Guo-Qiang Zhang' 'Satya S. Sahoo'] Data-driven neuroscience research is providing new insights in progression of neurological disorders and supporting the development of improved treatment approaches. However, the volume, velocity, and variety of neuroscience data generated from…
Harald Foidl, Valentina Golendukhina, Rudolf Ramler, Michael Felderer
'Michael Felderer'] Data pipelines are an integral part of various modern data-driven systems. However, despite their importance, they are often unreliable and deliver poor-quality data. A critical step toward improving this situation is a solid understanding of the aspects contributing to the quality of data…
Haruna Isah, Farhana Zulkernine
—An essential part of building a data-driven organization is the ability to handle and process continuous streams of data to discover actionable insights. The explosive growth of interconnected devices and the social Web has led to a large volume of data being generated on a continuous basis. Streaming data sources…
Anthony J. Kenyon, David Elizondo, Lipika Deka
— Typical event datasets such as those used in network intrusion detection comprise hundreds of thousands, sometimes millions, of discrete packet events. These datasets tend to be high dimensional, stateful, and time-series in nature, holding complex local and temporal feature associations. Packet data can be…
Klaus Kammerer, Rüdiger Pryss, Burkhard Hoppenstedt, Kevin Sommer + 1 more
'Manfred Reichert'] For machine manufacturing companies, besides the production of high quality and reliable machines, requirements have emerged to maintain machine-related aspects through digital services. The development of such services in the field of the Industrial Internet of Things (IIoT) is dealing with…
Sahil Seth, Samir Amin, Xingzhi Song, Xizeng Mao + 3 more
Bioinformatics analyses have become increasingly intensive computing processes, with lowering costs and increasing numbers of samples. Each laboratory spends time creating and maintaining a set of pipelines, which may not be robust, scalable, or efficient. Further, the existence of different computing environments…
Yue Chen, Jian Lu, Zhihong (Arry) Yao
The quality of traffic flow data is very important to the effective management and operation of urban traffic system. At present, most traffic flow data used in traffic flow research come from road sensors, but the shortcomings of long sampling period and sparse sampling points affect the quality control of traffic…
Samuel Lampa, Martin Dahlö, Jonathan Alvarsson, Ola Spjuth
The complex nature of biological data has driven the development of specialized software tools. Scientific workflow management systems simplify the assembly of such tools into pipelines and assist with job automation and aids reproducibility of analyses. Many contemporary workflow tools are specialized and not designed…
Michael Statt, Kristopher Brown, Santosh Suram, Linda Hung + 3 more
In this work, we present DBgen, a Python library that provides a framework for defining extract-transform-load (ETL) pipelines to create and populate SQL databases. DBgen is most useful when the underlying data has complex relationships, requires multi-step analysis, is large-scale, and the type of data being collected…
Marek Amanowicz, Damian Jankowski, Joanna Kolodziej
The increasing availability of mobile devices and applications, the progress in virtualisation technologies, and advances in the development of cloud-based distributed data centres have significantly stimulated the growing interest in the use of software-defined networks (SDNs) for both wired and wireless applications.…
Hasan Asy’ari Arief, Tomasz Wiktorski, Peter James Thomas, Nikolai Ushakov + 2 more
'Nikolai Ushakov' 'Leonid B. Liokumovich' 'Arthur H. Hartog'] Real-time monitoring of multiphase fluid flows with distributed fibre optic sensing has the potential to play a major role in industrial flow measurement applications. One such application is the optimization of hydrocarbon production to maximize short-term…
Authors not listed
Self-driving laboratories (SDLs) promise accelerated scientific discovery and product development by closing the loop between robotic execution and AI/ML-driven decision making. In practice, however, SDL orchestration remains fragmented; workflows are typically encoded as laboratory-specific scripts or bespoke…
Andrei Paleyes, Christian Cabrera, Neil D. Lawrence
As use of data driven technologies spreads, software engineers are more often faced with the task of solving a business problem using data-driven methods such as machine learning (ML) algorithms. Deployment of ML within large software systems brings new challenges that are not addressed by standard engineering…
Yu Qian, Olga Tchuvatkina, Josef Spidlen, Peter Wilkinson + 6 more
'Maura Gasparetto' 'Andrew R Jones' 'Frank J Manion' 'Richard H Scheuermann' 'Rafick-Pierre Sekaly' 'Ryan R Brinkman'] Background Flow cytometry technology is widely used in both health care and research. The rapid expansion of flow cytometry applications has outpaced the development of data storage and analysis tools.…
Taylor Reiter, Phillip T. Brooks, Luiz Irber, Shannon E.K. Joslin + 4 more
As the scale of biological data generation has increased, the bottleneck of research has shifted from data generation to analysis. Researchers commonly need to build computational workflows that include multiple analytic tools and require incremental development as experimental insights demand tool and parameter…
P. Anjali, A Binu
Big data analysis has become much popular in the present day scenario and the manipulation of big data has gained the keen attention of researchers in the field of data analytics. Analysis of big data is currently considered as an integral part of many computational and statistical departments. As a result, novel…
Mahnoor Zulfiqar, Michael R. Crusoe, Birgitta König-Ries, Christoph Steinbeck + 2 more
Scientific workflows facilitate the automation of data analysis tasks by integrating various software and tools executed in a particular order. To enable transparency and reusability in workflows, it is essential to implement the FAIR principles. Here, we describe our experiences implementing the FAIR principles for…
Aleksandar Jagličić, Torben Gädt, Matthias Hofmann
Isothermal heat flow calorimetry is a powerful method for studying chemical processes. In cement research, it has become indispensable for quantifying the heat release during cement hydration. It is used to study the reactivity of cementitious binders and the effect of admixture chemistry and dosage. Most isothermal…
Jason P. Kurs, Manuele Simi, Fabien Campagne
Computational workflows and pipelines are often created to automate series of processing steps. For instance, workflows enable one to standardize analysis for large projects or core facilities, but are also useful for individual biologists who need to perform repetitive data processing. Some workflow systems, designed…
Olukayode Majekodunmi, Sara Hashmi
In suspension flows through microchannels with parallel walls, rigid particles form clogs that grow continuously in the upstream direction. However, introducing a slight taper to channel walls leads to a qualitatively different clogging mechanism. Clogs of rigid particles do not grow continuously in these tapered…
Authors not listed
Real-world datasets in chemical engineering and bioengineering processes--such as those from catalytic reactors, multiphase flows, polymerization reactors, bioreactors, and clinical trials--can often be unlabelled or disorganized, rendering the training of existing supervised learning models ineffective at learning the…
Dominik Balazka, Dario Rodighiero
Starting from an analysis of frequently employed definitions of big data, it will be argued that, to overcome the intrinsic weaknesses of big data, it is more appropriate to define the object in relational terms. The excessive emphasis on volume and technological aspects of big data, derived from their current…