18 papers · ranked by Valyu relevance
Antonio Capizzi, Salvatore Distefano, Manuel Mazzara
DevOps is a quite effective approach for managing software development and operation, as confirmed by plenty of success stories in real applications and case studies. DevOps is now becoming the mainstream solution adopted by the software industry in development, able to reduce the time to market and costs while…
Jia Xu, Humza Naseer, Sean B. Maynard, Justin Fillipou
Digital business transformation has become increasingly important for organizations. Since transforming business digitally is an ongoing process, it requires an integrated and disciplined approach. Data Operations (DataOps), emerging in practice, can provide organizations with such an approach to leverage data and…
Scrocca, Mario, Grassi, Marco + 8 more
The implementation of AI-based applications in complex environments often requires the collaboration of several devices spanning from edge to cloud. Identifying the required devices and configuring them to collaborate is a challenge relevant to different scenarios, like industrial shopfloors, road infrastructures, and…
Dmytro Valiaiev
The proliferation of SQL for data processing has often occurred without the rigor of traditional software development, leading to siloed efforts, logic replication, and increased risk. This ad-hoc approach hampers data governance and makes validation nearly impossible. Organizations are adopting DataOps, a methodology…
Raúl Miñón, Josu Diaz-de-Arcaya, Ana I. Torre-Bastida, Juan López-de-Armentia + 4 more
'Juan López-de-Armentia' 'Gorka Zarate' 'Lander Bonilla' 'Asier Garcia-Perez' 'Jon Aguirre-Usandizaga'] Machine learning is already integrated in diverse domains enhancing their performance and decision support. For laboratories, this approach is normally sufficient. However, in real environments, these models can not…
Anthony Mammoliti, Petr Smirnov, Minoru Nakano, Zhaleh Safikhani + 3 more
Reproducibility is essential to Open Science, as there is limited relevance for finding that cannot be reproduced by independent research groups, regardless of its validity. It is therefore crucial for scientists to describe their experiments in sufficient detail so they can be reproduced, challenged, and built upon.…
Nicolas Le Guillarme, Wilfried Thuiller
With the rapid accumulation of biodiversity data, data integration has emerged as a hot topic in soil ecology. Data integration has indeed the potential to advance our knowledge of global patterns in soil biodiversity by facilitating large-scale meta-analytical studies of soil ecosystems. However, ecologists are still…
Michael Statt, Kristopher Brown, Santosh Suram, Linda Hung + 3 more
In this work, we present DBgen, a Python library that provides a framework for defining extract-transform-load (ETL) pipelines to create and populate SQL databases. DBgen is most useful when the underlying data has complex relationships, requires multi-step analysis, is large-scale, and the type of data being collected…
Robert W Reid, Jacob W Ferrier, Jeremy J Jay
Databio is capable of providing fast and accurate annotation of gene-oriented data sets, coupled with an integrated identifier conversion service to empower downstream data mining and computational analysis. Databio is enabled by fast real-time data structures applied to over 137 million unique identifiers, and uses…
Yannick Marcon, Tom Bishop, Demetris Avraam, Xavier Escriba-Montagut + 5 more
'Patricia Ryser-Welch' 'Stuart Wheater' 'Paul Burton' 'Juan R. González' 'Dina Schneidman-Duhovny'] Combined analysis of multiple, large datasets is a common objective in the health- and biosciences. Existing methods tend to require researchers to physically bring data together in one place or follow an analysis plan…
Michael Statt, Brian Rohr, Dan Guevarra, Ja'Nya Breeden + 2 more
Materials knowledge is inherently hierarchical. While high-level descriptors such as composition and structure are valuable for contextualizing materials data, the data must ultimately be considered in the context of its low-level acquisition details. Graph databases offer an opportunity to represent hierarchical…
Chih Chuan Shih, Jieqi Chen, Ai Shan Lee, Nicolas Bertin + 34 more
Genomic researchers are increasingly utilizing commercial cloud platforms (CCPs) to manage their data and analytics needs. Commercial clouds allow researchers to grow their storage and analytics capacity on demand, keeping pace with expanding project data footprints and enabling researchers to avoid large capital…
Petar Radanliev, David De Roure
With the increased digitalisation of our society, new and emerging forms of data present new values and opportunities for improved data driven multimedia services, or even new solutions for managing future global pandemics (i.e., Disease X). This article conducts a literature review and bibliometric analysis of…
Dani Arribas-Bel, Mark Green, Francisco Rowe, Alex Singleton
This paper develops the notion of “open data product”. We define an open data product as the open result of the processes through which a variety of data (open and not) are turned into accessible information through a service, infrastructure, analytics or a combination of all of them, where each step of development is…
Authors not listed
Raman spectroscopy is an increasingly powerful and fast-growing analytical technique across diverse disciplines, from materials science and chemistry to biology and medicine, thanks to advances in Raman instrumentation and greatly supported by the flourishing of chemometrics and artificial intelligence (AI). However…
Juha-Pekka Soininen, Carlos Fernández Sánchez, Stefano Modafferi, Stuart Campbell + 5 more
This paper presents an interoperability architecture for data spaces (DSIA), allowing members of the different data spaces using different technologies and services to exchange data. The DSIA builds on the existing trust of data space members and extends it through shared agreements and additional federation and…
Rosie Higman, Marta Teperek, Danny Kingsley
Research Data Management (RDM) presents an unusual challenge for service providers in Higher Education. There is increased awareness of the need for training in this area but the nature of discipline-specific practices involved make it difficult to provide training across a multi-disciplinary organisation. Whilst most…
Authors not listed
The nanosafety domain has seen significant advancements in data generation and sharing, yet challenges remain in ensuring data interoperability and reuse. This article focuses on developing a semantic interoperability framework for nanosafety data to maximize the FAIRness (Findability, Accessibility, Interoperability…