24 papers · ranked by Valyu relevance
Amin Sahebi, Marco Barbone, Marco Procaccini, Wayne Luk + 2 more
'Georgi Gaydadjiev' 'Roberto Giorgi'] Processing large-scale graphs is challenging due to the nature of the computation that causes irregular memory access patterns. Managing such irregular accesses may cause significant performance degradation on both CPUs and GPUs. Thus, recent research trends propose graph…
Shantenu Jha, Daniel S. Katz, André Luckow, Omer Rana + 2 more
'Yogesh Simmhan' 'Neil Chue Hong'] A common feature across many science and engineering applications is the amount and diversity of data and computation that must be integrated to yield insights. Data sets are growing larger and becoming distributed; and their location, availability and properties are often…
Raphael Eidenbenz, Thomas Locher
—There is a growing demand for live, on-the-fly processing of increasingly large amounts of data. In order to ensure the timely and reliable processing of streaming data, a variety of distributed stream processing architectures and platforms have been developed, which handle the fundamental tasks of (dynamically)…
Martin Werner
This paper provides an abstract analysis of parallel processing strategies for spatial and spatio-temporal data. It isolates aspects such as data locality and computational locality as well as redundancy and locally sequential access as central elements of parallel algorithm design for spatial data. Furthermore, the…
Łukasz Świerczewski
—Paper describes the theoretical and practical aspects of the proposed model that uses distributed computing to a global network of Internet communication. Distributed computing are widely used in modern solutions such as research, where the requirement is very high processing power, which can not be placed in one…
Haotian Li
Machine learning and deep learning are novel and trending approaches to solving real-world scientific problems. Graph machine learning is dedicated to performing learning methods, such as graph neural networks, on non-Euclidean data such as graphs. Molecules, with their natural graph structures, could be analyzed by…
Sachin Lakra, Deepak Kumar Sharma
Distributed Software Development today is in its childhood and not too widespread as a method of developing software in the global IT Industry. In this context, Petrinets are a mathematical model for describing distributed systems theoretically, whereas AttNets are one of their offshoots. But development of true…
Xianzhi Cao, Chong Chen, Shiwei Li, Chang Lv + 2 more
With the explosive growth of terminal devices, scheduling massive parallel task streams has become a core challenge for distributed platforms. For computing resource providers, enhancing reliability, shortening response times, and reducing costs are significant challenges, particularly in achieving energy efficiency…
Yunhong Gu, Robert L. Grossman
Cloud computing has demonstrated that processing very large datasets over commodity clusters can be done simply, given the right programming model and infrastructure. In this paper, we describe the design and implementation of the Sector storage cloud and the Sphere compute cloud. By contrast with the existing storage…
Aneesh Khole, Atharva Thakar, Avadhoot Kulkarni, Hrithik Jadhav + 2 more
'Shreyas Shende' 'Varad Karajkhede'] Abstract— Computer systems have evolved over the years starting from sizable, single-user, slow, and expensive machines to multi-user, fast, cheaper, and small-sized machines. The use of multi-user computer networks has given rise to a new paradigm of computing known as Distributed…
Afshin Zafari, Elisabeth Larsson
In this paper, we derive and investigate approaches to dynamically load balance a distributed task parallel application software. The load balancing strategy is based on task migration. Busy processes export parts of their ready task queue to idle processes. Idle–busy pairs of processes find each other through a random…
Wilfried Yves Hamilton Adoni, Tarik Nahhal, Moez Krichen, Abdeltif El byed + 1 more
'Abdeltif El byed' 'Ismail Assayad'] Big graphs are part of the movement of “Not Only SQL” databases (also called NoSQL) focusing on the relationships between data, rather than the values themselves. The data is stored in vertices while the edges model the interactions or relationships between these data. They offer…
Jamie Alnasir, Hugh P. Shanahan
The paper reviews the use of the Hadoop platform in Structural Bioinformatics applications. Specifically, we review a number of implementations using Hadoop of high-throughput analyses, e.g. ligand-protein docking and structural alignment, and their scalability in comparison with other batch schedulers and MPI. We find…
Ivan Rodriguez-Conde, Celso Campos, Florentino Fdez-Riverola, Antonio Fernández-Caballero + 1 more
'Antonio Fernández-Caballero' 'Juan M. Corchado'] Motivated by the pervasiveness of artificial intelligence (AI) and the Internet of Things (IoT) in the current “smart everything” scenario, this article provides a comprehensive overview of the most recent research at the intersection of both domains, focusing on the…
Authors not listed
Machine learning models are transforming data-driven research across scientific disciplines, yet their deployment as accessible and reliable web services remains a significant challenge. We introduce the NERDD framework, a scalable, maintainable, and secure microservices platform designed to support the sustainable…
David Ferere, Irvin Dongo, Yudith Cardinale, Claudia Campolo
The increasing evolution of computing technologies has fostered the new intelligent concept of Ubiquitous computing (Ubicomp). Ubicomp environments encompass the introduction of new paradigms, such as Internet of Things (IoT), Mobile computing, and Wearable computing, into communication networks, which demands more…
Authors not listed
This comprehensive review examines the evolution of autonomous materials synthesis laboratories that integrate artificial intelligence with advanced robotics to accelerate discovery. Traditional materials development pipelines typically require 10-20 years, but self-driving laboratories (SDLs) and Materials…
Ben Blamey, Salman Toor, Martin Dahlö, Håkan Wieslander + 6 more
This paper introduces the HASTE Toolkit, a cloud-native software toolkit capable of partitioning data streams in order to prioritize usage of limited resources. This in turn enables more efficient data-intensive experiments. We propose a model that introduces automated, autonomous decision making in data pipelines…
Moslem Amiri, Fahad Manzoor Siddiqui, Colm Kelly, Roger Woods + 2 more
'Karen Rafferty' 'Burak Bardak'] With security and surveillance, there is an increasing need to process image data efficiently and effectively either at source or in a large data network. Whilst a Field-Programmable Gate Array has been seen as a key technology for enabling this, the design process has been viewed as…
Authors not listed
Machine olfaction—the artificial replication of the sense of smell—faces significant challenges due to the absence of large, standardized training datasets. Unlike vision, language, and audio models, which benefit from extensive corpora such as ImageNet, GLUE, and AudioSet, olfaction lacks scaled equivalents and…
Jason P. Kurs, Manuele Simi, Fabien Campagne
Computational workflows and pipelines are often created to automate series of processing steps. For instance, workflows enable one to standardize analysis for large projects or core facilities, but are also useful for individual biologists who need to perform repetitive data processing. Some workflow systems, designed…
Azza E Ahmed, Joshua M Allen, Tajesvi Bhat, Prakruthi Burra + 16 more
The changing landscape of genomics research and clinical practice has created a need for computational pipelines capable of efficiently orchestrating complex analysis stages while handling large volumes of data across heterogeneous computational environments. Workflow Management Systems (WfMSs) are the software…
Gaurav Kaushik, Sinisa Ivkovic, Janko Simonovic, Nebojsa Tijanic + 2 more
As biomedical data becomes increasingly easy to generate in large quantities, the methods used to analyze it have proliferated rapidly. However, for the insights gained from these analyses to be meaningful, the analysis methods themselves must be transparent and reproducible. To address this issue, numerous groups have…
Peter Kraus, Edan Bainglass, Francisco F. Ramirez, Enea Svaluto-Ferro + 7 more
Compliance with good research data management practices means trust in the integrity of the data, and it is achievable by a full control of the data gathering process. In this work, we demonstrate tooling which bridges these two aspects, and illustrate its use in a case study of automated battery cycling. We…