14 papers · ranked by Valyu relevance
John Kruper, Ariel Rokem
Tractography based on diffusion-weighted MRI (dMRI) is the predominant in vivo method for mapping the brain’s white matter. However, it is also one of the most computationally demanding steps in neuroimaging data analysis-requiring the generation and filtering of millions of streamlines per subject. Over the past…
Mohammed Alaa Ala’anzy, Nurdaulet Tolendi, Baizhan Baubek, Abdulmohsen Algarni + 1 more
Sorting can be approached in two main ways: sequentially and in parallel. In sequential sorting, data is processed in a single-threaded manner, which can be slow for large datasets. However, parallel sorting divides the task across multiple processing units, enabling faster results by processing data simultaneously.…
Bahman Arasteh, Seyed Salar Sefati, Huseyin Kusetogullari, Farzad Kiani + 3 more
Efficient task scheduling remains a key challenge in High-Performance Computing and Internet of Things (IoT) systems, where the sequential execution of nested loops often limits parallelism. This paper proposes a hybrid approach that dynamically parallelizes nested loops in heterogeneous IoT environments. The suggested…
Guangfeng You, Chao Qian, Ouling Wu, Hongsheng Chen
The proliferation of deep learning applications has intensified the demand for electronic hardware with low energy consumption and fast computing speed. Neuromorphic photonics have emerged as a viable alternative to process high-throughput information at the physical space. However, the simultaneous attainment of high…
Kevin Garner, Polykarpos Thomadakis, Nikos Chrisochoides
This paper presents a distributed memory method for anisotropic mesh adaptation that is designed to avoid the use of collective communication and global synchronization techniques. In the presented method, meshing functionality is separated from performance aspects by utilizing a separate entity for each - a multicore…
Mateusz Gruzewski, Marek Palkowski, Ramon Antonio Rodriges Zalipynis
In this article, we present an efficient and concise OpenMP implementation of the Nussinov RNA folding algorithm, a well-known representative of non-serial polyadic dynamic programming (NPDP). Our goal is to develop an optimized implementation that can serve as a template for related dynamic programming applications.…
Christoph Stelz, Lukas Hübner, Alexandros Stamatakis, Russell Schwartz
Given an input array of elements $E=[e_{0},e_{1},…,e_{n-1}]$, distributed across p processing elements (PEs; e.g. processes or threads), we desire to compute $r=e_{0}⊕e_{1}⊕…⊕e_{n-1}$, where $⊕$ denotes a binary, associative operation (e.g. summation or multiplication). In a distributed reduction, we return the result…
Kecong Tang, Ardalan Naseri, Degui Zhi, Shaojie Zhang + 1 more
At each site indexed by k, the algorithm maintains a prefix array $P_{k}$, which encodes the co-lexicographic ordering of haplotypes based on their reversed prefixes up to the current site. Formally, let ${x_{i}}$, where $i=1$ to M, denote M haplotype sequences over N sites. At site k, the prefix array $P_{k}$ is…
Mohammad Abdur Rob, Md. Zakir Hossen, Md. Kamal Hossen, Md. Mithun Ali + 2 more
Sorting algorithms play a crucial role in computing, but most are designed with rigid structure that are only efficient under certain conditions. Although some sorting algorithms perform well in some circumstances, they do not perform well on some resistant platforms. This study introduces Wall-L Merge Sort, which…
Hong-Zhe Yang, Jian-Peng Dou, Feng Lu, Xiao-Wen Shang + 4 more
In-memory computing, which enables computation directly within memory, represents an efficient approach to processing massively parallel computation tasks that are intractable for conventional computers. However, implementations of in-memory computing have been primarily limited to the classical regime, with its…
Qiujiang Liang, Jun Yang
for Performant Large-Scale Ab Initio Calculations Authors: Qiujiang Liang, Jun Yang Computational acceleration of orbital-invariant local correlation methods on graphics processing units (GPUs) has remained largely unexplored due to substantial algorithmic complexities. The runtime efficiency of GPU-implemented local…
Jinli Chen, Akinsanmi S. Ige, Keith Runge, Pierre A. Deymier + 1 more
The human brain performs complex, high-dimensional (HD) computations, such as causal reasoning, counterfactual thinking, and abstraction, with ~1011 neurons while consuming ~20 watts of power. Neuromorphic computing seeks similar efficiency, but current devices face bottlenecks in bandwidth, energy, wiring, footprint…
Yong Yang, Daying Sun, Zhiyuan Ma, Wenhua Gu + 1 more
The gravity forward modeling algorithm is a compute-intensive method and is widely used in scientific computing, particularly in geophysics, to predict the impact of subsurface structures on surface gravity fields. Traditional implementations rely on CPUs, where performance gains are mainly achieved through algorithmic…
Renée A. Sirbu, Luciano Floridi
Biological computing (biocomputing) leverages biologically derived materials and processes, such as DNA and protein synthesis, to perform computational tasks. Biocomputing offers significant advantages over traditional silicon-based systems in terms of scalability, energy efficiency, computational flexibility, and…