13 papers · ranked by Valyu relevance
Jose L Figueroa, Richard Allen White
We now exist in the era of massive datasets from genomics, large language models, and all the known knowledge of humanity right at our fingertips. Much of this data is becoming more accessible; however, processing such data remains an ongoing issue across systems including high performance computing (HPC)…
Haymo Kutschbach
This work introduces a self-optimizing virtual processor (VP) for numerical array programs that shifts parallelization from a manual developer task to a cooperative, agent-like runtime mechanism. Instead of relying on centralized task-graph scheduling, static compiler optimization, or explicitly annotated parallel…
Stavrinides, Georgios L., Karatza, Helen D.
With the explosive growth of big data, workloads tend to get more complex and computationally demanding. Such applications are processed on distributed interconnected resources that are becoming larger in scale and computational capacity. Data-intensive applications may have different degrees of parallelism and must…
Ge Zhang
bcftools is the standard toolkit for handling VCF and BCF variant files, but it processes records on a single core; its --threads option speeds up only compression of the output, not the work done on variant records. Processing large call sets is therefore slow, and users often divide the genome and reassemble the…
G. Kandemir, D. H. Duncan, D. van Moorselaar, J. Theeuwes
For almost half a century, target-distractor similarity has been known to induce different visual search modes. When a target is highly salient, it can pop out, suggesting parallel processing of all items irrespective of set size. By contrast, high similarity among items requires item-by-item comparison with an…
Tasmia Jannat, Michael Gowanlock, Satish Puri
The growing volume of data in scientific domains has made spatial query processing increasingly challenging due to high data transfer costs across the memory hierarchy and limited memory bandwidth. To address these bottlenecks and reduce the energy consumed on data movement, this work explores Processing-in-Memory…
Ran Ginosar
I have greatly enjoyed spending many years in studying parallel computing. My journey goes thorough MP-C, PLURAL, Async Plural, HAL, RC64 and more. As a PhD student at Princeton I studied a combination of shared memory and message passing, motivated by algorithms and the ease of programming. While at the Technion, a…
Kenshin Obi, Ryo Yoshinaka, Fujimoto, Hiroshi + 1 more
—In recent years, the complexity and scale of embedded systems, especially in the rapidly developing field of autonomous driving systems, have increased significantly. This has led to the adoption of software and hardware approaches such as Robot Operating System (ROS) 2 and multi-core processors. Traditional manual…
Farzad Razi, Mehran Moghadam, Sercan Aygun, M. Hassan Najafi + 1 more
Today's high-performance architectures are increasingly constrained by data movement latency and energy overhead, as the slowdown of single-core performance scaling coincides with the rise of highly data-intensive workloads. In-memory architectures have emerged as a complementary solution to conventional von Neumann…
Rob Patro, Siddhant Bharti, Prajwal Singhania, Rakrish Dhakal + 2 more
The FASTQ file format is the lingua franca of primary data distribution and processing across most of bioinformatics. Over time, the compression, storage, transmission, and decompression of gzip compressed fastq.gz files has become a substantial scalability bottleneck in the modern world of fast and massively parallel…
Drew E. Winters
Studying flexible, adaptive transitions between cognitive tasks and serial-parallel processing under changing task demands has been a central focus for understanding human cognition. Advances in neuroimaging analysis have improved the ability to link cognition with brain function, providing a foundation for developing…
Rubén Langarita, Jesús Alastruey-Benedé, Pablo Ibáñez, Santiago Marco‐Sola + 2 more
—Multiple HPC applications are often bottlenecked by compute-intensive kernels implementing complex dependency patterns (data-dependency bound). Traditional general-purpose accelerators struggle to effectively exploit fine-grain parallelism due to limitations in implementing convoluted data-dependency patterns (like…
Noam Teyssier, Alexander Dobin
Single-cell genomics is rapidly scaling toward billion-cell atlases, but computational analysis has become a critical bottleneck. Processing multiplexed datasets with existing tools requires substantial computational resources and runtime that become prohibitive at scale. Here we present cyto, an ultra highthroughput…