24 papers · ranked by Valyu relevance
Ling Chen, Houming Wu, Wenjie Yu
Pipeline parallelism is essential for large-scale model training, but existing asynchronous approaches often degrade convergence due to parameter mismatch between forward and backward passes. We propose Asynchronous Multi-Directional Pipeline parallelism (AMDP) to mitigate this issue while sustaining high utilization.…
Ari Peden-Asarch, Meredith Weinstock, Kevin R. Coffey, John F. Neumaier
Miniscope calcium imaging provides a unique window into the activity of neurons during behavior while enabling spatial localization of individual cells across time. Despite its potential to revolutionize in vivo imaging alongside the rise of optogenetic tools, miniscopes remain underutilized. This gap may stem from the…
Erwan Tanguy-Legac, Tommaso Belvedere, Gianluca Corsini, Marco Tognon + 1 more
—Accurately controlling a robotic system in real time is a challenging problem. To address this, the robotics community has adopted various algorithms, such as Model Predictive Control (MPC) and Model Predictive Path Integral (MPPI) control. The first is difficult to implement on non-linear systems such as unmanned…
Ge Zhang
bcftools is the standard toolkit for handling VCF and BCF variant files, but it processes records on a single core; its --threads option speeds up only compression of the output, not the work done on variant records. Processing large call sets is therefore slow, and users often divide the genome and reassemble the…
Ruitao Liu, Xinyang Tian, Shuo Chen, Tingrui Zhang + 3 more
Pipeline parallelism is a key technique for scaling large-model training, but modern workloads exhibit runtime variability in computation and communication. Existing pipeline systems typically consume static, profiled, or adaptively generated schedules as pre-committed execution orders. When realized task readiness…
Authors not listed
Recent advances in machine learning force fields (MLFF) have significantly extended the reach of atomistic simulations. Continuous progress in this field requires reliable reference datasets, accurate MLFF architectures, and efficient active learning strategies to enable robust modeling of complex molecular and…
Rob Patro, Siddhant Bharti, Prajwal Singhania, Rakrish Dhakal + 2 more
The FASTQ file format is the lingua franca of primary data distribution and processing across most of bioinformatics. Over time, the compression, storage, transmission, and decompression of gzip compressed fastq.gz files has become a substantial scalability bottleneck in the modern world of fast and massively parallel…
Tingkai Liu, Muralidhar Andoorveedu, Sanjoy Das, Sanjay Patel + 1 more
The evolution of compute infrastructure has transformed multi-GPU systems into tightly integrated shared-memory structures. However, current software still mostly treats these coherent interconnects simply as high-speed networks. Simultaneously, the demand for serving Large Language Models under latency constraints has…
Mohamed Salim Nasser Al Hinai, Zakira Naureen, Syed Abdullah Gilani
Since, Bioinformatic research is getting attraction that can be seen with sudden increase in development of tools as well as publications. However, there are challenges in bioinformatics which are making the results difficult to obtain while facing reproducibility, scalability, and accessibility of computational…
Eric Simon, Renato B. Hoffmann, Lucas Alf, Dalvan Griebler
This paper introduces LOG.io, a comprehensive solution designed for correct rollback recovery and fine-grain data lineage capture in distributed data pipelines. It is tailored for serverless scalable architectures and uses a log-based rollback recovery protocol. LOG.io supports a general programming model…
Mohammed Alaa Ala’anzy, Nurdaulet Tolendi, Baizhan Baubek, Abdulmohsen Algarni + 1 more
Sorting can be approached in two main ways: sequentially and in parallel. In sequential sorting, data is processed in a single-threaded manner, which can be slow for large datasets. However, parallel sorting divides the task across multiple processing units, enabling faster results by processing data simultaneously.…
Authors not listed
Background: Pharmaceutical batch scheduling in multi-reactor configurations presents complex optimization challenges under operational uncertainty, yet limited research addresses how parallel processing capacity affects heuristic performance and predictive modeling. Objectives: This study investigated scheduling…
Bahman Arasteh, Seyed Salar Sefati, Huseyin Kusetogullari, Farzad Kiani + 3 more
Efficient task scheduling remains a key challenge in High-Performance Computing and Internet of Things (IoT) systems, where the sequential execution of nested loops often limits parallelism. This paper proposes a hybrid approach that dynamically parallelizes nested loops in heterogeneous IoT environments. The suggested…
Benjamin C. Perry, Jeonghyeon Kim, Philip A. Romero
AlphaFold 3 (AF3) enables accurate biomolecular modeling but is limited by slow, CPU-bound multiple sequence alignment (MSA) generation. We introduce AlphaFast, a drop-in framework that integrates GPU-accelerated MMseqs2 sequence search to remove this bottleneck. AlphaFast achieves a 68.5× speedup in MSA construction…
Authors not listed
Bayesian optimization (BO) has become increasingly important for experimental optimization across scientific domains, yet implementing BO pipelines requires significant programming expertise and familiarity with specialized frameworks. This creates a barrier for domain experts who could benefit from BO but lack the…
Authors not listed
This comprehensive review examines the evolution of autonomous materials synthesis laboratories that integrate artificial intelligence with advanced robotics to accelerate discovery. Traditional materials development pipelines typically require 10-20 years, but self-driving laboratories (SDLs) and Materials…
John Kruper, Ariel Rokem
Tractography based on diffusion-weighted MRI (dMRI) is the predominant in vivo method for mapping the brain’s white matter. However, it is also one of the most computationally demanding steps in neuroimaging data analysis-requiring the generation and filtering of millions of streamlines per subject. Over the past…
Harikrishna Tummalapalli, Christine M. Simpson, Riccardo Balin, Vitali A. Morozov + 3 more
Scientific computing is increasingly shifting from monolithic applications to coupled simulation-AI workflows composed of highly heterogeneous tasks with diverse hardware, scale, and runtime requirements. As these workflows scale to leadership-class systems, the resulting extreme ensemble sizes and task variability can…
Sangjin Lee, Sunggon Kim, Yongseok Son, Agbotiname Lucky Imoize
We propose ScaleDefrag, a parallel and asynchronous defragmentation tool that reduces defragmentation time by up to 3.8× compared to e4defrag, while improving scalability on multi-core systems. Flash-based solid-state drives (SSDs) have been widely adopted in various large-scale storage systems including cloud and HPC…
Bishwa Ghimire, Nicholas Booth, Tapio Lönnberg, Tero Aittokallio + 1 more
Nextpie provides improved insights in terms of computational resource usage for pipeline developers and system administrators by providing comprehensive aggregated visualizations. The wide range of deployment methods offers flexibility and, when deployed using a prebuilt Docker image, it offers a low threshold for…
Authors not listed
Modeling of chemical reactions is essential for understanding kinetic mechanisms and predicting possible outcomes of reacting systems. Quantum mechanical calculations are accurate but often prohibitively expensive. Deep learning has emerged as a faster alternative, but progress is slowed by a fragmented software…
Zhixiang Liu, Wentao Hu, Chunxue Xie, Lei Zhang
Pipeline installation robots install pipelines on both sides of mine roadways using robotic arms; the geometric shape and length of the pipelines affect the movement trajectory of the robotic arms, and the working environment is complex and changeable. Aiming at the issues of the mechanical arm’s long motion…
Yuma Osako, Aineias Arango, Toshitake Asabuki
Animals flexibly combine learned behaviors into novel actions without practicing their combinations, yet the computational mechanisms that enable independently acquired computations to be expressed in parallel remain unclear. Here we show that feedback geometry during learning determines whether recurrent dynamics can…
Authors not listed
Self-driving laboratories (SDLs) promise accelerated scientific discovery and product development by closing the loop between robotic execution and AI/ML-driven decision making. In practice, however, SDL orchestration remains fragmented; workflows are typically encoded as laboratory-specific scripts or bespoke…