19 papers · ranked by Valyu relevance
Yun Sil Chang, Hsin‐I Wu, Ren‐Song Tsay
—We propose an effective parallel program debugging approach based on the timing annotation technique. With prevalent multi-core platforms, parallel programming is required to fully utilize the computing power. However, the nondeterminism property and the associated concurrency bugs are notorious and remain to be great…
Aleem Akhtar, Aamir Shafi, Mohsan Jameel
MPJ Express is a messaging system that allows computational scientists to write and execute parallel Java applications on High Performance Computing (HPC) hardware. Despite its successful adoption in the Java HPC community, the MPJ Express software currently does not provide any support for debugging and profiling…
God'salvation F. Oguibe, Vinodh Kumaran Jayakumar, Tongping Liu, Andrew Lan + 1 more
Concurrent programming is a core component of Computer Science curricula, yet remains notoriously difficult for students to master due to its inherent complexity and the nondeterministic nature of concurrency bugs such as deadlocks and race conditions. In this work, we present ParaView, an educational tool designed to…
Matteo Marra, Guillermo Polito, Elisa Gonzalez Boix
- b Univ. Lille, CNRS, Centrale Lille, Inria, UMR 9189 CRIStAL Centre de Recherche en Informatique Signal et Automatique de Lille, F-59000 Lille, France Abstract Context Recent studies show that developers spend most of their programming time testing, verifying and debugging software. As applications become more and…
Fábio Petrillo, Yann‐Gaël Guéhéneuc, Marcelo Soares Pimenta, Carla Dal Sasso Freitas + 1 more
'Carla Dal Sasso Freitas' 'Foutse Khomh'] Abstract One of the most important tasks in software maintenance is debugging. To start an interactive debugging session, developers usually set breakpoints in an integrated development environment and navigate through different paths in their debuggers. We started our work by…
Authors not listed
Recent advances in machine learning force fields (MLFF) have significantly extended the reach of atomistic simulations. Continuous progress in this field requires reliable reference datasets, accurate MLFF architectures, and efficient active learning strategies to enable robust modeling of complex molecular and…
Pau Andrio, Adam Hospital, Cristian Ramon-Cortes, Javier Conejero + 4 more
The usage of workflows has led to progress in many fields of science, where the need to process large amounts of data is coupled with difficulty in accessing and efficiently using High Performance Computing platforms. On the one hand, scientists are focused on their problem and concerned with how to process their data.…
Ben Langmead, Christopher Wilks, Valentin Antonescu, Rone Charles
General-purpose processors can now contain many dozens of processor cores and support hundreds of simultaneous threads of execution. To make best use of these threads, genomics software must contend with new and subtle computer architecture issues. We discuss some of these and propose methods for improving thread…
Authors not listed
With the ever-increasing demand for atomistic structures representative of real-life systems as well as the ad-vent of exascale computers, it has now become necessary and possible to use advanced global optimization (GO) techniques to intelligently sample the potential energy surface (PES). Given the previous studies…
Kecong Tang, Ahsan Sanaullah, Degui Zhi, Shaojie Zhang
Durbin’s positional Burrows-Wheeler transform (PBWT) enables algorithms with the optimal time complexity of O(MN) for reporting all vs all haplotype matches in a population panel with M haplotypes and N variant sites. However, even this efficiency may still be too slow when the number of haplotypes reaches millions. To…
Stuart Byma, Akash Dhasade, Adrian Altenhoff, Christophe Dessimoz + 1 more
This paper presents a new, parallel implementation of clustering and demonstrates its utility in greatly speeding up the process of identifying homologous proteins. Clustering is a technique to reduce the number of comparison needed to find similar pairs in a set of n elements such as protein sequences. Precise…
Sikao Guo, Nenad Korolija, Kent Milfeld, Adip Jhaveri + 3 more
Particle-based reaction-diffusion models offer a high-resolution alternative to the continuum reaction-diffusion approach, capturing the discrete and volume-excluding nature of molecules undergoing stochastic dynamics. These methods are thus uniquely capable of simulating explicit self-assembly of particles into…
Authors not listed
Protein conformational landscapes contain the functionally relevant information useful for understanding biological processes. Mapping out conformational landscapes provides valuable insights into protein behaviors and biological phenomena, and has relevance to therapeutic design. While experimental structural biology…
Manuel Carrer, Henrique Musseli Cezar, Sigbjørn Løland Bore, Morten Ledum + 1 more
We develop #-HylleraasMD (#-HyMD), a fully end-to-end differentiable molecular dynamics software based on the Hamiltonian hybrid particle-field formalism, and use it to establish a protocol for automated optimization of force field parameters. #-HyMD is templated on the recently established HylleraaasMD software, while…
Ido Ben-Shalom, Charles Lin, Brian Radak, Woody Sherman + 1 more
Molecular dynamics (MD) simulations of proteins are commonly used to sample from the Boltzmann distribution of conformational states, with wide-ranging applications spanning chemistry, biophysics, and drug discovery. However, MD can be inefficient at equilibrating water occupancy for buried cavities in proteins that…
Authors not listed
Bayesian optimization (BO) has become increasingly important for experimental optimization across scientific domains, yet implementing BO pipelines requires significant programming expertise and familiarity with specialized frameworks. This creates a barrier for domain experts who could benefit from BO but lack the…
A. Lucas Martins, Alexandre Laborde, Michael Orger
We present Sardine, a software framework built with.NET to control experimental setups through the reliable execution of dynamic networks of independent modules, where each module can interface with a hardware device (e.g., camera, motor) or represent an operation over data (e.g., image filter, data stream). The…
Azza E. Ahmed, Jacob Heldenbrand, Yan Asmann, Faisal M. Fadlelmola + 10 more
Genomic variant discovery is frequently performed using the GATK Best Practices variant calling pipeline, a complex workflow with multiple steps, fans/merges, and conditionals. This complexity makes management of the workflow difficult on a computer cluster, especially when running in parallel on large batches of data…
Peter Kraus, Edan Bainglass, Francisco F. Ramirez, Enea Svaluto-Ferro + 7 more
Compliance with good research data management practices means trust in the integrity of the data, and it is achievable by a full control of the data gathering process. In this work, we demonstrate tooling which bridges these two aspects, and illustrate its use in a case study of automated battery cycling. We…