27 papers · ranked by Valyu relevance
Ha-Myung Park, Namyong Park, Sung-Hyon Myaeng, U Kang + 1 more
'Tatsuro Kawamoto'] A connected component in a graph is a set of nodes linked to each other by paths. The problem of finding connected components has been applied to diverse graph analysis tasks such as graph partitioning, graph compression, and pattern recognition. Several distributed algorithms have been proposed to…
Pooja Yadav, Sriniwas Pandey, Sraban Kumar Mohanty
Clustering is an unsupervised learning technique in which data or objects are grouped into sets based on some similarity measure. Most of the clustering algorithms assume that the main memory is infinite and can accommodate the set of patterns. In reality many applications give rise to a large set of patterns which…
Roumaissa Ghlib, Rania Bouhadouza, Faicel Hnaien
Designing compact and efficient quantum circuits that are compatible with Noisy Intermediate-Scale Quantum (NISQ) hardware remains a central challenge in quantum computing. Most existing optimization approaches rely on fidelity-based fitness functions that require computing the full unitary matrix of the circuit.…
Alexander Ulanov, Andrey Simanovsky, Manish Marwah
—Present day machine learning is computationally intensive and processes large amounts of data. It is implemented in a distributed fashion in order to address these scalability issues. The work is parallelized across a number of computing nodes. It is usually hard to estimate in advance how many nodes to use for a…
Hassan Mushtaq, Sajid Gul Khawaja, Muhammad Usman Akram, Amanullah Yasin + 3 more
Clustering is the most common method for organizing unlabeled data into its natural groups (called clusters), based on similarity (in some sense or another) among data objects. The Partitioning Around Medoids (PAM) algorithm belongs to the partitioning-based methods of clustering widely used for objects categorization…
Andrew Simmonett, Bernard Brooks, Thomas Darden
Evaluation of noncovalent electrostatic interactions is the dominant bottleneck in classical molecular dynamics simulations, and evaluation of Coulombic matrix elements similarly limits quantum mechanical self consistent field calculations. These difficulties are a result of the Coulomb operator’s slow decay, which…
Yun Xu, Wenhua Cheng, Pengyu Nie, Fengfeng Zhou + 1 more
Haplotype phasing represents an essential step in studying the association of genomic polymorphisms with complex genetic diseases, and in determining targets for drug designing. In recent years, huge amounts of genotype data are produced from the rapidly evolving high-throughput sequencing technologies, and the data…
Mahmudur Rahman Hera, David Koslicki, Conrado Martínez
With the surge in sequencing data generated from an ever-expanding range of biological studies, designing scalable computational techniques has become essential. One effective strategy to enable large-scale computation is to split long DNA or protein sequences into k-mers, and summarize large k-mer sets into compact…
Ragnar Groot Koerkamp, Igor Martayan
Because of the rapidly-growing amount of sequencing data, computing sketches of large textual datasets has become an essential preprocessing task. These sketches are typically much smaller than the input sequences, but preserve sufficient information for downstream analysis. Minimizers are an especially popular…
Anuj Sharma, Syed Mohammed Arshad Zaidi
Graphs and their traversal is becoming significant as it is applicable to various areas of mathematics, science and technology. Various problems in fields as varied as biochemistry (genomics), electrical engineering (communication networks), computer science (algorithms and computation) can be modeled as Graph…
Jamshed Khan, Laxman Dhulipala, Rob Patro
The rapid growth of genomic data over the past decade has made scalable and efficient sequence analysis algorithms, particularly for constructing de Bruijn graphs and their colored and compacted variants critical components of many bioinformatics pipelines. Colored compacted de Bruijn graphs condense repetitive…
Michael Bar-Sinai
Storing and manipulating Big Data relies on various data structures, algorithms and technologies. Some of these are new, while others have existed for quite a while (the Bloom filter was presented in 1970) and are now making their way into mainstream software engineering. We present those algorithms and technologies…
Authors not listed
We present a fast, asymptotically linear-scaling implementation of the perturbative quadruples energy correction in coupled-cluster theory using local natural orbitals. Our work follows the domain-based local pair natural orbital (DLPNO) approach previously applied to lower levels of excitations in coupled-cluster…
Pier Paolo Poier, Louis Lagardère, Jean-Philip Piquemal
We propose a new strategy to solve the Tkatchenko-Scheffler Many-Body Dispersion (MBD) model’s equations. Our approach overcomes the original O(N**3) computational complexity that limits its applicability to large molecular systems within thecontext of O(N) Density Functional Theory (DFT). First, in order to generate…
Hsiang-Huang Wu, Chien‐Min Wang, Hsuan-Chi Kuo, Wei-Chun Chung + 1 more
'Jan-Ming Ho'] Abstract—Suffix Array (SA) is a cardinal data structure in many pattern matching applications, including data compression, plagiarism detection and sequence alignment. However, as the volumes of data increase abruptly, the construction of SA is not amenable to the current large-scale data processing…
Yuanchao Zhang, Deanne M. Taylor
In single-cell RNA-seq (scRNA-seq) experiments, the number of individual cells has increased exponentially due to significant improvements on single-cell isolation and massively parallel sequencing technologies. However, computational methods have not scaled to the same order, presenting analytical challenges to the…
Pier Paolo Poir, Louis Lagardère, Jean-Philip Piquemal
We propose a new strategy to solve the Tkatchenko-Scheffler Many-Body Dispersion (MBD) model’s equations. Our approach overcomes the original O(N**3) computational complexity that limits its applicability to large molecular systems within thecontext of O(N) Density Functional Theory (DFT). First, in order to generate…
Muhammad Idris, Shujaat Hussain, Muhammad Hameed Siddiqi, Waseem Hassan + 3 more
'Waseem Hassan' 'Hafiz Syed Muhammad Bilal' 'Sungyoung Lee' 'Christophe Antoniewski'] Large quantities of data have been generated from multiple sources at exponential rates in the last few years. These data are generated at high velocity as real time and streaming data in variety of formats. These characteristics give…
Saurav Prakash, Amirhossein Reisizadeh, Ramtin Pedarsani, Salman Avestimehr
'Salman Avestimehr'] To combat the growing demands for efficient processing of large scale graph-structured datasets, many distributed graph computing systems have been developed recently. As these systems require many messages to be exchanged among computing machines at each step of the computation, communication…
Manuel Penschuck
Shuffling is the process of rearranging a sequence of elements into a random order such that any permutation occurs with equal probability. It is an important building block in a plethora of techniques used in virtually all scientific areas. Consequently considerable work has been devoted to the design and…
Tizian Schulz, Paul Medvedev
Given a sequencing read, the broad goal of read mapping is to find the location(s) in the reference genome that have a “similar sequence”. Traditionally, “similar sequence” was defined as having a high alignment score and read mappers were viewed as heuristic solutions to this well-defined problem. For sketch-based…
Authors not listed
Identifying synthesis routes from knowledge graphs poses challenges beyond retrosynthesis, including path–finding artifacts and data issues. We introduce “SynGPS”, a novel algorithm that overcomes these limitations by identifying viable routes even with common artifacts. SynGPS can resolve nonsensical cycles…
Mary Pitman, David Hahn, Gary Tresadern, David Mobley
Drug discovery is accelerated with computational methods such as alchemical simulations to estimate ligand affinities. In particular, relative binding free energy (RBFE) simulations are beneficial for lead optimization. To use RBFE simulations to compare prospective ligands in silico, researchers first plan the…
Silu Huang, Ada Wai-Chee Fu
As computer clusters are found to be highly effective for handling massive datasets, the design of efficient parallel algorithms for such a computing model is of great interest. We consider (α, k)-minimal algorithms for such a purpose, where α is the number of rounds in the algorithm, and k is a bound on the deviation…
Ricardo Villanueva-Polanco
In this paper, we will study the key enumeration problem, which is connected to the key recovery problem posed in the cold boot attack setting. In this setting, an attacker with physical access to a computer may obtain noisy data of a cryptographic secret key of a cryptographic scheme from main memory via this data…
Alex M. Ascension, Marcos J. Araúzo-Bravo
Big Data analysis is a discipline with a growing number of areas where huge amounts of data is extracted and analyzed. Parallelization in Python integrates Message Passing Interface via mpi4py module. Since mpi4py does not support parallelization of objects greater than 2^31^ bytes, we developed BigMPI4py, a Python…
Martin Werner
This paper provides an abstract analysis of parallel processing strategies for spatial and spatio-temporal data. It isolates aspects such as data locality and computational locality as well as redundancy and locally sequential access as central elements of parallel algorithm design for spatial data. Furthermore, the…