14 papers · ranked by Valyu relevance
Botao Peng, Panagiota Fatourou, Themis Palpanas
Data series similarity search is a core operation for several data series analysis applications across many different domains. Nevertheless, even state-of-the-art techniques cannot provide the time performance required for large data series collections. We propose ParIS and ParIS+, the first disk-based data series…
Carl Kugblenu, Petri Vuorimaa
Production vector search systems often fan out each query across parallel lanes (threads, replicas, or shards) to meet latency service-level objectives (SLOs). In practice, these lanes rediscover the same candidates, so extra compute does not increase coverage. We present a coordination-free lane partitioner that turns…
Thiago Teixeira, George Teodoro, Eduardo Valle, Joel Saltz
—Similarity search is critical for many database applications, including the increasingly popular online services for Content-Based Multimedia Retrieval (CBMR). These services, which include image search engines, must handle an overwhelming volume of data, while keeping low response times. Thus, scalability is…
Xiandong Meng, Yanqing Ji
This paper focuses on the latest research and critical reviews on modern computing architectures, software and hardware accelerated algorithms for bioinformatics data analysis with an emphasis on one of the most important sequence analysis applications-hidden Markov models (HMM). We show the detailed performance…
Ashish Chapagain, Dima Abuoliem, In Ho Cho, Tongbiao Wang
Multifunctional nanosurfaces receive growing attention due to their versatile properties. Capillary force lithography (CFL) has emerged as a simple and economical method for fabricating these surfaces. In recent works, the authors proposed to leverage the evolution strategies (ES) to modify nanosurface characteristics…
Eray Özkural, Cevdet Aykanat
All-pairs similarity problem asks to find all vector pairs in a set of vectors the similarities of which surpass a given similarity threshold, and it is a computational kernel in data mining and information retrieval for several tasks. We investigate the parallelization of a recent fast sequential algorithm. We propose…
Botao Peng, Panagiota Fatourou, Themis Palpanas
Data series similarity search is a core operation for several data series analysis applications across many different domains. However, the state-of-the-art techniques fail to deliver the time performance required for interactive exploration, or analysis of large data series collections. In this work, we propose MESSI…
Zeyu Xia, Canqun Yang, Chenchen Peng, Yifei Guo + 3 more
'Tao Tang' 'Yingbo Cui'] Background The advent of Single Molecule Real-Time (SMRT) sequencing has overcome many limitations of second-generation sequencing, such as limited read lengths, PCR amplification biases. However, longer reads increase data volume exponentially and high error rates make many existing alignment…
Soo Hee Han, Joon Heo, Hong Gyoo Sohn, Kiyun Yu
In this study, a parallel processing method using a PC cluster and a virtual grid is proposed for the fast processing of enormous amounts of airborne laser scanning (ALS) data. The method creates a raster digital surface model (DSM) by interpolating point data with inverse distance weighting (IDW), and produces a…
Patrizio Dazzi
Embarrassingly parallel problems are characterised by a very small amount of information to be exchanged among the parts they are split in, during their parallel execution. As a consequence they do not require sophisticated, low-latency, high-bandwidth interconnection networks but can be efficiently computed in…
Yiqiu Wang, Yan Gu, Julian Shun
The DBSCAN method for spatial clustering has received significant attention due to its applicability in a variety of data analysis tasks. There are fast sequential algorithms for DB-SCAN in Euclidean space that take ( log) work for two dimensions, sub-quadratic work for three or more dimensions, and can be computed…
Johann M Kraus, Hans A Kestler
Background In recent years, the demand for computational power in computational biology has increased due to rapidly growing data sets from microarray and other high-throughput technologies. This demand is likely to increase. Standard algorithms for analyzing data, such as cluster algorithms, need to be parallelized…
Sirilak Ketchaya, Apisit Rattanatranurak
Quicksort is an important algorithm that uses the divide and conquer concept, and it can be run to solve any problem. The performance of the algorithm can be improved by implementing this algorithm in parallel. In this paper, the parallel sorting algorithm named the Multi-Deque Partition Dual-Deque Merge Sorting…
Rajendra Purohit, K. R. Chowdhary, Sunıl Dutt Purohıt
—Arrival of multicore systems has enforced a new scenario in computing, the parallel and distributed algorithms are fast replacing the older sequential algorithms, with many challenges of these techniques. The distributed algorithms provide distributed processing using distributed file systems and processing units…