11 papers · ranked by Valyu relevance
Mohammed Alaa Ala’anzy, Nurdaulet Tolendi, Baizhan Baubek, Abdulmohsen Algarni + 1 more
Sorting can be approached in two main ways: sequentially and in parallel. In sequential sorting, data is processed in a single-threaded manner, which can be slow for large datasets. However, parallel sorting divides the task across multiple processing units, enabling faster results by processing data simultaneously.…
Bérenger Bramas, Muhammad Aleem
The way developers implement their algorithms and how these implementations behave on modern CPUs are governed by the design and organization of these. The vectorization units (SIMD) are among the few CPUs’ parts that can and must be explicitly controlled. In the HPC community, the x86 CPUs and their vectorization…
Suluk Chaikhan, Suphakant Phimoltares, Chidchanok Lursinsap, Marcin Woźniak
'Marcin Woźniak'] Tremendous quantities of numeric data have been generated as streams in various cyber ecosystems. Sorting is one of the most fundamental operations to gain knowledge from data. However, due to size restrictions of data storage which includes storage inside and outside CPU with respect to the massive…
Suluk Chaikhan, Suphakant Phimoltares, Chidchanok Lursinsap, Mohamed Hammad
'Mohamed Hammad'] Big streaming data environment concerns a complicated scenario where data to be processed continuously flow into a processing unit and certainly cause a memory overflow problem. This obstructs the adaptation of deploying all existing classic sorting algorithms because the data to be sorted must be…
Md. Shafiqul Islam, Md. Khaledur Rahman, M. Sohel Rahman
A transposition is an operation that exchanges two adjacent blocks in a permutation. A prefix transposition always moves a prefix of the permutation to another location. In this article, we use a data structure, called the permutation tree, to improve the running time of the best known approximation algorithm (with…
Sirilak Ketchaya, Apisit Rattanatranurak
Quicksort is an important algorithm that uses the divide and conquer concept, and it can be run to solve any problem. The performance of the algorithm can be improved by implementing this algorithm in parallel. In this paper, the parallel sorting algorithm named the Multi-Deque Partition Dual-Deque Merge Sorting…
Patrick Kunzmann
Alignment searches are fast heuristic methods to identify similar regions between two sequences. This group of algorithms is ubiquitously used in a myriad of software to find homologous sequences or to map sequence reads to genomes. Often the first step in alignment searches is k-mer decomposition: listing all…
Alejandro Gonzales-Irribarren
The General Transfer Format (GTF) is a widely used format for gene annotation data, integral to various downstream analyses. Efficient management and sorting of GTF data are relevant, as unsorted data can intere with computational efficiency and interpretability. Sorted GTF data, on the other hand, enable more…
Heng Li, Jiazhen Rong
We present bedtk, a new toolkit for manipulating genomic intervals in the BED format. It supports sorting, merging, intersection, subtraction and the calculation of the breadth of coverage. Bedtk employs implicit interval tree, a new data structure for fast interval overlap queries. It is several to tens of times…
Md. Khaledur Rahman, M. Sohel Rahman
The genome rearrangement problem computes the minimum number of operations that are required to sort all elements of a permutation. A block-interchange operation exchanges two blocks of a permutation which are not necessarily adjacent and in a prefix block-interchange, one block is always the prefix of that…
Alessio Campanelli, Giulio Ermanno Pibiri, Jason Fan, Rob Patro
We describe lossless compressed data structures for the colored de Bruijn graph (or, c-dBG). Given a collection of reference sequences, a c-dBG can be essentially regarded as a map from k-mers to their color sets. The color set of a k-mer is the set of all identifiers, or colors, of the references that contain the…