25 papers · ranked by Valyu relevance
Maleeha Najam, Raihan Ur Rasool, Hafiz Farooq Ahmad, Usman Ashraf + 1 more
'Asad Waqar Malik'] Storing and processing of large DNA sequences has always been a major problem due to increasing volume of DNA sequence data. However, a number of solutions have been proposed but they require significant computation and memory. Therefore, an efficient storage and pattern matching solution is…
Kai Blin, Wolfgang Wohlleben, Tilmann Weber
Patterns in biological sequences frequently signify interesting features in the underlying molecule. Many tools exist to search for well-known patterns. Less support is available for exploratory analysis, where no well-defined patterns are known yet. PatScanUI ([https://patscan.secondarymetabolites.org/]()) provides a…
Jinane Bazzi, Jana Sweidan, Mohammed E. Fouda, Rouwaida Kanj + 1 more
'Ahmed M. Eltawil'] Abstract—DNA pattern matching is essential for many widely used bioinformatics applications. Disease diagnosis is one of these applications, since analyzing changes in DNA sequences can increase our understanding of possible genetic diseases. The remarkable growth in the size of DNA datasets has…
J. A. M. Rexie, Kumudha Raimond, Mythily Murugaaboopathy, D. Brindha + 1 more
'Henock Mulugeta'] An area of medical science, that is, gaining prominence, is DNA sequencing. Genetic mutations responsible for the disease have been detected using DNA sequencing. The research is focusing on pattern identification methodologies for dealing with DNA-sequencing problems relating to various…
Janja Paliska Soldo, Ana Sović Kržić, Damir Seršić
This paper focuses on pattern matching in the DNA sequence. It was inspired by a previously reported method that proposes encoding both pattern and sequence using prime numbers. Although fast, the method is limited to rather small pattern lengths, due to computing precision problem. Our approach successfully deals with…
Matan Drory Retwitzer, Maya Polishchuk, Elena Churkin, Ilona Kifer + 2 more
'Ilona Kifer' 'Zohar Yakhini' 'Danny Barash'] Searching for RNA sequence-structure patterns is becoming an essential tool for RNA practitioners. Novel discoveries of regulatory non-coding RNAs in targeted organisms and the motivation to find them across a wide range of organisms have prompted the use of computational…
Christos Papalitsas, Ioannis Mouratidis, Michail Patsakis, Evangelos Stogiannos + 2 more
The exponential growth of publicly available genomic data has created unprecedented opportunities for sequence-based discovery. Locating specific k-mers is fundamental to diverse applications, including metagenomic classification, pathogen and cancer detection, and variant calling yet efficient identification of…
Will Solow, Matthew Barich, Brendan Mumey
The Exact Circular Pattern Matching (ECPM) problem consists of reporting every occurrence of a rotation of a pattern P in a text T. In many real-world applications, specifically in computational biology, circular rotations are of interest because of their prominence in virus DNA. Thus, given no restrictions on…
Suejb Memeti, Sabri Pllana
—Rapid analysis of DNA sequences is important in preventing the evolution of different viruses and bacteria during an early phase, early diagnosis of genetic predispositions to certain diseases (cancer, cardiovascular diseases), and in DNA forensics. However, real-world DNA sequences may comprise several Gigabytes and…
P Giacomelli
—In this paper we will describe a new approach on the well-known suffix-array algorithm using Big Table Data Technology. We will demonstrate how it is possible to refactor a well-known algorithm coupled by taking advantage of an highperformance distributed datastore, to illustrate the advantages of using datastore…
Konstantinos F. Xylogiannopoulos
Pattern detection and string matching are fundamental problems in computer science and the accelerated expansion of bioinformatics and computational biology have made them a core topic for both disciplines. The SARS-CoV-2 pandemic has made such problems more demanding with hundreds or thousands of new genome variants…
Konstantinos F. Xylogiannopoulos
—Exact string matching has been a fundamental problem in computer science for decades because of many practical applications. Some are related to common procedures, such as searching in files and text editors, or, more recently, to more advanced problems such as pattern detection in Artificial Intelligence and…
Yusei Kobori, Satoshi Mizuta
Graphical representation of DNA sequences is one of the most popular techniques of alignment-free sequence comparison. In this article, we propose a new method for extracting features of DNA sequences represented by binary images, in which we estimate the similarity between DNA sequences by the frequency histograms of…
Luca Renders, Lore Depuydt, Travis Gagie, Jan Fostier
Aligning sequencing reads to reference genomes is a fundamental task in bioinformatics. Aligners can be classified as lossy or lossless: lossy aligners prioritize speed by reporting only one or a few high-scoring alignments, whereas lossless aligners output all optimal alignments, ensuring completeness and sensitivity.…
Aimin Yang, Wei Zhang, Jiahao Wang, Ke Yang + 2 more
Deoxyribonucleic acid (DNA) is a biological macromolecule. Its main function is information storage. At present, the advancement of sequencing technology had caused DNA sequence data to grow at an explosive rate, which has also pushed the study of DNA sequences in the wave of big data. Moreover, machine learning is a…
P. Pandiselvam, T. Marimuthu, R. Lawrance
String matching algorithm plays the vital role in the Computational Biology. The functional and structural relationship of the biological sequence is determined by similarities on that sequence. For that, the researcher is supposed to aware of similarities on the biological sequences. Pursuing of similarity among…
Ariel Chernomoretz, Manuel Balparda, Laura La Grutta, Andres Calabrese + 3 more
GENis is an open source multi-tier information system developed to run a forensic DNA database at local, regional and national levels.^1^ It was conceived as a highly customizable system, enforcing several security policies including: data encryption, double factor identification, structure of user’s roles and…
Hongyi Xin, Jeremie Kim, Sunny Nahar, Carl Kingsford + 2 more
Approximate String Matching is a pivotal problem in the field of computer science. It serves as an integral component for many string algorithms, most notably, DNA read mapping and alignment. The improved LV algorithm proposes an improved dynamic programming strategy over the banded Smith-Waterman algorithm but suffers…
HANIYEH ABDOLLAHZADEH, Tonya Peeples, Mohammad Shahcheraghi
DNA-based nanomaterials have shown great potential in numerous applications, thanks to their unique properties including DNA's various molecular interactions, programmability, and versatility with biological modules. Meanwhile, the DNA origami platforms have shown promise in the creation of drug carriers. This…
Nicholas Tjahjono, Evgeni Penev, Boris Yakobson
Programmable self-assembly provides a promising avenue to improve upon traditional synthesis and create multi-component materials with emergent properties and arbitrary nanoscale complexity. However, its most successful realizations utilizing DNA often use complicated arduous procedures that result in low yields. Here…
Cindy Ng, Anirban Samanta, Ole Aalund Mandrup, Emily Tsang + 3 more
The folding of double-stranded DNA around histones is a central mechanism in eukaryotic cells for compacting the genetic information into chromosomes. Very few artificial methods are available for controlling the shape of dsDNA at any level, whereas several artificial methods have been developed to efficiently organize…
Jürgen Behr, Timm Michel, Maya Giridhar, Santra Santhosh + 11 more
Large- to ultra-large-scale synthesis of nucleic acids is becoming an increasingly important tool for understanding and manipulating biological systems, as well as for developing new technologies based on engineered biological materials, including DNA-based nanofabrication, aptamers and writing digital data at the…
Authors not listed
Sequence is the critical determinant of macromolecular function, yet current polymer design approaches often optimize monomer composition and ratios while ignoring sequence. This creates poorly defined design spaces for active learning that miss the vast combinatorial landscape of sequence possibilities. We introduce…
Authors not listed
Large-scale de novo nucleic acid synthesis is a powerful tool enabling researchers to better understand and engineer biological systems. Fields ranging from genomics to nucleic acid therapeutics to synthetic biology make use of high-throughput experimental approaches requiring access to large pools or libraries of DNA…
Authors not listed
DNA-encoded libraries (DELs) have emerged as an effective and efficient selection strategy for lead compound discovery in academia and industry over the past decades. Despite recent advancements in this field, DEL is still limited by sensitive DNA-based constructs, especially low selection success rates from random…