Search · four archives
Search · four archives
19 papers · ranked by Valyu relevance
Ana H. M. P. Tavares, Armando J. Pinho, Raquel M. Silva, João M. O. S. Rodrigues + 3 more
'João M. O. S. Rodrigues' 'Carlos A. C. Bastos' 'Paulo J. S. G. Ferreira' 'Vera Afreixo'] We address the problem of discovering pairs of symmetric genomic words (i.e., words and the corresponding reversed complements) occurring at distances that are overrepresented. For this purpose, we developed new procedures to…
Ana Tavares, Jakob Raymaekers, Peter J. Rousseeuw, Raquel M. Silva + 4 more
'Carlos A. C. Bastos' 'Armando J. Pinho' 'Paula Brito' 'Vera Afreixo'] In this work we study reverse complementary genomic word pairs in the human DNA, by comparing both the distance distribution and the frequency of a word to those of its reverse complement. Several measures of dissimilarity between distance…
G. Nallappa Bhavithran, R. Selvakumar
The biggest challenge when using DNA as a storage medium is maintaining its stability. The relative occurrence of Guanine (G) and Cytosine (C) is essential for the longevity of DNA. In addition to that, reverse complementary base pairs should not be present in the code. These challenges are overcome by a proper choice…
Jens Zentgraf, Sven Rahmann
Motivation Short DNA sequences of length k that appear in a single location (e.g., at a single genomic position, in a single species from a larger set of species, etc.) are called unique k-mers. They are useful for placing sequenced DNA fragments at the correct location without computing alignments and without…
Alon Kafri, Benny Chor, David Horn
Background Inversion Symmetry is a generalization of the second Chargaff rule, stating that the count of a string of k nucleotides on a single chromosomal strand equals the count of its inverse (reverse-complement) k-mer. It holds for many species, both eukaryotes and prokaryotes, for ranges of k which may vary from 7…
Fabian Klötzl
Bits, nucleotides and speed.
Virginia Iannibelli, Isabella Caranzano, Giovanni Birolo, Cesare Rollo + 5 more
Codon usage bias is a central record of mutation, selection, drift, and translational constraints, but it is usually treated separately from generalized Chargaff symmetry, the tendency for words and their reverse complements to occur at similar frequencies in long DNA sequences. Here we ask whether codon usage contains…
Shakir Ali, Amal S. Alali, Mohd Azeem, Atif Ahmad Khan + 2 more
Let e be a fixed positive integer and $n_{1},n_{2}$ be odd positive integers. The main objective of this article is to investigate the algebraic structure of double cyclic codes of length $(n_{1},n_{2})$ over the finite chain ring $(Re = F_{4e}+vF_{4e})$, where $v2=0$. Building upon this structural framework, we…
Mohammad Saifur Rahman, Ali Alatabbi, Tanver Athar, Maxime Crochemore + 1 more
'Maxime Crochemore' 'M. Sohel Rahman'] Background An absent word with respect to a sequence is a word that does not occur in the sequence as a factor; an absent word is minimal if all its factors on the other hand occur in that sequence. In this paper we explore the idea of using minimal absent words (MAW) to compute…
Ryosuke Yamano, Tetsuo Shibuya
The Shortest Common Superstring (SCS) problem asks for the shortest string that contains each of a given set of strings as a substring. Its reverse-complement variant, the Shortest Common Superstring problem with Reverse Complements (SCS-RC), naturally arises in bioinformatics applications, where for each input string…
Adrian Korban, Serap Şahinkaya, Deniz Üstün
In this paper, we give a matrix construction method for designing DNA codes that come from group matrix rings. We show that with our construction one can obtain reversible Gk -codes of length kn, where k, n ∈ N, over the finite commutative Frobenius ring R. We employ our construction method to obtain many DNA codes…
Roland Wittler
To index or compare sequences efficiently, often k-mers, i.e., substrings of fixed length k, are used. For efficient indexing or storage, k-mers are often encoded as integers, e.g., applying some bijective mapping between all possible σ^k^ k-mers and the interval [0, σ^k^ −1], where σ is the alphabet size. In many…
Avanti Shrikumar, Peyton Greenside, Anshul Kundaje
Deep learning approaches that have produced breakthrough predictive models in computer vision, speech recognition and machine translation are now being successfully applied to problems in regulatory genomics. However, deep learning architectures used thus far in genomics are often directly ported from computer vision…
Sukhamoy Pattanayak, Abhay Kumar Singh
In this paper, we develop the theory for constructing DNA cyclic codes of odd length over R = Z 4 [ u ] / h u 2 − 1 i based on the deletion distance. Firstly, we relate DNA pairs with a special 16 elements of ring R. Cyclic codes of odd length over R satisfy the reverse constraint and the reverse-complement constraint…
Abdullah Dertli, Yasemin Çengellenmiş
The structures of cyclic DNA codes of odd length over the finite rings R = Z4 + wZ4, w 2 = 2 and S = Z4 + wZ4 + vZ4 + wvZ4, w2 = 2, v2 = v, wv = vw are studied. The links between the elements of the rings R, S and 16 and 256 codons are established, respectively. Cyclic codes of odd length over the finite ring R…
Krishna Gopal Benerjee, Manish Gupta
DNA strings and their properties are widely studied since last 20 years due to its applications in DNA computing. In this area, one designs a set of DNA strings (called DNA code) which satisfies certain thermodynamic and combinatorial constraints such as reverse constraint, reverse-complement constraint, -content…
Shibsankar Das, Krishna Gopal Benerjee, Adrish Banerjee
—In this paper, we present a novel design strategy of DNA codes with length 3n over the non-chain ring R = Z4 + uZ4 + u 2Z4 with 64 elements and u 3 = 1, where n denotes the length of a code over R. We first study and analyze a distance conserving map defined over the ring R into the length-3 DNA sequences. Then, we…
Authors not listed
A framework for catalysis based on categorical aperture selection rather than temporal acceleration is presented. Traditional catalysis theory describes catalysts as agents that accelerate reactions by lowering activation energies, implicitly treating time as the fundamental variable and reaction rate enhancement as…
Sarwan Ali, Haris Mansoor, Prakash Chourasia, Imdad Ulla Khan + 1 more
The analysis of large volumes of molecular (genomic, proteomic, etc.) sequences has become a significant research field, especially after the recent coronavirus pandemic. Although it has proven beneficial to sequence analysis, machine learning (ML) is not without its difficulties, particularly when the feature space…