24 papers · ranked by Valyu relevance
Elif Eser, Tolga Can, Hakan Ferhatosmanoğlu, Manuela Helmer-Citterich
'Manuela Helmer-Citterich'] Sequence similarity tools, such as BLAST, seek sequences most similar to a query from a database of sequences. They return results significantly similar to the query sequence and that are typically highly similar to each other. Most sequence analysis tasks in bioinformatics require an…
Vaitea Opuu
Machine learning (ML) methods for proteins and RNAs rely on multiple sequence alignments (MSAs) and related datasets such as experimental mutagenesis libraries, yet the amount of usable information they contain remains unclear. Here, a spectral measure of information is recast into an interpretable quantity for MSAs…
Kaitlin Chaung, Tavor Baharav, Ivan Zheludev, Julia Salzman
We present a unifying statistical formulation for many fundamental problems in genome science and develop a reference-free, highly efficient algorithm that solves it. Sequence diversification – nucleic acid mutation, rearrangement, and reassortment – is necessary for the differentiation and adaptation of all…
Gavin Huttley, Katherine Caley, Robert McArthur
The algorithms required for phylogenetics — multiple sequence alignment and phylogeny estimation — are both compute intensive. As the size of DNA sequence datasets continues to increase, there is a need for a tool that can effectively lessen the computational burden associated with this widely used analysis.…
Cory D. Dunn
Phylogenetic analyses can take advantage of multiple sequence alignments as input. These alignments typically consist of homologous nucleic acid or protein sequences, and the inclusion of outlier or aberrant sequences can compromise downstream analyses. Here, I describe a program, SequenceBouncer, that uses the Shannon…
Authors not listed
Large language models (LLMs) have shown promising potential across diverse chemistry tasks, including forward reaction prediction, retrosynthesis, and property prediction. However, their ability to capture the intrinsic chemistry of molecules remains unclear. To study this, we evaluate the consistency of…
Yihuang Kang, Vladimir Zadorozhny
Process Monitoring involves tracking a system's behaviors, evaluating the current state of the system, and discovering interesting events that require immediate actions. In this paper, we consider monitoring temporal system state sequences to help detect the changes of dynamic systems, check the divergence of the…
Chris Cundy, Stefano Ermon
In many domains, autoregressive models can attain high likelihood on the task of predicting the next observation. However, this maximum-likelihood (MLE) objective does not necessarily match a downstream use-case of autoregressively generating high-quality sequences. The MLE objective weights sequences proportionally to…
Tomohiro Nishiyama
Divergences are quantities that measure discrepancy between two probability distributions and play an important role in various fields such as statistics and machine learning. Divergences are non-negative and are equal to zero if and only if two distributions are the same. In addition, some important divergences such…
Marcelo Losada, Víctor A. Penas, Federico Holik, Pedro W. Lamberti
The Jensen-Shannon divergence has been successfully applied as a segmentation tool for symbolic sequences, that is to separate the sequence into subsequences with the same symbolic content. In this work, we propose a method, based on the the Jensen-Shannon divergence, for segmentation of what we call quantum generated…
K V Harsha, Jithin Ravi, Tobias Koch
—In two-sampling testing, one observes two independent sequences of independent and identically distributed random variables distributed according to the distributions P 1 and P 2 and wishes to decide whether P 1 = P 2 (null hypothesis) or P 1 ̸= P 2 (alternative hypothesis). The Gutman test for this problem compares…
Nicola De Maio, Alexander V. Alekseyenko, William J. Coleman-Smith, Fabio Pardi + 4 more
Many important applications in bioinformatics, including sequence alignment and protein family profiling, employ sequence weighting schemes to mitigate the effects of non-independence of homologous sequences and under- or over-representation of certain taxa in a dataset. These schemes aim to assign high weights to…
Valmir C. Barbosa
The COI mitochondrial gene is present in all animal phyla and in a few others, and is the leading candidate for species identification through DNA barcoding. Calculating a generalized form of total correlation on publicly available data on the gene yields distinctive information-theoretic descriptors of the phyla…
Authors not listed
Sequence is the critical determinant of macromolecular function, yet current polymer design approaches often optimize monomer composition and ratios while ignoring sequence. This creates poorly defined design spaces for active learning that miss the vast combinatorial landscape of sequence possibilities. We introduce…
Jie Ren, Xin Bai, Yang Young Lu, Kujin Tang + 3 more
'Gesine Reinert' 'Fengzhu Sun'] Genome and metagenome comparisons based on large amounts of next generation sequencing (NGS) data pose significant challenges for alignment-based approaches due to the huge data size and the relatively short length of the reads. Alignment-free approaches based on the counts of word…
Zeehasham Rasheed, Huzefa Rangwala, Daniel Barbará
Background Advances in biotechnology have changed the manner of characterizing large populations of microbial communities that are ubiquitous across several environments."Metagenome" sequencing involves decoding the DNA of organisms co-existing within ecosystems ranging from ocean, soil and human body. Several…
Adrià Antich, Creu Palacín, Xavier Turon, Owen S. Wangensteen + 1 more
'Joseph Gillespie'] DNA metabarcoding is broadly used in biodiversity studies encompassing a wide range of organisms. Erroneous amplicons, generated during amplification and sequencing procedures, constitute one of the major sources of concern for the interpretation of metabarcoding results. Several denoising programs…
Andrzej Zielezinski, Hani Z. Girgis, Guillaume Bernard, Chris-Andre Leimeister + 15 more
Alignment-free (AF) sequence comparison is attracting persistent interest driven by data-intensive applications. Hence, many AF procedures have been proposed in recent years, but a lack of a clearly defined benchmarking consensus hampers their performance assessment. Here, we present a community resource…
Xavier F. Cadet, Reda Dehak, Sang Peter Chin, Miloud Bessafi
The nature of changes involved in crossed-sequence scale and inner-sequence scale is very challenging in protein biology. This study is a new attempt to assess with a phenomenological approach the non-stationary and nonlinear fluctuation of changes encountered in protein sequence. We have computed fluctuations from an…
Fabio Pardi, Nick Goldman, Leonid Kruglyak
Several projects investigating genetic function and evolution through sequencing and comparison of multiple genomes are now underway. These projects consume many resources, and appropriate planning should be devoted to choosing which species to sequence, potentially involving cooperation among different sequencing…
Alexander Solovyov, W Ian Lipkin
Background Many problems in computational biology require alignment-free sequence comparisons. One of the common tasks involving sequence comparison is sequence clustering. Here we apply methods of alignment-free comparison (in particular, comparison using sequence composition) to the challenge of sequence clustering.…
Babu Bassa
In this communication the author describes a software tool named "ChameleonSort". The software program, developed by the present author is useful in the sorting of biological sequence variants like those accumulating mutations while diverging from the common ancestors. Examples include viral protein variants, protein…
Authors not listed
Cyclic peptides become attractive therapeutic candidates due to their diverse biological activities. However, existing deep learning-based sequence design models, such as ProteinMPNN, are primarily optimized using cross-entropy loss and often overlook the unique topological constraints of cyclic peptides. This limits…
Samuel Renaud, Rachael Mansbach
Current antibacterial treatments cannot overcome the rapidly growing resistance of bacteria to antibiotic drugs, and novel treatment methods are required. One option is the development of new antimicrobial peptides (AMPs), to which bacterial resistance build-up is comparatively slow. Deep generative models have…