Search · four archives
Search · four archives
22 papers · ranked by Valyu relevance
Andrzej Zielezinski, Susana Vinga, Jonas Almeida, Wojciech M. Karlowski
'Wojciech M. Karlowski'] Alignment-free sequence analyses have been applied to problems ranging from whole-genome phylogeny to the classification of protein families, identification of horizontally transferred genes, and detection of recombined sequences. The strength of these methods makes them particularly useful for…
Cameron M. Nugent, Sarah J. Adamowicz
Characterization of biodiversity from environmental DNA samples and bulk metabarcoding data is hampered by off-target sequences that can confound conclusions about a taxonomic group of interest. Existing methods for isolation of target sequences rely on alignment to existing reference barcodes, but this can bias…
Daniel J. van Zyl, Marcel Dunaiski, Houriiyah Tegally, Cheryl Baxter + 2 more
The rapid increase in nucleotide sequence data generated by next-generation sequencing (NGS) technologies demands efficient computational tools for sequence comparison. Alignment-based methods, such as BLAST, are increasingly overwhelmed by the scale of contemporary datasets due to their high computational demands for…
Gurjit S. Randhawa, Kathleen A. Hill, Lila Kari
MLDSP-GUI (Machine Learning with Digital Signal Processing) is an open-source, alignment-free, ultrafast, computationally lightweight, standalone software tool with an interactive Graphical User Interface (GUI) for comparison and analysis of DNA sequences. MLDSP-GUI is a general-purpose tool that can be used for a…
Gurjit S. Randhawa, Kathleen A. Hill, Lila Kari
Background Although software tools abound for the comparison, analysis, identification, and classification of genomic sequences, taxonomic classification remains challenging due to the magnitude of the datasets and the intrinsic problems associated with classification. The need exists for an approach and software tool…
Amanda Araújo Serrão de Andrade, Marco Grivet, Otávio Brustolini, Ana Tereza Ribeiro Vasconcelos + 1 more
Alignment-free feature extraction methods are well-established in comparing genomes that do not share an alignable set of common genes (). These methods assess sequence similarity without resorting to sequence alignment and overcome the limitations of well-established alignment-based methods (; ). These alignment-free…
Saeedeh Akbari Rokn Abadi, Azam Sadat Abdosalehi, Faezeh Pouyamehr, Somayyeh Koohi
'Somayyeh Koohi'] Bio-sequence comparators are one of the most basic and significant methods for assessing biological data, and so, due to the importance of proteins, protein sequence comparators are particularly crucial. On the other hand, the complexity of the problem, the growing number of extracted protein…
Gurjit S. Randhawa, Kathleen A. Hill, Lila Kari
Although methods and software tools abound for the comparison, analysis, identification, and taxonomic classification of the enormous amount of genomic sequences that are continuously being produced, taxonomic classification remains challenging. The difficulty lies within both the magnitude of the dataset and the…
Mohamed Amine Remita, Abdoulaye Baniré Diallo
—Viral sequence classification is an important task in pathogen detection, epidemiological surveys and evolutionary studies. Statistical learning methods are widely used to classify and identify viral sequences in samples from environments. These methods face several challenges associated with the nature and properties…
Deborah Galpert, Alberto Fernández, Francisco Herrera, Agostinho Antunes + 2 more
'Agostinho Antunes' 'Reinaldo Molina-Ruiz' 'Guillermin Agüero-Chapin'] Background The development of new ortholog detection algorithms and the improvement of existing ones are of major importance in functional genomics. We have previously introduced a successful supervised pairwise ortholog classification approach…
Varsha Achuthan, Deeptangi Mangsuli, Neelam Sinha, Shweta Ramdas
This paper introduces ”SVD-FCGR,” a scalable and efficient frame-work for phylogenetic analysis using Singular Value Decomposition (SVD) on Frequency Chaos Game Representation (FCGR). Unlike traditional MSA techniques, SVD-FCGR handles large datasets with lower computational complexity. It supports both genome-wide and…
Michael Höhl, Isidore Rigoutsos, Mark A. Ragan
We have developed an alignment-free method that calculates phylogenetic distances using a maximum-likelihood approach for a model of sequence change on patterns that are discovered in unaligned sequences. To evaluate the phylogenetic accuracy of our method, and to conduct a comprehensive comparison of existing…
Ivan Borozan, Stuart Watt, Vincent Ferretti
Motivation: Alignment-based sequence similarity searches, while accurate for some type of sequences, can produce incorrect results when used on more divergent but functionally related sequences that have undergone the sequence rearrangements observed in many bacterial and viral genomes. Here, we propose a…
Troy Hernandez, Jie Yang
The typical process for classifying and submitting a newly sequenced virus to the NCBI database involves two steps. First, a BLAST search is performed to determine likely family candidates. That is followed by checking the candidate families with the Pairwise Sequence Alignment tool for similar species. The submitter's…
T. Wang, Mark Herbster, Shahzad I. Mian
The International Committee on Taxonomy of Viruses (ICTV) develops, refines and maintains a universal virus taxonomy; Order is the highest taxon in the branching hierarchy of recognised viral taxa. Historically, ICTV (sub)committees have classified viruses on the basis of morphological characteristics and various other…
Jie Ren, Xin Bai, Yang Young Lu, Kujin Tang + 3 more
'Gesine Reinert' 'Fengzhu Sun'] Genome and metagenome comparisons based on large amounts of next generation sequencing (NGS) data pose significant challenges for alignment-based approaches due to the huge data size and the relatively short length of the reads. Alignment-free approaches based on the counts of word…
Jian Chen, Le Yang, Lu Li, Steve Goodison + 1 more
Sequence comparison is a fundamental problem in bioinformatics and plays a key role in a wide range of applications. Alignment-free methods provide a computationally efficient alternative to alignment-based methods for large-scale sequence analysis. Several neural network-based methods have recently been developed for…
K P Sanil Shanker, Jim Austin, Elizabeth Sherly
This paper proposes an algorithm for alignmentfree sequence comparison using Logical Match. Here, we compute the score using fuzzy membership values which generate automatically from the number of matches and mismatches. We demonstrate the method with both the artificial and real datum. The results show the uniqueness…
Gihad N. Sohsah, Ali Reza Ibrahimzada, Huzeyfe Ayaz, Ali Cakmak
Taxonomy of living organisms gains major importance in making the study of vastly heterogeneous living things easier. In addition, various fields of applied biology (e.g., agriculture) depend on classification of living creatures. Specific fragments of the DNA sequence of a living organism have been defined as DNA…
Authors not listed
In molecular machine learning, the choice of the representation of molecules can have a significant impact on model performance. However, understanding the root causes of these performance differences often proves challenging. One promising approach to explore model behavior is representational alignment, which…
Joseph Redshaw, Darren Ting, Alex Brown, Jonathan Hirst + 1 more
Antimicrobial peptides (AMPs) represent a potential solution to the growing problem of antimicrobial resistance, yet their identification through wet-lab experiments is a costly and timeconsuming process. Accurate computational predictions would allow rapid in silico screening of candidate AMPs, thereby accelerating…
Authors not listed
Machine learning models are increasingly applied to heterogeneous materials datasets spanning different synthesis routes, measurement protocols, and structural classes. Although multi-task and representation-learning approaches are commonly used to improve predictive performance, the latent representations learned by…