16 papers · ranked by Valyu relevance
Aimin Yang, Wei Zhang, Jiahao Wang, Ke Yang + 2 more
Deoxyribonucleic acid (DNA) is a biological macromolecule. Its main function is information storage. At present, the advancement of sequencing technology had caused DNA sequence data to grow at an explosive rate, which has also pushed the study of DNA sequences in the wave of big data. Moreover, machine learning is a…
Chunyan Ao, Shihu Jiao, Yansu Wang, Liang Yu + 1 more
With the rapid development of biotechnology, the number of biological sequences has grown exponentially. The continuous expansion of biological sequence data promotes the application of machine learning in biological sequences to construct predictive models for mining biological sequence information. There are many…
Hemalatha Gunasekaran, K. Ramalakshmi, A. Rex Macedo Arokiaraj, S. Deepa Kanmani + 2 more
'S. Deepa Kanmani' 'Chandran Venkatesan' 'C. Suresh Gnana Dhas'] In a general computational context for biomedical data analysis, DNA sequence classification is a crucial challenge. Several machine learning techniques have used to complete this task in recent years successfully. Identification and classification of…
Antonino Fiannaca, Laura La Paglia, Massimo La Rosa, Giosue’ Lo Bosco + 4 more
'Giosue’ Lo Bosco' 'Giovanni Renda' 'Riccardo Rizzo' 'Salvatore Gaglio' 'Alfonso Urso'] Background An open challenge in translational bioinformatics is the analysis of sequenced metagenomes from various environmental samples. Of course, several studies demonstrated the 16S ribosomal RNA could be considered as a barcode…
Eduardo Corel, Florian Pitschi, Ivan Laprevotte, Gilles Grasseau + 2 more
'Gilles Didier' 'Claudine Devauchelle'] Background While multiple alignment is the first step of usual classification schemes for biological sequences, alignment-free methods are being increasingly used as alternatives when multiple alignments fail. Subword-based combinatorial methods are popular for their low…
Hao Xiong, Daniel Capurso, Śaunak Sen, Mark R. Segal + 1 more
Most existing methods for sequence-based classification use exhaustive feature generation, employing, for example, all -mer patterns. The motivation behind such (enumerative) approaches is to minimize the potential for overlooking important features. However, there are shortcomings to this strategy. First, practical…
Gundolf Schenk, Thomas Margraf, Andrew E Torda
Background Protein structure alignments are usually based on very different techniques to sequence alignments. We propose a method which treats sequence, structure and even combined sequence + structure in a single framework. Using a probabilistic approach, we calculate a similarity measure which can be applied to…
Giulia Fiscon, Emanuel Weitschek, Eleonora Cella, Alessandra Lo Presti + 7 more
'Alessandra Lo Presti' 'Marta Giovanetti' 'Muhammed Babakir-Mina' 'Marco Ciotti' 'Massimo Ciccozzi' 'Alessandra Pierangeli' 'Paola Bertolazzi' 'Giovanni Felici'] Background Continuous improvements in next generation sequencing technologies led to ever-increasing collections of genomic sequences, which have not been…
Michal Ziemski, Treepop Wisanwanichthan, Nicholas A. Bokulich, Benjamin D. Kaehler
'Benjamin D. Kaehler'] Naive Bayes classifiers (NBC) have dominated the field of taxonomic classification of amplicon sequences for over a decade. Apart from having runtime requirements that allow them to be trained and used on modest laptops, they have persistently provided class-topping classification accuracy. In…
Saeedeh Akbari Rokn Abadi, Amirhossein Mohammadi, Somayyeh Koohi
Background The prevalence of the COVID-19 disease in recent years and its widespread impact on mortality, as well as various aspects of life around the world, has made it important to study this disease and its viral cause. However, very long sequences of this virus increase the processing time, complexity of…
Pavel Kuksa, Vladimir Pavlovic
Background In this work we consider barcode DNA analysis problems and address them using alternative, alignment-free methods and representations which model sequences as collections of short sequence fragments (features). The methods use fixed-length representations (spectrum) for barcode sequences to measure…
Ritwika Das, Anil Rai, Dwijesh Chandra Mishra, Juan Francisco Martín
Fungal species identification from metagenomic data is a highly challenging task. Internal Transcribed Spacer (ITS) region is a potential DNA marker for fungi taxonomy prediction. Computational approaches, especially deep learning algorithms, are highly efficient for better pattern recognition and classification of…
Abhigyan Nath, Karthikeyan Subbiah
To counter the host RNA silencing defense mechanism, many plant viruses encode RNA silencing suppressor proteins. These groups of proteins share very low sequence and structural similarities among them, which consequently hamper their annotation using sequence similarity-based search methods. Alternatively the machine…
Ting Lin, Miao Wang, Min Yang, Xu Yang + 3 more
With the exponential growth of data, solving classification or regression tasks by mining time series data has become a research hotspot. Commonly used methods include machine learning, artificial neural networks, and so on. However, these methods only extract the continuous or discrete features of sequences, which…
Caroline König, Martha I Cárdenas, Jesús Giraldo, René Alquézar + 1 more
'Alfredo Vellido'] Background The characterization of proteins in families and subfamilies, at different levels, entails the definition and use of class labels. When the adscription of a protein to a family is uncertain, or even wrong, this becomes an instance of what has come to be known as a label noise problem.…
Rhydon Jackson, Debra Knisley, Cecilia McIntosh, Phillip Pfeiffer
Machine learning was applied to a challenging and biologically significant protein classification problem: the prediction of avonoid UGT acceptor regioselectivity from primary sequence. Novel indices characterizing graphical models of residues were proposed and found to be widely distributed among existing amino acid…