24 papers · ranked by Valyu relevance
Aimin Yang, Wei Zhang, Jiahao Wang, Ke Yang + 2 more
Deoxyribonucleic acid (DNA) is a biological macromolecule. Its main function is information storage. At present, the advancement of sequencing technology had caused DNA sequence data to grow at an explosive rate, which has also pushed the study of DNA sequences in the wave of big data. Moreover, machine learning is a…
Zengyou He, Guangyao Xu, Chaohua Sheng, Bo Xu + 1 more
Sequence classification is an important data mining task in many real-world applications. Over the past few decades, many sequence classification methods have been proposed from different aspects. In particular, the pattern-based method is one of the most important and widely studied sequence classification methods in…
Severin Gsponer, Luca Costabello, Chan Le Van, Sumit Pai + 3 more
'Christophe Guéret' 'Georgiana Ifrim' 'Freddy Lécué'] Abstract. Sequence classification is the supervised learning task of building models that predict class labels of unseen sequences of symbols. Although accuracy is paramount, in certain scenarios interpretability is a must. Unfortunately, such trade-off is often…
Chunyan Ao, Shihu Jiao, Yansu Wang, Liang Yu + 1 more
With the rapid development of biotechnology, the number of biological sequences has grown exponentially. The continuous expansion of biological sequence data promotes the application of machine learning in biological sequences to construct predictive models for mining biological sequence information. There are many…
Mohamed Amine Remita, Ahmed Halioui, Abou Abdallah Malick Diouara, Bruno Daigle + 2 more
Advances in cloning and sequencing technology yielded a massive number of genome of virus strains. The classification and annotation of these genomes constitute important assets in the discovery of genomic variability, taxonomic characteristics and disease mechanisms. Existing classification methods are often designed…
Roman Samarev, Andrey Vasnetsov, Elizaveta Smelkova
The article deals with the issue of modification of metric classification algorithms. In particular, it studies the algorithm k-Nearest Neighbours for its application to sequential data. A method of generalization of metric classification algorithms is proposed. As a part of it, there has been developed an algorithm…
Sarwan Ali, Haris Mansoor, Prakash Chourasia, Murray Patterson
Biological sequence classification is vital in various fields, such as genomics and bioinformatics. The advancement and reduced cost of genomic sequencing have brought the attention of researchers for protein and nucleotide sequence classification. Traditional approaches face limitations in capturing the intricate…
Hemalatha Gunasekaran, K. Ramalakshmi, A. Rex Macedo Arokiaraj, S. Deepa Kanmani + 2 more
'S. Deepa Kanmani' 'Chandran Venkatesan' 'C. Suresh Gnana Dhas'] In a general computational context for biomedical data analysis, DNA sequence classification is a crucial challenge. Several machine learning techniques have used to complete this task in recent years successfully. Identification and classification of…
Hao Xiong, Daniel Capurso, Śaunak Sen, Mark R. Segal + 1 more
Most existing methods for sequence-based classification use exhaustive feature generation, employing, for example, all -mer patterns. The motivation behind such (enumerative) approaches is to minimize the potential for overlooking important features. However, there are shortcomings to this strategy. First, practical…
Gundolf Schenk, Thomas Margraf, Andrew E Torda
Background Protein structure alignments are usually based on very different techniques to sequence alignments. We propose a method which treats sequence, structure and even combined sequence + structure in a single framework. Using a probabilistic approach, we calculate a similarity measure which can be applied to…
Giulia Fiscon, Emanuel Weitschek, Eleonora Cella, Alessandra Lo Presti + 7 more
'Alessandra Lo Presti' 'Marta Giovanetti' 'Muhammed Babakir-Mina' 'Marco Ciotti' 'Massimo Ciccozzi' 'Alessandra Pierangeli' 'Paola Bertolazzi' 'Giovanni Felici'] Background Continuous improvements in next generation sequencing technologies led to ever-increasing collections of genomic sequences, which have not been…
Murilo Horacio Pereira da Cruz, Douglas Silva Domingues, Priscila Tiemi Maeda Saito, Alexandre Rossi Paschoal + 1 more
Transposable elements (TEs) are the most represented sequences occurring in eukaryotic genomes. They are capable of transpose and generate multiple copies of themselves throughout genomes. These sequences can produce a variety of effects on organisms, such as regulation of gene expression. There are several types of…
Suprativ Saha, Rituparna Chaki
Data mining techniques have been used by researchers for analyzing protein sequences. In protein analysis, especially in protein sequence classification, selection of feature is most important. Popular protein sequence classification techniques involve extraction of specific features from the sequences. Researchers…
Saeedeh Akbari Rokn Abadi, Amirhossein Mohammadi, Somayyeh Koohi
Background The prevalence of the COVID-19 disease in recent years and its widespread impact on mortality, as well as various aspects of life around the world, has made it important to study this disease and its viral cause. However, very long sequences of this virus increase the processing time, complexity of…
Suprativ Saha
Protein sequence classification involves feature selection for accurate classification. Popular protein sequence classification techniques involve extraction of specific features from the sequences. Researchers apply some well-known classification techniques like neural networks, Genetic algorithm, Fuzzy ARTMAP, Rough…
Ananya Basu, Suprativ Saha
In post genomic era with the advent of new technologies a huge amount of complex molecular data are generated with high throughput. The management of this biological data is definitely a challenging task due to complexity and heterogeneity of data for discovering new knowledge. Issues like managing noisy and incomplete…
Gihad N. Sohsah, Ali Reza Ibrahimzada, Huzeyfe Ayaz, Ali Cakmak
Taxonomy of living organisms gains major importance in making the study of vastly heterogeneous living things easier. In addition, various fields of applied biology (e.g., agriculture) depend on classification of living creatures. Specific fragments of the DNA sequence of a living organism have been defined as DNA…
H.M.Fazlul Haque, Fariha Arifin, Sheikh Adilina, Muhammod Rafsanjani + 1 more
The information of a cell is primarily contained in Deoxyribonucleic Acid (DNA). There is a flow of information of DNA to protein sequences via Ribonucleic acids (RNA) through transcription and translation. These entities are vital for the genetic process. Recent developments in epigenetic also show the importance of…
Authors not listed
Terminally labeled DNA oligonucleotides have wide applications in modern biology and biotechnological applications. It has been observed that the fluorescent intensity of light released from these fluorescent labels is heavily influenced by the terminal sequence of nucleotides. Recent studies have assayed and published…
Joseph Redshaw, Darren Ting, Alex Brown, Jonathan Hirst + 1 more
Antimicrobial peptides (AMPs) represent a potential solution to the growing problem of antimicrobial resistance, yet their identification through wet-lab experiments is a costly and timeconsuming process. Accurate computational predictions would allow rapid in silico screening of candidate AMPs, thereby accelerating…
Babu Bassa
In this communication the author describes a software tool named "ChameleonSort". The software program, developed by the present author is useful in the sorting of biological sequence variants like those accumulating mutations while diverging from the common ancestors. Examples include viral protein variants, protein…
Almas Jabeen, Nadeem Ahmad, Khalid Raza
RNA-Seq measures expression levels of several transcripts simultaneously. The identified reads can be gene, exon, or other region of interest. Various computational tools have been developed for studying pathogen or virus from RNA-Seq data by classifying them according to the attributes in several predefined classes…
Authors not listed
Sequence is the critical determinant of macromolecular function, yet current polymer design approaches often optimize monomer composition and ratios while ignoring sequence. This creates poorly defined design spaces for active learning that miss the vast combinatorial landscape of sequence possibilities. We introduce…
Daniel Probst
Last year, a preprint gained notoriety, proposing that a k-nearest neighbour classifier is able to outperform large-language models using compressed text as input and normalised compression distance (NCD) as a metric. In chemistry and biochemistry, molecules are often represented as strings, such as SMILES for small…