23 papers · ranked by Valyu relevance
Shashi Dhanasekar, Akash Saranathan, Pengtao Xie
Accurately predicting gene function from DNA sequences remains a fundamental challenge in genomics, particularly given the limited experimental annotation available for most genes. Existing computational approaches often formulate function prediction as a classification task over predefined categories, limiting their…
Omkar Chandra, Madhu Sharma, Neetesh Pandey, Indra Prakash Jha + 3 more
The number of annotated genes in the human genome has increased tremendously, and understanding their biological role is challenging through experimental methods alone. There is a need for a computational approach to infer the function of genes, particularly for non-coding RNAs, with reliable explainability. We have…
Ren Qi, Arun Kumar Sangaiah, Dariusz Mrozek, Quan Zou
Predicting the function of genes is a critical problem in biology. The current generation rate of new gene sequences is too fast to discover and validate them experimentally, emphasizing the importance of machine learning. Machine learning techniques have advanced our understanding of gene function, which have been…
Elena Rojano, Fernando M. Jabato, James R. Perkins, José Córdoba-Caballero + 5 more
'José Córdoba-Caballero' 'Federico García-Criado' 'Ian Sillitoe' 'Christine Orengo' 'Juan A. G. Ranea' 'Pedro Seoane-Zonjic'] Background Protein function prediction remains a key challenge. Domain composition affects protein function. Here we present DomFun, a Ruby gem that uses associations between protein domains and…
Danielle Miller, Ofir Arias, David Burstein, Janet Kelso
The “Sequence” search mode allows users to submit a sequence query in FASTA format either by uploading a file or directly pasting protein sequences into the search box. If a family with sufficient similarity is found (E-value < 10−4), the web server will provide information for the best hits to the proteins provided…
Rohan Shawn Sunil, Shan Chun Lim, Manoj Itharajula, Marek Mutwil
Elucidating gene function is one of the ultimate goals of plant science. Despite this, only ~15% of all genes in the model plant Arabidopsis thaliana have comprehensively experimentally verified functions. While bioinformatical gene function prediction approaches can guide biologists in their experimental efforts…
Aysun Urhan, Bianca-Maria Cosma, Abigail L. Manson, Thomas Abeel
Today, we know the function of only a small fraction of all known protein sequences identified. This problem is even more salient in bacteria as human-centric studies are prioritized in the field and there is much to uncover in the bacterial genetic repertoire. Conventional approaches to bacterial gene annotation are…
Constance J. Jeffery
In recent years, improvements in protein function prediction methods have led to increased success in annotating protein sequences. However, the functions of over 30% of protein-coding genes remain unknown for many sequenced genomes. Protein functions vary widely, from catalyzing chemical reactions to binding DNA or…
Rund Tawfiq, Maxat Kulmanov, Robert Hoehndorf
Protein function annotation has traditionally followed a reductionist approach, assigning functions to individual proteins acting in isolation. This paradigm treats each annotation as an independent fact, disconnected from the broader biological system. However, proteins operate within integrated cellular networks…
Divyanshu Aggarwal, Yasha Hasija
—Deep Learning and big data have shown tremendous success in bioinformatics and computational biology in recent years; artificial intelligence methods have also significantly contributed in the task of protein function classification. This review paper analyzes the recent developments in approaches for the task of…
Alexander Adrian-Hamazaki, Paul Pavlidis
It is widely accepted in genomics that coexpression of RNA transcripts suggests a commonality of function. This intuition is explicitly leveraged in machine learning methods that predict gene function, where it is often combined with other features such as protein interactions and sequence similarity. For example…
Bas Stringer, Annika Jacobsen, Qingzhen Hou, Hans de Ferrante + 5 more
'Olga Ivanova' 'Katharina Waury' 'Jose Gavaldá-García' 'Sanne Abeln' 'K. Anton Feenstra'] | 11 | Function Prediction | 1 | | --- | --- | --- | | | Bas Stringer Annika Jacobsen Qingzhen Hou | | | | Hans de Ferrante Olga Ivanova Katharina Waury | | | | Jose Gavald´a-Garc´ıa Sanne Abeln K. Anton Feenstra | | | | 1…
Christopher A Mancuso, Patrick S Bills, Douglas Krum, Jacob Newsted + 2 more
Biomedical researchers take advantage of high-throughput, high-coverage technologies to routinely generate sets of genes of interest across a wide range of biological conditions. Although these technologies have directly shed light on the molecular underpinnings of various biological processes and diseases, the list of…
Abigail Djossou, Wend Yam D D Ouedraogo, Aida Ouangraoua, Alex Bateman
In gene prediction tool development, the choice of information and the algorithm used to process it are crucial. Gene structures are typically predicted using four types of information: signal sensors, content sensors, gene similarity, and experimental data (). The first two are considered intrinsic information, while…
Yan, Jingquan, Yuwei Miao, Lei Yu + 4 more
Exploring how genetic sequences shape phenotypes is a fundamental challenge in biology and a key step toward scalable, hypothesis-driven experimentation. The task is complicated by the large modality gap between sequences and phenotypes, as well as the pleiotropic nature of gene–phenotype relationships. Existing…
Chengyu Liu, Wei Wang
Developing models with high interpretability and even deriving formulas to quantify relationships between biological data is an emerging need. We propose here a framework for ab initio derivation of sequence motifs and linear formula using a new approach based on the interpretable neural network model called contextual…
Morgan N. Price, Adam P. Arkin
Automated annotations of protein function are error-prone because of our lack of knowledge of protein functions. For example, it is often impossible to predict the correct substrate for an enzyme or a transporter. Furthermore, much of the knowledge that we do have about the functions of proteins is missing from the…
Jesús Antonio Motta, Pedro David Gómez
In this work, we present a highly efficient machine learning method for identifying DNA sequences that code for genes. The learning process is based on Human Genome Build 38 (GRCh38) sequences extracted from various specialized databases. The sequences were then translated into amino acid sequences and used to build…
Sheikh Sunzid Ahmed
Francisella tularensis Schu S4 is the causal agent of a sporadic zoonotic disease known as Tularemia, which has shown epidemic outbreaks recently in certain parts of the world. This pathogen is a potential agent of biowarfare or bioterrorism and is classified as a category A pathogen by the National Institute of…
Lingmin Zhan, Yuanyuan Zhang, Yingdong Wang, Aoyi Wang + 5 more
genetic perturbations from multiplex biological networks Authors: ['Lingmin Zhan' 'Yuanyuan Zhang' 'Yingdong Wang' 'Aoyi Wang' 'Caiping Cheng' 'Jinzhong Zhao' 'Wuxia Zhang' 'P.H. Hiller L. Lia' 'Jianxin Chen'] Systematic characterization of biological effects to genetic perturbation is essential to the application of…
Authors not listed
One aim of the international Human Proteome Organization (HUPO) Human Proteome Project (HPP) is to obtain high-confidence translation evidence for every human protein-coding gene established in its target list of 19433 entries based on the protein-coding genes from Ensembl-GENCODE. However, 76 are annotated in…
Authors not listed
The development of RNA-based therapeutics has significantly expanded the landscape of drug discovery by enabling precise modulation of gene expression. These approaches offer the potential to target previously "undruggable" genes, overcoming limitations inherent to traditional small molecule therapies. However, the…
Authors not listed
Integrating machine learning (ML) into drug discovery has ushered in a new era of innovation, dramatically enhancing the efficiency and precision of identifying and developing new therapeutics. This review provides a comprehensive analysis of the current applications of machine learning in drug discovery, focusing on…