20 papers · ranked by Valyu relevance
Adam Frankish, Mark Diekhans, Anne-Maud Ferreira, Rory Johnson + 51 more
'Irwin Jungreis' 'Jane Loveland' 'Jonathan M Mudge' 'Cristina Sisu' 'James Wright' 'Joel Armstrong' 'If Barnes' 'Andrew Berry' 'Alexandra Bignell' 'Silvia Carbonell\xa0Sala' 'Jacqueline Chrast' 'Fiona Cunningham' 'Tomás Di\xa0Domenico' 'Sarah Donaldson' 'Ian T Fiddes' 'Carlos García\xa0Girón' 'Jose Manuel Gonzalez'…
Adam Frankish, Sílvia Carbonell-Sala, Mark Diekhans, Irwin Jungreis + 58 more
'Jane\xa0E Loveland' 'Jonathan\xa0M Mudge' 'Cristina Sisu' 'James\xa0C Wright' 'Carme Arnan' 'If Barnes' 'Abhimanyu Banerjee' 'Ruth Bennett' 'Andrew Berry' 'Alexandra Bignell' 'Carles Boix' 'Ferriol Calvet' 'Daniel Cerdán-Vélez' 'Fiona Cunningham' 'Claire Davidson' 'Sarah Donaldson' 'Cagatay Dursun' 'Reham Fatima'…
Danielle Thierry-Mieg, Jean Thierry-Mieg
Background Regions covering one percent of the genome, selected by ENCODE for extensive analysis, were annotated by the HAVANA/Gencode group with high quality transcripts, thus defining a benchmark. The ENCODE Genome Annotation Assessment Project (EGASP) competition aimed at reproducing Gencode and finding new genes.…
Jonathan M Mudge, Sílvia Carbonell-Sala, Mark Diekhans, Jose Gonzalez Martinez + 52 more
'Jose\xa0Gonzalez Martinez' 'Toby Hunt' 'Irwin Jungreis' 'Jane\xa0E Loveland' 'Carme Arnan' 'If Barnes' 'Ruth Bennett' 'Andrew Berry' 'Alexandra Bignell' 'Daniel Cerdán-Vélez' 'Kelly Cochran' 'Lucas\xa0T Cortés' 'Claire Davidson' 'Sarah Donaldson' 'Cagatay Dursun' 'Reham Fatima' 'Matthew Hardy' 'Prajna Hebbar' 'Zoe…
Adam Frankish, Barbara Uszczynska, Graham RS Ritchie, Jose M Gonzalez + 7 more
'Jose M Gonzalez' 'Dmitri Pervouchine' 'Robert Petryszak' 'Jonathan M Mudge' 'Nuno Fonseca' 'Alvis Brazma' 'Roderic Guigo' 'Jennifer Harrow'] Background A vast amount of DNA variation is being identified by increasingly large-scale exome and genome sequencing projects. To be useful, variants require accurate functional…
Zeming Dong, Qiang Hu, Xiaofei Xie, Maxime Cordy + 2 more
'Jianjun Zhao'] Pre-trained code models lead the era of code intelligence. Many models have been designed with impressive performance recently. However, one important problem, data augmentation for code data that automatically helps developers prepare training data lacks study in the field of code learning. In this…
Siebren Frölich, Maarten van der Sande, Tilman Schäfers, Simon J. van Heeringen
'Simon J. van Heeringen'] Analyzing a functional genomics experiment, such as ATAC-, ChIP- or RNA-sequencing, requires reference data including a genome assembly and gene annotation. These resources can generally be retrieved from different organizations and in different versions. Most bioinformatic workflows require…
Irwin Jungreis, Michael L. Tress, Jonathan Mudge, Cristina Sisu + 11 more
In a 2018 paper posted to bioRxiv, Pertea et al. presented the CHESS database, a new catalog of human gene annotations that includes 1,178 new protein-coding predictions. These are based on evidence of transcription in human tissues and homology to earlier annotations in human and other mammals. Here, we reanalyze the…
Ales Varabyou, Beril Erdogdu, Steven L. Salzberg, Mihaela Pertea
ORFanage is a system designed to assign open reading frames (ORFs) to both known and novel gene transcripts while maximizing similarity to annotated proteins. The primary intended use of ORFanage is the identification of ORFs in the assembled results of RNA sequencing (RNA-seq) experiments, a capability that most…
S. Taylor Head, Aryun Nemani, Yung-Han Chang, Tabitha A. Harrison + 5 more
Most eQTL and TWAS analyses quantify expression using aggregate, tissue-agnostic transcript annotations and ignore isoform-level regulation, potentially obscuring or misattributing regulatory mechanisms. Here, we developed a framework leveraging publicly available long-read RNA-seq data to perform tissue-informed…
Raghunandan Wable, Achuth Suresh Nair, Anirudh Pappu, Widnie Pierre-Louis + 6 more
Timely understanding of biological secrets of complex diseases will ultimately benefit millions of individuals by reducing the high risks for mortality and improving the quality of life with personalized diagnoses and treatments. Due to the advancements in sequencing technologies and reduced cost, genomics data is…
Miguel Maquedano, Daniel Cerdán-Vélez, Michael L. Tress
In 2018 we analysed the three main repositories for the human proteome, Ensembl/GENCODE, RefSeq and UniProtKB. They disagreed on the coding status of one of every eight annotated coding genes. The analysis inspired bilateral collaborations between annotation groups. Here we have repeated our analysis with updated…
Zhiwen Pan, Jan Dellith, Lothar Wondraczek
Understanding the multivariate origin of physical properties is particularly complex for polyionic glasses. As a concept, the term genome has been used to describe the entirety of structure-property relations in solid materials, based on functional genes acting as descriptors for a particular property, for example, for…
Robersy Sanchez, Jesús Barreto
Experimental studies reveal that genome architecture splits into natural domains suggesting a well-structured genomic architecture, where, for each species, genome populations are integrated by individual mutational variants. Herein, we show that the architecture of population genomes from the same or closed related…
Authors not listed
One aim of the international Human Proteome Organization (HUPO) Human Proteome Project (HPP) is to obtain high-confidence translation evidence for every human protein-coding gene established in its target list of 19433 entries based on the protein-coding genes from Ensembl-GENCODE. However, 76 are annotated in…
Yueming Wu, Chengwei Liu, Yang Liu
With the boom in modern software development, open-source software has become an integral part of various industries, driving progress in computer science. However, the immense complexity and diversity of the open-source ecosystem also pose a series of challenges, including issues of quality, security, management…
Miloje Rakočević
In some previous works (2018a,b; 2019, 2021a,b, 2022) we presented a new type of mirror symmetry, expressed in the set of protein amino acids; such a symmetry, that it simultaneously represents the semiotic essence of the genetic code. In this paper we provide new evidences that the genetic code represents the unity of…
Wei Ma, Mengjie Zhao, Ezekiel Soremekun, Qiang Hu + 5 more
'Mike Papadakis' 'Maxime Cordy' 'Xiaofei Xie' 'Yves Le Traon'] Code embedding is a keystone in the application of machine learning on several Software Engineering (SE) tasks. To effectively support a plethora of SE tasks, the embedding needs to capture program syntax and semantics in a way that is generic. To this end…
Belfiore, Asia, Jonathan Passerat‐Palmbach, Dmitrii Usynin
The increased availability of genetic data has transformed genomics research, but raised many privacy concerns regarding its handling due to its sensitive nature. This work explores the use of language models (LMs) for the generation of synthetic genetic mutation profiles, leveraging differential privacy (DP) for the…
Babu Bassa
In this communication the author describes a software tool named "ChameleonSort". The software program, developed by the present author is useful in the sorting of biological sequence variants like those accumulating mutations while diverging from the common ancestors. Examples include viral protein variants, protein…