24 papers · ranked by Valyu relevance
Yutaka Satou, Katsuhiko Mineta, Michio Ogasawara, Yasunori Sasakura + 16 more
'Eiichi Shoguchi' 'Keisuke Ueno' 'Lixy Yamada' 'Jun Matsumoto' 'Jessica Wasserscheid' 'Ken Dewar' 'Graham B Wiley' 'Simone L Macmil' 'Bruce A Roe' 'Robert W Zeller' 'Kenneth EM Hastings' 'Patrick Lemaire' 'Erika Lindquist' 'Toshinori Endo' 'Kohji Hotta' 'Kazuo Inaba'] An improved assembly of the Ciona intestinalis…
Bernardo P. de Almeida, Hugo Dalla-Torre, Guillaume Richard, Christopher Blum + 14 more
Genome annotation models that directly analyze DNA sequences are indispensable for modern biological research, enabling rapid and accurate identification of genes and other functional elements. This capability is paramount as the volume of sequenced genomes rapidly expands, making the need for efficient and accurate…
Siyuan Li, Zedong Wang, Zicheng Liu, Di Wu + 4 more
Genomic Sequence Modeling Authors: ['Siyuan Li' 'Zedong Wang' 'Zicheng Liu' 'Di Wu' 'Cheng Tan' 'Jiangbin Zheng' 'Yufei Huang' 'Stan Z. Li'] Similar to natural language models, pre-trained genome language models are proposed to capture the underlying intricacies within genomes with unsupervised sequence modeling. They…
Thomas M Asbury, Matt Mitman, Jijun Tang, W Jim Zheng
Background New technologies are enabling the measurement of many types of genomic and epigenomic information at scales ranging from the atomic to nuclear. Much of this new data is increasingly structural in nature, and is often difficult to coordinate with other data sets. There is a legitimate need for integrating and…
Tianyu Liu, Xiangyu Zhang, Rex Ying, Hongyu Zhao
Sequence-to-function models can predict gene expression from sequence data and be used to link genetic information with transcriptomics data to understand regulatory processes and their effects on complex phenotypes. The genomic language models are pre-trained with large-scale DNA sequences and can generate robust…
Gonzalo Benegas, Chengzhong Ye, Carlos Albors, Jianan Canal Li + 1 more
'Yun S. Song'] Large language models (LLMs) are having transformative impacts across a wide range of scientific fields, particularly in the biomedical sciences. Just as the goal of Natural Language Processing is to understand sequences of words, a major objective in biology is to understand biological sequences.…
Zicheng Liu, Jiahui Li, Siyuan Li, Zelin Zang + 4 more
Foundation Models Authors: ['Zicheng Liu' 'Jiahui Li' 'Siyuan Li' 'Zelin Zang' 'Cheng Tan' 'Yufei Huang' 'Yajing Bai' 'Stan Z. Li'] The Genomic Foundation Model (GFM) paradigm is expected to facilitate the extraction of generalizable representations from massive genomic data, thereby enabling their application across a…
David R. Kelley, Yakir A. Reshef, David Belanger, Cory Y. McLean + 2 more
Models for predicting phenotypic outcomes from genotypes have important applications to understanding genomic function and improving human health. Here, we develop a machine-learning system to predict cell type-specific epigenetic and transcriptional profiles in large mammalian genomes from DNA sequence alone. Using…
Samuel A. Cushman
Integration of genomic and epigenomic datasets with genetic modeling is improving model testing, validation and calibration (Figure [F1]). Spatially explicit, individual based simulation modeling has advanced such that relationships between environmental characteristics, population structure and the genetic or…
Frederikke Isa Marin, Felix Teufel, Marc Horrender, Dennis Madsen + 3 more
'Dennis Pultz' 'Ole Winther' 'Wouter Boomsma'] The genome sequence contains the blueprint for governing cellular processes. While the availability of genomes has vastly increased over the last decades, experimental annotation of the various functional, non-coding and regulatory elements encoded in the DNA sequence…
Salman Mohamadi, Farhang Yeganegi, Hamidreza Amindavar
This paper provides a framework in order to statistically model sequences from human genome, which is allowing a formulation to synthesize gene sequences. We start by converting the alphabetic sequence of genome to decimal sequence by Huffman coding. Then, this decimal sequence is decomposed by HP filter into two…
Fang Liu, Eivind Tøstesen, Jostein K Sundet, Tor-Kristian Jenssen + 5 more
'Christoph Bock' 'Geir Ivar Jerstad' 'William G Thilly' 'Eivind Hovig' 'Yves van de Peer'] In a living cell, the antiparallel double-stranded helix of DNA is a dynamically changing structure. The structure relates to interactions between and within the DNA strands, and the array of other macromolecules that constitutes…
Pooja Kathail, Ayesha Bajwa, Nilah M. Ioannidis
prediction Authors: ['Pooja Kathail' 'Ayesha Bajwa' 'Nilah M. Ioannidis'] The majority of genetic variants identified in genome-wide association studies of complex traits are non-coding, and characterizing their function remains an important challenge in human genetics. Genomic deep learning models have emerged as a…
Sayan Ghosal, Youssef Barhomi, Tejaswini Ganapathi, Amy Krystosik + 5 more
Accurately predicting gene expression from DNA sequence remains a central challenge in human genetics. Current sequence-based models overlook natural genetic variation across individuals, while population-based models are restricted to variants observed within specific cohorts. Here, we present VariantFormer, a…
Michael J Gilchrist
The Xenopus community has made concerted efforts over the last 10-12 years systematically to improve the available sequence information for this amphibian model organism ideally suited to the study of early development in vertebrates. Here I review progress in the collection of both sequence data and physical clone…
Bilal Wajid, Erchin Serpedin, Mohamed Nounou, Hazem Nounou
Reference assisted assembly requires the use of a reference sequence, as a model, to assist in the assembly of the novel genome. The standard method for identifying the best reference sequence for the assembly of a novel genome aims at counting the number of reads that align to the reference sequence, and then choosing…
Mateusz Chiliński, Dariusz Plewczynski
Prediction of chromatin interactions from DNA sequence has been a significant research challenge in the last couple of years. Several solutions have been proposed, most of which are based on encoder-decoder architecture, where 1D sequence is convoluted, encoded into the latent representation, and then decoded using 2D…
Ethalinda K S Cannon, David C Molik, Adam J Wright, Huiting Zhang + 4 more
'Loren Honaas' 'Kapeel Chougule' 'Sarah Dyer' 'T Harris'] Title: Abstract The rapid increase in the number of reference-quality genome assemblies presents significant new opportunities for genomic research. However, the absence of standardized naming conventions for genome assemblies and annotations across datasets…
Kasper Munch, Anders Krogh
Background The number of sequenced eukaryotic genomes is rapidly increasing. This means that over time it will be hard to keep supplying customised gene finders for each genome. This calls for procedures to automatically generate species-specific gene finders and to re-train them as the quantity and quality of reliable…
Alex Lee, Joshua Rackers, William Bricker
One of the fundamental limitations of accurately modeling biomolecules like DNA is the inability to perform quantum chemistry calculations on large molecular structures. We present a machine learning model based on an equivariant Euclidean Neural Network framework to obtain quantum-accurate electron densities for…
Seongmun Jeong, Jae-Yoon Kim, Namshin Kim
CVRMS is an R package designed to extract marker subsets from repeated rank-based marker datasets generated from genome-wide association studies or marker effects for genome-wide prediction (https://github.com/lovemun/CVRMS). CVRMS provides an optimized genome-wide biomarker set with the best predictability of…
Hélène Lopez Maestre, Lilia Brinza, Camille Marchet, Janice Kielbassa + 11 more
SNPs (Single Nucleotide Polymorphisms) are genetic markers whose precise identification is a prerequisite for association studies. Methods to identify them are currently well developed for model species, but rely on the availability of a (good) reference genome, and therefore cannot be applied to non-model species.…
Authors not listed
The development of RNA-based therapeutics has significantly expanded the landscape of drug discovery by enabling precise modulation of gene expression. These approaches offer the potential to target previously "undruggable" genes, overcoming limitations inherent to traditional small molecule therapies. However, the…
Martin Šícho, Xuhan Liu, Daniel Svozil, Gerard van Westen
This manuscript describes the development and architecture of the GenUI software platform for integration of molecular generators. The source code for the components of the platform is available in the following repositories: https://github.com/martin-sicho/genui https://github.com/martin-sicho/genui-gui…