Search · four archives
Search · four archives
16 papers · ranked by Valyu relevance
Sasha Darmon, Arnaud Mary, Vincent Lacroix
Transcribed repeats represent a major challenge in the de novo assembly of transcriptomes from short RNA-seq reads. Young transposable elements (TEs) and more generally, inexact repeats, create dense and ambiguous regions in the assembly graph, preventing the correct assembly of transcripts. In this paper, we introduce…
Gerardo Patiño-Guillén, Jovan Pešović, Marko Panić, Max Earle + 5 more
Short tandem repeat expansions underlie a class of neurological and neuromuscular diseases known as repeat expansion disorders, yet the precise characterisation of these repeats remains technically challenging. Conventional amplification-based methods fail to resolve repeat length accurately due to amplification bias…
Jabale Rahmat, Tuan Pham, Amanda M. Larracuente
Highly repetitive sequences pose problems for genome assembly and analysis. While advances in long-read sequencing technologies have helped reveal the organization of repetitive genomic sequences at unprecedented resolution, their functional characterization remains difficult because molecular assays that probe…
Daniel Garcia-Ruano, Mikaël Georges, Saswat K. Mohanty, Rahma Baaziz + 3 more
Repetitive elements show different patterns between lncRNAs and protein-coding transcripts. Several studies show that lncRNAs associate closely with TEs: Kapusta et al. reported 75% of lncRNAs (GENCODE v13) harbor TE fragments, and Carlevaro-Fita et al. found ~83% contain TEs. Our RepeatMasker results indicate that 70%…
B. Poggiali, L. Putzeys, J. D. Andersen, A. Vidaki
The human genome is dominated by repetitive DNA, whose genetic and epigenetic variation plays a key role in gene regulation, genome stability, and disease. Recent advances in long-read sequencing now enable large-scale, haplotype-resolved, and DNA methylation-informative analysis of the human genome, including on…
Brett N Adey, Danielle J Maddock, Sylvie Hermann-Le Denmat, Marcel E Dinger + 4 more
Large genomes such as the human genome are pervasively transcribed yet encode relatively few unambiguously functional elements. This has led to debate over whether pervasive transcription is indicative of large suites of uncharacterized functional elements or is simply background noise. Here, we used a deep-learning…
Yoojung Han, Ja-Hyun Jang, Hyeshik Chang
Tandem repeat expansion disorders can be difficult to diagnose when expansions exceed 200 repeats, as standard methods (for example, Southern blot and modified PCR) often fail. We present a Cas9-targeted nanopore sequencing workflow and an automated analysis pipeline, RepeatLab, for accurate repeat-length estimation…
Siyuan Li, Kai Yu, Anna Wang, Zicheng Liu + 6 more
Modeling genomic sequences faces two unsolved challenges: the information density varies widely across different regions, while there is no clearly defined minimum vocabulary unit. Relying on either four primitive bases or independently designed DNA tokenizers, existing approaches with naive masked language modeling…
Alexander Sweeten, Michael C. Schatz, Adam M. Phillippy
Satellite DNA has long posed challenges for genome assembly and analysis due to its low sequence complexity and poor mappability. These large heterochromatic arrays of tandem repeats are ubiquitous across eukaryotic genomes, yet remain understudied. Current methods for annotating satellite regions, and other classes of…
Rui Zhu, Xiaopu Zhou, Haixu Tang, Stephen W. Scherer + 1 more
Trained on massive cross-species DNA corpora, DNA Large language Models (LLMs) learned the fundamental "grammar" and evolutionary patterns of genomic sequences. This makes them powerful priors for DNA sequence modeling, particularly across long distances. Yet, two major constraints hinder their use: the quadratic…
Authors not listed
RNA–RNA interactions drive the formation of biomolecular condensates via liquid–liquid phase separation (LLPS), but the molecular mechanisms governing this phenomenon remain poorly understood. Here, we employ Martini 3 coarse-grained simulations to investigate phase transitions of G4C2 RNA repeats—sequences implicated…
Atsushi Takeda, Tsukasa Fukunaga, Michiaki Hamada
We developed BWR-finder (Burrows–Wheeler transform-based Repeat finder), a new software tool for database-free detection of interspersed repeats in large genomes of tens of gigabases. BWR-finder employs a BWT-based seed-and-extend repeat detection algorithm and parallelized extension computation, improving both runtime…
Xianghao Zhan, Jingyu Xu, Yuanning Zheng, Zinaida Good + 1 more
Spatial transcriptomics enables spatial gene expression profiling, motivating computational models that capture spatially conditioned regulatory relationships. We introduce SAGE-FM, a lightweight spatial transcriptomics foundation model based on graph convolutional networks (GCN) trained with a masked-central-spot…
Authors not listed
The WRN helicase has recently emerged as a promising therapeutic target for microsatellite instability (MSI)-high cancers. Here, we report LXW-P1, a potent WRN degrader derived from marine bromotyrosine alkaloids. Its molecular target was identified using an AI-guided, pathway-informed perturbation transcriptomics…
Selçuk Korkmaz
Data leakage remains a recurrent source of optimistic bias in biomedical machine learning studies. Standard row-wise cross-validation and globally estimated preprocessing steps are often inappropriate for data with repeated measurements, study-level heterogeneity, batch effects, or temporal dependencies. This paper…
Jianyu Yang, Shaun Mahony
Interpreting genomics deep learning models remains challenging. Existing feature attribution methods are largely restricted to one-hot DNA inputs and therefore cannot assess the influence of more general genomic features such as chromatin states or genomic repeats. Concept attribution methods offer an input-agnostic…