9 papers · ranked by Valyu relevance
Peter Heringer, Daniel Doerr
Pangenome graphs offer a compact and comprehensive representation of genomic diversity, improving tasks such as variant calling, genotyping, and other downstream analyses. Although the underlying graph structures scale sublinearly with the number of haplotypes, the widely used GFA file format suffers from rapidly…
Stuart Lee, Di Cook, Michael Lawrence
The Bioconductor project has created many useful data abstractions for analysing high-throughput genomics experiments. However, there is a cognitive load placed on a user in learning a data abstraction and understanding its appropriate use. Through-out a standard workflow, a user must navigate and know many of these…
Amy E. Ramage, Kaila Cote, Jill C. Thorson, Katelyn Lerner + 2 more
Language rehabilitation centers on modifying its use through experience-based neuroplasticity. Implicit statistical learning of language is essential to its acquisition and likely its rehabilitation following brain injury, but its corresponding brain networks remain elusive. Coordinate-based meta-analyses were…
R. Daniel Kortschak, David L. Adelson
bíogo is a framework designed to ease development and maintenance of computationally intensive bioinformatics applications. The library is written in the Go programming language, a garbage-collected, strictly typed compiled language with built in support for concurrent processing, and performance comparable to C and…
Melissa Sanabria, Jonas Hirsch, Anna R. Poetsch
Large Language Models (LLMs) on natural language have achieved a level of performance that allows the generation of coherent and syntactically correct text. DNA sequence of genomes follows rules similar to natural language, but a distinguishing factor is the absence of a concept analogous to words. We established…
Julie D. Thompson, Raymond Ripp, Claudine Mayer, Olivier Poch + 1 more
The X circular code is a set of 20 trinucleotides (codons) that has been identified in the protein-coding genes of most organisms (bacteria, archaea, eukaryotes, plasmids, viruses). It has been shown previously that the X circular code has the important mathematical property of being an error-correcting code. Thus…
Yi Jiang, Shuang Wang, Shaohong Feng, Cankun Wang + 5 more
Foundation models have transformed AI by leveraging large-scale data to efficiently perform diverse tasks, and their applications in bioinformatics are primarily focused on data-centric tasks like cell type annotation and gene expression analysis. However, their potential extends beyond data analysis, offering…
Jacqueline A Jansen, Artür Manukyan, Nour Al Khoury, Altuna Akalin
Data analysis is constrained by a shortage of skilled experts, particularly in biology, where detailed data interpretation is vital for understanding complex biological processes and developing new treatments and diagnostics. To address this, we developed mergen, an R package that leverages Large Language Models (LLMs)…
Huifang Ma, Zhicheng Ji
Large language models have shown remarkable capabilities in algorithm design, but their effectiveness in solving data science challenges remains poorly understood. We conducted a classroom experiment in which graduate students used large language models (LLMs) to solve biomedical data science challenges on Kaggle.…