5 papers · ranked by Valyu relevance
Giansalvo Cirrincione, Elisa Ficarra, Marta Lovino
Protein language models (pLMs) such as ESM-2 and ProtBERT rely on pretraining corpora of tens to hundreds of millions of sequences and on encoder architectures whose depth, width and number of attention heads are chosen by the practitioner and never revisited during training. The entry cost of state-of-the-art pLMs is…
Amélie Barozet, Vincent Cabeli, Jean Ogier du Terrail, Alexey Rukhovich + 6 more
The development of climate-resilient crops would be greatly accelerated by models able to reason directly over plant genomic sequences and to pinpoint trait-associated regions or loci. Anticipating the impact of DNA base changes (variants) remains challenging, and understanding regulatory mechanisms is still an active…
Jackie Rao, Muntadher Jihad, Giulia Biffi, Paul D.W. Kirk
Identifying cell types from single-cell RNA sequencing (scRNA-seq) data typically requires several separate and often uninterpretable steps: dimensionality reduction, batch-correction, clustering, marker-gene identification and the discovery of finer-grained structure. Here we introduce scFLAME (single-cell Factor…
Brett Kiyota, Chaehyeon Lee, Haoyang Yao, Nozomu Yachie
The rapid expansion of single-cell genomic datasets has led to the compilation of biological resources comprising hundreds of millions of cells across tissues, developmental stages, and disease states. This has underscored the need for scalable and interpretable data representations that preserve the complex…
Francesco Carli, Polina Rusina, Lun Ai, Leonie Küchenhoff + 7 more
Language models and agents are increasingly used in biomedicine, but current benchmarks reward correct answers even when the underlying reasoning is flawed. Here we introduce Karenina, an open-source framework that turns expert knowledge into multi-dimensional evaluations of questions, conversations and autonomous…