6 papers · ranked by Valyu relevance
Yiming Xue, Xiaojian Liu, Weimin Zhu, Shengfan Wang + 2 more
While protein-RNA interactions are fundamental to post-transcriptional processes, achieving a holistic understanding of their regulatory logic remains challenging. Current computational models often treat binding affinity, interface mapping, and RNA design as isolated tasks, thereby failing to provide a unified…
Hao Ding, Nannan Wu, Tianyi Qiu
DNA foundation models such as Evo2 7B adopt hybrid Hyena/attention architectures (Striped-Hyena2) whose single-stream autoregressive decoding is bounded by weight bandwidth at ∼45 tok/s. Speculative decoding on such hybrids faces a systems problem that prior SSM work solves only partially: after a draft is verified…
Eliezer Masliah
How transient neural representations become integrated and stable enough to function as internal neural models remains incompletely understood. Grounded in efficient coding, Bayesian and predictive frameworks, recurrent and attractor dynamics, neural state-space models, and systems neuroscience, the Principle of…
Giansalvo Cirrincione, Elisa Ficarra, Marta Lovino
Protein language models (pLMs) such as ESM-2 and ProtBERT rely on pretraining corpora of tens to hundreds of millions of sequences and on encoder architectures whose depth, width and number of attention heads are chosen by the practitioner and never revisited during training. The entry cost of state-of-the-art pLMs is…
Minha Park, Shinwoo Kim, Seokhyun Moon, Hyeongwoo Kim + 2 more
Co-folding models accurately predict biomolecular complexes, but how their internal representations support joint structure prediction remains unclear. We analyze the Boltz-1 trunk by decomposing its pair representation into intra-chain and inter-chain blocks and applying ablation, layer-wise representation geometry…
Matteo Farina, Pietro Zamberlan, Arno Onken, Ulisse Ferrari
For datasets with thousands of neurons and images, vision transformers have proven successful at predicting neural responses to stimuli. However, they are expected to underperform in low-data regimes, where CNNs and Gaussian processes are considered more effective. We ask whether transformers can be made competitive…