24 papers · ranked by Valyu relevance
Chengwei Wei, Yun-Cheng Wang, Bin Wang, C.‐C. Jay Kuo
Language modeling studies the probability distributions over strings of texts. It is one of the most fundamental tasks in natural language processing (NLP). It has been widely used in text generation, speech recognition, machine translation, etc. Conventional language models (CLMs) aim to predict the probability of…
Michael R. Douglas
Artificial intelligence is making spectacular progress, and one of the best examples is the development of large language models (LLMs) such as OpenAI's GPT series. In these lectures, written for readers with a background in mathematics or physics, we give a brief history and survey of the state of the art, and…
Jack Grieve, Sara Bartl, Matteo Fuoli, Jason Grafmiller + 6 more
In this article, we introduce a sociolinguistic perspective on language modeling. We claim that language models in general are inherently modeling varieties of language, and we consider how this insight can inform the development and deployment of language models. We begin by presenting a technical definition of the…
Sofia Serrano, Zander Brumbaugh, Noah A. Smith
| 1 | Introduction | | 2 | | --- | --- | --- | --- | | 2 | | Background: Natural language processing concepts and tools | 3 | | | 2.1 | Taskification: Defining what we want a system to do | 4 | | | | 2.1.1 Abstract vs. concrete system capabilities | 4 | | | | 2.1.2 We need data and an evaluation method for research…
Aisha Khatun, Anisur Rahman, Hemayet Ahmed Chowdhury, Md. Saiful Islam + 1 more
'Md. Saiful Islam' 'Ayesha Tasnim'] Abstract. Language models are at the core of natural language processing. The ability to represent natural language gives rise to its applications in numerous NLP tasks including text classification, summarization, and translation. Research in this area is very limited in Bangla due…
Martin Schrimpf, Idan Blank, Greta Tuckute, Carina Kauf + 4 more
The neuroscience of perception has recently been revolutionized with an integrative modeling approach in which computation, brain function, and behavior are linked across many datasets and many computational models. By revealing trends across models, this approach yields novel insights into cognitive and neural…
Jose Juan Almagro Armenteros, Alexander Rosenberg Johansen, Ole Winther, Henrik Nielsen
Language modelling (LM) on biological sequences is an emergent topic in the field of bioinformatics. Current research has shown that language modelling of proteins can create context-dependent representations that can be applied to improve performance on different protein prediction tasks. However, little effort has…
Paul Azunre, Salomey Osei, Salomey Afua Addo, Lawrence Adu-Gyamfi + 23 more
'Stephen Moore' 'Bernard Adabankah' 'Bernard Opoku' 'Clara Asare-Nyarko' 'Samuel Nyarko' 'Cynthia Amoaba' 'Esther Dansoa Appiah' 'Felix Akwerh' 'Richard Nii Lante Lawson' 'Joel Budu' 'Emmanuel Debrah' 'Nana Boateng' 'Wisdom Ofori' 'Edwin Buabeng-Munkoh' 'Franklin Adjei' 'Isaac K. E. Ampomah' 'Joseph Otoo' 'Reindorf…
Alessandro Lopopolo, Evelina Fedorenko, Roger Levy, Milena Rabovsky
In this section, we provide a concise, accessible overview of key computational concepts (language models and word embeddings) and methods which are, with all the due differences and peculiarities, integral to many of the papers featured in this special issue. Word embeddings, also known as word vector representations…
Fudong Zhang, Bo Chai, Yujie Wu, Wai Ting Siok + 1 more
Elucidating the language-brain relationship requires bridging the methodological gap between linguistics’ abstract theoretical frameworks and neuroscience’s empirical neural data. As an interdisciplinary cornerstone, computational neuroscience formalizes language’s hierarchical and dynamic structures into testable…
Wenzhe Yang
In this paper, we discuss how pure mathematics and theoretical physics can be applied to the study of language models. Using set theory and analysis, we formulate mathematically rigorous definitions of language models, and introduce the concept of the moduli space of distributions for a language model. We formulate a…
Sanjar Adilov
Generative neural networks have shown promising results in de novo drug design. Recent studies suggest that one of the efficient ways to produce novel molecules matching target properties is to model SMILES sequences using deep learning in a way similar to language modeling in natural language processing. In this…
Ismail Ismail
Smoothing is one technique to overcome data sparsity in statistical language model. Although in its mathematical definition there is no explicit dependency upon specific natural language, different natures of natural languages result in different effects of smoothing techniques. This is true for Russian language as…
Kristijan Armeni, Roel M. Willems, Stefan Frank
Cognitive neuroscientists of language comprehension study how neural computations relate to cognitive computations during comprehension. On the cognitive part of the equation, it is important that the computations and processing complexity are explicitly defined. Probabilistic language models can be used to give a…
Nikola Kölbl, Stefan Rampp, Martin Kaltenhäuser, Konstantin Tziridis + 5 more
Language comprehension involves continuous prediction of upcoming words, with syntactic structure and semantic meaning intertwined in the human brain. To date, few studies have used combined magnetoencephalography (MEG) and electroen-cephalography (EEG) measurements to investigate how syntactic processing, predictive…
Shuya Nakata, Yoshiharu Mori, Shigenori Tanaka
Ultra-large virtual chemical spaces have emerged as a valuable resource for drug discovery, providing access to billions of make-on-demand compounds with high synthetic success rates. Chemical language models can potentially accelerate the exploration of these vast spaces through direct compound generation. However…
Authors not listed
Predicting molecular properties is a key challenge in drug discovery. Machine learning models, especially those based on transformer architectures, are increasingly used to make these predictions from chemical structures. Inspired by recent progress in natural language processing, many studies have adopted encoder-only…
Ping Li, Xiaowei Zhao
Connectionist models have had a profound impact on theories of language. While most early models were inspired by the classic parallel distributed processing architecture, recent models of language have explored various other types of models, including self-organizing models for language acquisition. In this paper, we…
Michael Moret, Lukas Friedrich, Francesca Grisoni, Daniel Merk + 1 more
Generative machine learning models sample drug-like molecules from chemical space without the need for explicit design rules. A deep learning framework for customized compound library generation is presented, aiming to enrich and expand the pharmacologically relevant chemical space with new molecular entities ‘on…
Abigail G. Toth, Petra Hendriks, Niels A. Taatgen, Jacolien van Rij
During real-time language processing, people rely on linguistic and non-linguistic biases to anticipate upcoming linguistic input. One of these linguistic biases is known as the implicit causality bias, wherein language users anticipate that certain entities will be rementioned in the discourse based on the entity's…
Authors not listed
Large Language Models (LLMs) based on transformer architectures excel at internet-scale tasks. However, real-world scientific scenarios—such as synthetic chemistry laboratories and autonomous experimental setups—typically involve incremental data generation in batches as new chemical reactions are conducted, unlike…
Daniel Mitropolsky, Christos H. Papadimitriou
Despite tremendous progress in neuroscience, we do not have a compelling narrative for the precise way whereby the spiking of neurons in our brain results in high-level cognitive phenomena such as planning and language. We introduce a simple mathematical formulation of six basic and broadly accepted principles of…
Authors not listed
Scientific modeling often requires navigating a trade-off between physical interpretability and empirical accuracy—a task that can take weeks of iteration, especially in systems with partial observability, structural complexity, and experimental errors. Here, we show how a state-of-the-art agentic reasoning-and-coding…
Nathan Frey, Ryan Soklaski, Simon Axelrod, Siddharth Samsi + 3 more
Massive scale, both in terms of data availability and computation, enables significant breakthroughs in key application areas of deep learning such as natural language processing (NLP) and computer vision. There is emerging evidence that scale may be a key ingredient in scientific deep learning, but the importance of…