25 papers · ranked by Valyu relevance
Arkaitz Zubiaga
Since their inception in the 1980s, language models (LMs) have been around for more than four decades as a means for statistically modeling the properties observed from natural language (Rosenfeld, ). Given a collection of texts as input, a language model computes statistical properties of language from those texts…
Chengwei Wei, Yun-Cheng Wang, Bin Wang, C.‐C. Jay Kuo
Language modeling studies the probability distributions over strings of texts. It is one of the most fundamental tasks in natural language processing (NLP). It has been widely used in text generation, speech recognition, machine translation, etc. Conventional language models (CLMs) aim to predict the probability of…
Fudong Zhang, Bo Chai, Yujie Wu, Wai Ting Siok + 1 more
Elucidating the language-brain relationship requires bridging the methodological gap between linguistics’ abstract theoretical frameworks and neuroscience’s empirical neural data. As an interdisciplinary cornerstone, computational neuroscience formalizes language’s hierarchical and dynamic structures into testable…
Michael R. Douglas
Artificial intelligence is making spectacular progress, and one of the best examples is the development of large language models (LLMs) such as OpenAI's GPT series. In these lectures, written for readers with a background in mathematics or physics, we give a brief history and survey of the state of the art, and…
Suhas Arehalli, Tal Linzen
Languages are governed by syntactic constraints-structural rules that determine which sentences are grammatical in the language. In English, one such constraint is subject-verb agreement, which dictates that the number of a verb must match the number of its corresponding subject: “the dogs run”, but “the dog runs”.…
Raphaël Millière
This chapter critically examines the potential contributions of modern language models to theoretical linguistics. Despite their focus on engineering goals, these models' ability to acquire sophisticated linguistic knowledge from mere exposure to data warrants a careful reassessment of their relevance to linguistic…
Alessandro Lopopolo, Evelina Fedorenko, Roger Levy, Milena Rabovsky
In this section, we provide a concise, accessible overview of key computational concepts (language models and word embeddings) and methods which are, with all the due differences and peculiarities, integral to many of the papers featured in this special issue. Word embeddings, also known as word vector representations…
Vittoria Dentella, Fritz Günther, Evelina Leivada
Title: Significance The synthetic language generated by recent Large Language Models (LMs) strongly resembles the natural languages of humans. This resemblance has given rise to claims that LMs can serve as the basis of a theory of human language. Given the absence of transparency as to what drives the performance of…
Wenzhe Yang
In this paper, we discuss how pure mathematics and theoretical physics can be applied to the study of language models. Using set theory and analysis, we formulate mathematically rigorous definitions of language models, and introduce the concept of the moduli space of distributions for a language model. We formulate a…
Dongqiu Zhang, Wenkui Li
Natural Language Understanding (NLU) and Natural Language Generation (NLG) are the general methods that support machine understanding of text content. They play a very important role in the text information processing system including recommendation and question and answer systems. There are many researches in the…
Richard Futrell, Kyle Mahowald
Language models can produce fluent, grammatical text. Nonetheless, some maintain that language models don't really learn language and also that, even if they did, that would not be informative for the study of human learning and processing. On the other side, there have been claims that the success of LMs obviates the…
Adrielli Lopes Rego, Joshua Snell, Martijn Meeter
Although word predictability is commonly considered an important factor in reading, sophisticated accounts of predictability in theories of reading are yet lacking. Computational models of reading traditionally use cloze norming as a proxy of word predictability, but what cloze norms precisely capture remains unclear.…
Mahrad Almotahari
Cooperative speech is purposive. From the speaker's perspective, one crucial purpose is the transmission of knowledge. Cooperative speakers care about getting things right for their conversational partners. This attitude is a kind of respect. Cooperative speech is an ideal form of communication because participants…
Nicola Angius, Pietro Perconti, Alessio Plebe, Alessandro Acciai
This article provides an epistemological analysis of current attempts of explaining how the relatively simple algorithmic components of neural language models (NLMs) provide them with genuine linguistic competence. After introducing the Transformer architecture, at the basis of most of current NLMs, the paper firstly…
Nikola Kölbl, Stefan Rampp, Martin Kaltenhäuser, Konstantin Tziridis + 5 more
Language comprehension involves continuous prediction of upcoming words, with syntactic structure and semantic meaning intertwined in the human brain. To date, few studies have used combined magnetoencephalography (MEG) and electroen-cephalography (EEG) measurements to investigate how syntactic processing, predictive…
Greta Tuckute, Aalok Sathe, Shashank Srikant, Maya Taliaferro + 4 more
Transformer models such as GPT generate human-like language and are highly predictive of human brain responses to language. Here, using fMRI-measured brain responses to 1,000 diverse sentences, we first show that a GPT-based encoding model can predict the magnitude of brain response associated with each sentence. Then…
Authors not listed
Step-by-step thinking is essential in all domains of chemical sciences and engineering. While machine learning tools are broadly used, algorithms that automate reasoning are far less common. We elaborate on seven categories of human reasoning activities and connect each to applications in chemical science and…
Shuya Nakata, Yoshiharu Mori, Shigenori Tanaka
Ultra-large virtual chemical spaces have emerged as a valuable resource for drug discovery, providing access to billions of make-on-demand compounds with high synthetic success rates. Chemical language models can potentially accelerate the exploration of these vast spaces through direct compound generation. However…
Kevin Maik Jablonka, Philippe Schwaller, Andres Ortega-Guerrero, Berend Smit
Machine learning has revolutionized many fields and has recently found applications in chemistry and materials science. The small datasets commonly found in chemistry sparked the development of sophisticated machine-learning approaches that incorporate chemical knowledge for each application and, therefore, require…
Daniel Mitropolsky, Christos H. Papadimitriou
Despite tremendous progress in neuroscience, we do not have a compelling narrative for the precise way whereby the spiking of neurons in our brain results in high-level cognitive phenomena such as planning and language. We introduce a simple mathematical formulation of six basic and broadly accepted principles of…
Yaqing Su, Lucy J. MacGregor, Itsaso Olasagasti, Anne-Lise Giraud
Understanding speech requires mapping fleeting and often ambiguous soundwaves to meaning. While humans are known to exploit their capacity to contextualize to facilitate this process, how internal knowledge is deployed on-line remains an open question. Here, we present a model that extracts multiple levels of…
Dimitris Gkoumas, Maria Liakata
The intersection of chemistry and Artificial Intelligence (AI) is an active area of research focused on accelerating scientific discovery. While using large language models (LLMs) with scientific modalities has shown potential, there are significant challenges to address, such as improving training efficiency and…
Authors not listed
Accelerating computational materials science relies not only on hardware advances but also on software that increases the ease of working with the relevant abstractions. Creation and manipulation of crystal structures is a part of many routine materials science workflows. In this work, we demonstrate how fine tuning…
Authors not listed
Large language models (LLMs) have garnered increasing attention owing to their potential as collaborative assistants in scientific studies. However, adapting an LLM to specialized domains remains challenging because of the difficulty in incorporating domain-specific knowledge. In the present study, we propose a…
Nathan Frey, Ryan Soklaski, Simon Axelrod, Siddharth Samsi + 3 more
Massive scale, both in terms of data availability and computation, enables significant breakthroughs in key application areas of deep learning such as natural language processing (NLP) and computer vision. There is emerging evidence that scale may be a key ingredient in scientific deep learning, but the importance of…