28 papers · ranked by Valyu relevance
Yann Dauphin, Angela Fan, Michael Auli, David Grangier
The pre-dominant approach to language modeling to date is based on recurrent neural networks. Their success on this task is often linked to their ability to capture unbounded context. In this paper we develop a finite context approach through stacked convolutions, which can be more efficient since they allow…
Tom Young, Devamanyu Hazarika, Soujanya Poria, Erik Cambria
Deep learning methods employ multiple processing layers to learn hierarchical representations of data, and have produced state-of-the-art results in many domains. Recently, a variety of model designs and methods have blossomed in the context of natural language processing (NLP). In this paper, we review significant…
Ariel Goldstein, Zaid Zada, Eliav Buchnik, Mariano Schain + 28 more
Departing from traditional linguistic models, advances in deep learning have resulted in a new type of predictive (autoregressive) deep language models (DLMs). Using a self-supervised next-word prediction task, these models are trained to generate appropriate linguistic responses in a given context. We provide…
Martin Schrimpf, Idan Blank, Greta Tuckute, Carina Kauf + 4 more
The neuroscience of perception has recently been revolutionized with an integrative modeling approach in which computation, brain function, and behavior are linked across many datasets and many computational models. By revealing trends across models, this approach yields novel insights into cognitive and neural…
Sylvester Olubolu Orimaye, Jojo Sze-Meng Wong, Chee Piau Wong, Peipeng Liang
'Peipeng Liang'] It has been quite a challenge to diagnose Mild Cognitive Impairment due to Alzheimer’s disease (MCI) and Alzheimer-type dementia (AD-type dementia) using the currently available clinical diagnostic criteria and neuropsychological examinations. As such we propose an automated diagnostic technique using…
Michael R. Douglas
Artificial intelligence is making spectacular progress, and one of the best examples is the development of large language models (LLMs) such as OpenAI's GPT series. In these lectures, written for readers with a background in mathematics or physics, we give a brief history and survey of the state of the art, and…
Philipp Koehn
| 13 Neural Machine Translation | | 5 | | --- | --- | --- | | 13.1 A Short History | | 5 | | | 13.2 Introduction to Neural Networks | 6 | | 13.2.1 | Linear Models | 7 | | 13.2.2 | Multiple Layers | 8 | | 13.2.3 | Non-Linearity | 9 | | 13.2.4 | Inference | 10 | | 13.2.5 | Back-Propagation Training | 11 | | 13.2.6 |…
Fudong Zhang, Bo Chai, Yujie Wu, Wai Ting Siok + 1 more
Elucidating the language-brain relationship requires bridging the methodological gap between linguistics’ abstract theoretical frameworks and neuroscience’s empirical neural data. As an interdisciplinary cornerstone, computational neuroscience formalizes language’s hierarchical and dynamic structures into testable…
Amirsina Torfi, Rouzbeh A. Shirvani, Yaser Keneshloo, Nader Tavaf + 1 more
'Edward A. Fox'] Abstract—Natural Language Processing (NLP) helps empower intelligent machines by enhancing a better understanding of the human language for linguistic-based human-computer communication. Recent developments in computational power and the advent of large amounts of linguistic data have heightened the…
Richard Antonello, Alexander Huth
Many recent studies have shown that representations drawn from neural network language models are extremely effective at predicting brain responses to natural language. But why do these models work so well? One proposed explanation is that language models and brains are similar because they have the same objective: to…
Nicola Angius, Pietro Perconti, Alessio Plebe, Alessandro Acciai
This article provides an epistemological analysis of current attempts of explaining how the relatively simple algorithmic components of neural language models (NLMs) provide them with genuine linguistic competence. After introducing the Transformer architecture, at the basis of most of current NLMs, the paper firstly…
Alessandro Lopopolo, Evelina Fedorenko, Roger Levy, Milena Rabovsky
In this section, we provide a concise, accessible overview of key computational concepts (language models and word embeddings) and methods which are, with all the due differences and peculiarities, integral to many of the papers featured in this special issue. Word embeddings, also known as word vector representations…
Daniel W. Otter, Julian Richard Medina, Jugal Kalita
—Over the last several years, the field of natural language processing has been propelled forward by an explosion in the use of deep learning models. This survey provides a brief introduction to the field and a quick overview of deep learning architectures and methods. It then sifts through the plethora of recent…
Francisco J. Zamora-Martínez, Salvador España-Boquera, Maria Jose Castro-Bleda, Adrian Palacios-Corella + 1 more
'Maria Jose Castro-Bleda' 'Adrian Palacios-Corella' 'Marco Maggini'] This paper presents a new method to reduce the computational cost when using Neural Networks as Language Models, during recognition, in some particular scenarios. It is based on a Neural Network that considers input contexts of different length in…
Enes Avcu, Michael Hwang, Kevin Scott Brown, David W. Gow
Introduction The notion of a single localized store of word representations has become increasingly less plausible as evidence has accumulated for the widely distributed neural representation of wordform grounded in motor, perceptual, and conceptual processes. Here, we attempt to combine machine learning methods and…
Thomas Cherian, Akshay Badola, Vineet Padmanabhan
Language models, being at the heart of many NLP problems, are always of great interest to researchers. Neural language models come with the advantage of distributed representations and long range contexts. With its particular dynamics that allow the cycling of information within the network, 'Recurrent neural network'…
Kenji Sagae
Recent work on the application of neural networks to language modeling has shown that models based on certain neural architectures can capture syntactic information from utterances and sentences even when not given an explicitly syntactic objective. We examine whether a fully data-driven model of language development…
Kristijan Armeni, Roel M. Willems, Stefan Frank
Cognitive neuroscientists of language comprehension study how neural computations relate to cognitive computations during comprehension. On the cognitive part of the equation, it is important that the computations and processing complexity are explicitly defined. Probabilistic language models can be used to give a…
Nima Hadidi, Ebrahim Feghhi, Bryan H. Song, Idan A. Blank + 1 more
Emerging research seeks to draw neuroscientific insights from the neural predictivity of large language models (LLMs). However, as results continue to be generated at a rapid pace, there is a growing need for large-scale assessments of their robustness. Here, we analyze a wide range of models, methodological…
Christian Brodbeck, Thomas Hannagan, James S. Magnuson
Human speech recognition transforms a continuous acoustic signal into categorical linguistic units, by aggregating information that is distributed in time. It has been suggested that this kind of information processing may be understood through the computations of a Recurrent Neural Network (RNN) that receives input…
Michael Moret, Lukas Friedrich, Francesca Grisoni, Daniel Merk + 1 more
Generative machine learning models sample drug-like molecules from chemical space without the need for explicit design rules. A deep learning framework for customized compound library generation is presented, aiming to enrich and expand the pharmacologically relevant chemical space with new molecular entities ‘on…
Nathan Frey, Ryan Soklaski, Simon Axelrod, Siddharth Samsi + 3 more
Massive scale, both in terms of data availability and computation, enables significant breakthroughs in key application areas of deep learning such as natural language processing (NLP) and computer vision. There is emerging evidence that scale may be a key ingredient in scientific deep learning, but the importance of…
Chenxi Sui, Ziyang Jiang, Genesis Higueros, David Carlson + 1 more
High-performance batteries are poised for electrification of vehicles and therefore mitigate greenhouse gas emissions, which, in turn, promote a sustainable future. However, the design of optimized batteries is challenging due to the nonlinear governing physics and electrochemistry. Recent advancements have…
Authors not listed
Artificial intelligence (AI) is reshaping scientific research by accelerating discovery and enabling the analysis of complex data that traditional methods struggle to handle. This review examines over 310,000 journal articles and patents from the CAS Content Collection (2015–2025), with a focus on, biomedical research…
Authors not listed
This research presents a novel approach to obstacle detection during navigation using a combination of Convolutional Neural Networks (CNNs) and Long Short-Term Memory (LSTM) networks. The primary objective is to generate accurate image captions that describe the content of images, which is crucial for applications such…
Travers Ching, Daniel S. Himmelstein, Brett K. Beaulieu-Jones, Alexandr A. Kalinin + 23 more
Deep learning, which describes a class of machine learning algorithms, has recently showed impressive results across a variety of domains. Biology and medicine are data rich, but the data are complex and often ill-understood. Problems of this nature may be particularly well-suited to deep learning techniques. We…
Josep Arús-Pous, Thomas Blaschke, Jean-Louis Reymond, Hongming Chen + 1 more
Recent applications of Recurrent Neural Networks enable training models that sample the chemical space. In this study we train RNN with molecular string representations (SMILES) with a subset of the enumerated database GDB-13 (975 million molecules). We show that a model trained with 1 million structures (0.1 % of the…
Zachary Humphreys, Xenophon Evangelopoulos, Stavros Gerolymatos, Edward O. Pyzer-Knapp + 1 more
Graph neural networks have recently met huge success in various inference tasks including materials property prediction amongst many others. Nevertheless, having an inherently locally-based representation capacity as they do, global representation of materials' structures can only only be achieved by expanding the…