006.35
Language Models
Reasoning, alignment, agents and the science of LLMs.
Drawer contents
Filled from
Search threads
large language model reasoning
LLM alignment and evaluation benchmark
006.35
Reasoning, alignment, agents and the science of LLMs.
Drawer contents
Filled from
Search threads
large language model reasoning
LLM alignment and evaluation benchmark
Ian B. de Haan, Peter van der Putten, Max van Duijn
Large language models (LLMs) have recently shown strong performance on Theory of Mind (ToM) tests, prompting debate about the nature and validity of the underlying capabilities. At the same time, reasoning-oriented LLMs trained via reinforcement learning with verifiable rewards have demonstrated notable improvements…
Eduardo Valle, Fergal Reid
We show that tiny transformers can profitably employ a simple form of Chain of Thought, which we call protoreasoning, allowing us to study step-by-step reasoning on ~1M-parameter models and opening up opportunities for much more detailed experimentation and analysis than is feasible for larger models. Current Large…
Noam Koren, Roy Bar-Haim, Abigail Goldsteen
Task-oriented conversational agents are evaluated using curated or automatically generated benchmarks, yet benchmark quality is rarely assessed. Poor benchmarks may contain inconsistent tasks, simplistic scenarios, or limited policy coverage, leading to unreliable evaluations. We introduce a reference-free framework…
Hamed Damirchi, Ignacio Meza De la Jara, Damith Ranasinghe, Yuhang Liu + 1 more
As language models are increasingly used for tasks that require verifiable reasoning, reliably distinguishing sound reasoning from flawed reasoning has become an important practical problem. Recent trajectory-based methods seek this signal in layerwise residual-stream displacements, which capture how representations…
Xingyu Guo, Wei Chen, Linlin Yang, Baochang Zhang
Search agents extend large language models beyond static parametric memory by enabling them to acquire and use ex ternal evidence during multi-step reasoning. For knowledge intensive tasks involving complex or evolving information, their reliability depends not only on retrieving relevant ev idence but also on using it…
Jiaoyang Li, Junhao Ruan, Shengwei Tang, Kaiyan Chang + 3 more
Large language models (LLMs) often generate inaccurate answers due to their reliance on static internal knowledge. Retrieval-augmented generation (RAG) addresses this limitation by integrating external knowledge and excelling at single-hop queries. However, it struggles with multi-hop questions that require…
Eunbi Choi, Kibong Choi, Sehyun Chun, Seokhee Hong + 73 more
This technical report presents K-EXAONE 2.0, an open-weight multilingual foundation model developed by LG AI Research as a step in our effort toward global frontier-scale foundation models. Rather than training from scratch, we upcycle K-EXAONE and expand its architecture, yielding a Mixture-of-Experts (MoE) model with…
Jonas Gann, Michael Gertz
Retrieval-augmented generation (RAG) improves question answering by grounding large language models (LLMs) in external knowledge such as text corpora. However, its reasoning process remains largely opaque: intermediate reasoning steps are difficult to verify and cannot be reliably attributed to specific evidence.…
Yangfan Jiang, Fei Wei, Ergute Bao, Xiaokui Xiao + 2 more
Direct preference optimization (DPO) is now a standard method for aligning large language models (LLMs) using human preference data. Each DPO example contains a prompt and a pair of candidate model responses. While prompts and responses are often public or model-generated, the relative preference between responses…
Jiahao Zhang, Yongzhi Tong, Zelin Fu, Pengde Zhao + 3 more
Existing personalized LLM benchmarks primarily rely on textual personas or isolated behavioral signals, providing limited evaluation of cross-domain behavioral personalization, where responses must be grounded in heterogeneous daily-life activities. To address this gap, we introduce LUNAR, the first benchmark for…
Mukhtiar Ali, Harsh Dubey, Sugam Mishra, Chulwoo Pack
Benchmarking video-language models has largely focused on short clips and single-sentence metrics, leaving open whether current systems can generate accurate long-form, paragraph-level descriptions. We introduce CLIP-CC-Bench, an evaluation suite for long-form video description built from 5 hours of movie content…
Réemi Andrieu, Damien Sileo
Reasoning about necessity and possibility depends on assumptions about accessibility between worlds and about which objects exist at each one. The same inference may therefore hold under one modal system and fail under another. Evaluating language models on such problems requires testing whether their judgments follow…
Keane Zhang, Varshini Chinta, Raj Sanjay Shah, Sashank Varma
Anaphors are expressions that refer to other expressions, called antecedents. The process of connecting the two is called resolution. Cognitive science has identified multiple factors that affect the speed and success of anaphor resolution, including discourse structure, situation-model properties, and semantic…
Yuma Asato, Kiyoaki Shirai, Natthawut Kertkeidkachorn
Large Language Models (LLMs) are often used as evaluators of text quality, known as LLM-as-a-Judge, which can outperform conventional automatic evaluation metrics that rely on reference texts. However, LLM evaluators tend to generate particular scores regardless of the context of the evaluated text, which is known as…
Tao Wang, Qihao Yang, Rongjiao Liang, Lianghong Lin + 3 more
Large language models (LLMs) increasingly support complex professional tasks, yet their capabilities in rule-intensive document review remain insufficiently evaluated. National standard documents, such as China GB/T standards, offer a representative testbed: they are lengthy, highly structured, and governed by explicit…
Shahed Masoudian, Passant Shafaei, Monorama Swain, Markus Schedl
Benchmark scores are reported as properties of a model, yet the inference framework used to produce them, such as HuggingFace, vLLM, or Ollama, are considered non-influential and their names and versions are almost never disclosed. In this work we investigate how much this choice can influence the model output. In a…
Shahd Gaben, Heba Sbahi, Samer Rashwani, Abdessalam Bouchekif + 4 more
Large language models (LLMs) are increasingly used for question answering, education, and research, including in religious and cultural domains where answers depend on specialised source traditions. Yet in Islamic Studies, key concepts, methods, and debates preserved in the authoritative scholarly tradition, known as…
Divyansh Singh
LLM judges are often asked to extract criteria and evidence before choosing between candidate answers. This workflow assumes that the intermediate record preserves the information needed for a later verdict. For reasoning-capable models, visible field order does not reveal internal decision order, so we test an…
Ziyun Zeng, Zixuan Wang, Yongsheng Yu, Hang Hua + 1 more
Evaluating generated videos remains challenging because existing benchmarks rely on fixed evaluation content, cover only a subset of generation and editing settings, and provide limited evidence for their scores. We introduce VideoArgus, a unified rubric-grounded framework covering five video generation and editing…
Pedro Ferreira, Wilker Aziz, Ivan Titov
Chain-of-thought (CoT) reasoning offers a window into the decision-making of large language models (LLMs), which can be monitored for target behaviors by reading the reasoning trace, motivating work on CoT monitorability. Latent CoT approaches, however, replace the explicit tokens with a small number of continuous…
Germana Bertoli, Ilaria Amelia Caggiano, Francesca Lagioia, Riccardo Rovatti + 2 more
The article reports on a blind Turing Test experiment, assessing the performance of out-of-the-box leading LLMs on three Italian legal professional exams: the Bar, Judges and Notary exams. Leading LLMs were asked to generate full written exam papers, which were made indistinguishable from human submissions and…
Enrico Mensa, Lorenzo Zane, Calogero Jerik Scozzaro, Matteo Delsanto + 2 more
Large Language Models (LLMs) have transformed computational linguistics and achieved remarkable performance across numerous natural language processing tasks, yet significant gaps persist in understanding how these systems process culturally embedded linguistic expressions. This paper introduces ProverbIT, a novel…
Yixuan Wang, Licheng Luo, Yu Fu, Kaidi Xu + 2 more
Translating natural language instructions into machine-interpretable formal specifications enables robots and autonomous systems to plan, reason, and formally verify their behavior. However, existing translation models typically generate a specification for every input, even when the result is unreliable or fails to…
Sajib Hossain, Md Kamrus Samad, Anan Ghosh, Labib Imam Chowdhury + 1 more
Deep neural networks have shown impressive success in NLP tasks owing to their complex structure and huge number of edges. Achieving state-of-the-art performance in natural language processing with a large pre-trained model such as BERT is expensive and time-consuming, carries a large carbon footprint, and is difficult…