Paraphernalia
AarXiv30 Oct 2024Cited 2×

Multi-Agent Large Language Models for Conversational Task-Solving

J. K. Becker

Abstract

In an era where single large language models have dominated the landscape of artificial intelligence for years, multi-agent systems arise as new protagonists in conversational task-solving. While previous studies have showcased their potential in reasoning tasks and creative endeavors, an analysis of their limitations concerning the conversational paradigms and the impact of individual agents is missing. It remains unascertained how multi-agent discussions perform across tasks of varying complexity and how the structure of these conversations influences the process. To fill that gap, this work systematically evaluates multi-agent systems across various discussion paradigms, assessing their strengths and weaknesses in both generative tasks (i.e., summarization, translation, and paraphrase type generation) and question-answering tasks (i.e., extractive, strategic, and ethical question-answering). Alongside the experiments, I propose a taxonomy of 20 multiagent research studies from 2022 to 2024, followed by the introduction of a framework for deploying multi-agent LLMs in conversational task-solving. I demonstrate that while multi-agent systems excel in complex reasoning tasks, outperforming a single model by leveraging expert personas, they fail on basic tasks like translation. Concretely, I identify three challenges that arise: problem drift, alignment collapse, and monopolization. Multi-agent systems discuss more difficult examples for longer until they reach a consensus, adapting to the complexity of a problem. While longer discussions enhance reasoning, agents fail to maintain conformity to strict task requirements, which leads to problem drift, making shorter conversations more effective for basic tasks. However, prolonged discussions also risk alignment collapse, raising new safety concerns for these systems. The discussion format and personas impact individual agents' response length. Moreover, I showcase discussion monopolization through long generations, posing the problem of fairness in decision-making for tasks like summarization.

A figure from Multi-Agent Large Language Models for Conversational Task-Solving
fig. from the paper

§ The Valyu brief

Reading the full paper and taking notes. This takes a few seconds…

§ Ask this paper

Ask a question about this paper

Valyu reads the full text and answers from what the paper actually says.

Q.

Searching the other archives…

Multi-Agent Large Language Models for Conversational Task-Solving · Paraphernalia