Multi-Agent Large Language Models for Conversational Task-Solving
J. K. Becker
Abstract
In an era where single large language models have dominated the landscape of artificial intelligence for years, multi-agent systems arise as new protagonists in conversational task-solving. While previous studies have showcased their potential in reasoning tasks and creative endeavors, an analysis of their limitations concerning the conversational paradigms and the impact of individual agents is missing. It remains unascertained how multi-agent discussions perform across tasks of varying complexity and how the structure of these conversations influences the process. To fill that gap, this work systematically evaluates multi-agent systems across various discussion paradigms, assessing their strengths and weaknesses in both generative tasks (i.e., summarization, translation, and paraphrase type generation) and question-answering tasks (i.e., extractive, strategic, and ethical question-answering). Alongside the experiments, I propose a taxonomy of 20 multiagent research studies from 2022 to 2024, followed by the introduction of a framework for deploying multi-agent LLMs in conversational task-solving. I demonstrate that while multi-agent systems excel in complex reasoning tasks, outperforming a single model by leveraging expert personas, they fail on basic tasks like translation. Concretely, I identify three challenges that arise: problem drift, alignment collapse, and monopolization. Multi-agent systems discuss more difficult examples for longer until they reach a consensus, adapting to the complexity of a problem. While longer discussions enhance reasoning, agents fail to maintain conformity to strict task requirements, which leads to problem drift, making shorter conversations more effective for basic tasks. However, prolonged discussions also risk alignment collapse, raising new safety concerns for these systems. The discussion format and personas impact individual agents' response length. Moreover, I showcase discussion monopolization through long generations, posing the problem of fairness in decision-making for tasks like summarization.

§ The Valyu brief
Reading the full paper and taking notes. This takes a few seconds…
§ Ask this paper
Ask a question about this paper
Valyu reads the full text and answers from what the paper actually says.
Searching the other archives…