Research

Microsoft researchers find AI degrades in multiturn chats

Microsoft researchers are exposing how AI performance degrades during complex, multiturn interactions, a finding that could reshape how developers build and evaluate agentic systems.

Microsoft Research Blog1 day agoResearch
Image: Microsoft Research Blog

Microsoft Research's AI Interaction and Learning team, led by Partner Research Manager Jennifer Neville, is shifting the focus of artificial intelligence evaluation from static benchmarks to realistic, long-horizon workflows. In a recent podcast interview, Neville and Principal Applied Scientist Chad Atalla discussed how current large language models struggle when tasks are stretched across multiple turns or collaborative editing sessions. The team's findings reveal that while models excel at highly optimized, single-turn benchmarks, their performance degrades significantly when users attempt to clarify or build upon instructions over time.

To study this phenomenon, Neville's team adapted public single-turn benchmark datasets. They used simulated users to break down fully specified instructions and deliver them as gradual clarifications over multiple conversational turns. The experiment exposed a stark contrast: models that achieved high accuracy on single-turn tasks frequently lost track of the context during multiturn interactions. For practitioners and users, Neville offers a simple workaround: if a model becomes confused after a long exchange, users should clear the chat history and feed the fully refined prompt back to the system in a single turn.

The team also investigated agentic knowledge work, where AI assistants perform repeated edits on documents over long-horizon workflows. In these scenarios, errors accumulate over time, leading to a subtle loss of semantic content that eventually confuses the AI. Neville noted that these "surprising failures" often occur because human expectations of AI capability are misaligned with how transformer architectures actually compute and retrieve information. Tasks that humans find simple, or tasks that a model successfully completes once, can still fail unexpectedly in subsequent iterations.

To identify these failure patterns without compromising user privacy, Microsoft analyzes consumer logs at scale using privacy-preserving methods that prevent researchers from directly viewing the data. The AI Interaction and Learning team is currently developing reinforcement learning methods to help models maintain context and learn more effectively in multiturn environments. Until these algorithmic improvements are integrated, Neville cautions that AI systems cannot be fully automated and that users must actively verify outputs.

This is our own summary of reporting by Microsoft Research Blog

More in Research