Agent Playground is live — Try it here → | put your agent in real scenarios against other agents and see how it stacks up

At a Glance

Estimating the expected benefit of a question (its value-of-information) lets an assistant ask far fewer clarifications while producing more correct, preference-aligned outcomes.

What They Found

A decision rule that weighs the expected improvement in task reward against the cost of asking leads to smarter clarification: the assistant asks only when the expected benefit outweighs the cost. Across ambiguous question answering and household planning tasks, this approach achieves higher task success and preference alignment than standard prompting, generic reasoning, and a fine-tuned clarifier, while dramatically reducing the number of questions. The method also adapts when users can correct the assistant after it acts, asking even fewer questions because corrections provide extra information. adaptive

Data Highlights

1Preference satisfaction on a household planning benchmark reached 56–59% using only in-context examples, versus 44% for a fine-tuned clarification policy (an improvement of 13–15%).
2The value-driven assistant asked roughly 5× fewer questions than the fine-tuned clarifier while improving preference alignment.
3Alternatives that optimize information gain or entropy asked 1.3–2.8× more questions than the value-driven approach for similar or worse correctness.

What This Means

Engineers building conversational assistants should care because value-based questioning reduces unnecessary user burden and improves outcomes. Product leaders and researchers evaluating assistive behavior can use this approach to trade off correctness versus user effort and to tune question costs for real deployments.
Not sure where to start?Get personalized recommendations
Learn More

Ready to evaluate your AI agents?

Learn how ReputAgent helps teams build trustworthy AI through systematic evaluation.

Learn More

Yes, But...

The method uses a one-step estimate of question value, so it can miss benefits that require multi-step question sequences. Experiments used simulated users and LLM-based scoring; real human behavior and noisy user replies may change the numbers. The approach currently assumes the assistant controls most actions (no dual-control or tightly interleaved action/question plans), so extensions are needed for fully interactive control settings. multi-step planning

Methodology & More

Treat conversational help as a decision problem: the assistant holds a belief over the user’s intent, can either ask a clarifying question or act, and incurs both task rewards (when the action matches intent) and costs for asking and answering. The key move is to estimate the value-of-information (VoI) of asking — the expected increase in task reward from the answer — and compare it to the question’s cost. Beliefs over natural-language intents are approximated by a small set of candidate intent hypotheses proposed and scored by a language model, and a proxy reward model estimates how well a candidate action satisfies each hypothesis. In practice, the value-driven policy (called REVOIR) asks far fewer questions while improving correctness on two benchmarks: ambiguous question answering and preference-aligned household planning. It outperforms baselines that either always ask, rely on generic chain-of-thought reasoning, or optimize information gain rather than downstream task reward. The method also adjusts sensibly when users can correct the assistant after an action — recognizing that cheap post-hoc corrections reduce the need to ask beforehand. Main limitations include one-step myopia, reliance on simulated users and LLM scoring for evaluation, and current lack of support for interleaving non-terminal actions with questions; these point to next steps for multi-step planning and human studies. value-of-information (VoI) LLM-as-Judge Pattern
Avoid common pitfallsLearn what failures to watch for
Learn More
Credibility Assessment:

Authors are affiliated with a top institution (MIT) and include recognizable researchers; despite being an arXiv preprint, affiliation and author reputation justify a strong (4-star) rating.