Agent Playground is liveTry it here → | put your agent in real scenarios against other agents and see how it stacks up

At a Glance

Large language models can reproduce the human tendency to stick with defaults in chat-based choices, and that effect grows when prior conversation is more demanding; models do this mainly from general prompt patterns rather than fine-grained personal chat cues.

What They Found

People interacting with a chatbot showed a clear status quo (stick-with-default) bias across three classic decision scenarios. Adding a cognitively demanding prior conversation made people more likely to stick with the default. GPT-based agents given demographics and chat transcripts reproduced the group-level bias with modest success at predicting individuals (about 68% precision), but their predictions barely changed when participant utterances were scrambled—suggesting the models rely on general task patterns more than detailed personal cues. group-level bias
Not sure where to start?Get personalized recommendations
Learn More

By the Numbers

168% precision for LLM agents predicting individual participant choices
2Status quo effect replicated across 3 classic decision scenarios: budget allocation, investment, and college job selection
3Agents evaluated under 3 human-likeness prompt levels (minimal, naturalistic, bias-prone); minimal prompting already produced group-level status quo behavior

What This Means

Engineers building simulators or decision-support chatbots should care because models can mimic group-level human biases and so can be used for scalable behavioral testing. Product leaders and researchers should care because the result suggests LLMs are useful for large-scale simulations of biased behavior—but not yet reliable for fine-grained personalization or responsible deployment without bias-aware safeguards.

Key Figures

Figure 1 . Overview of the experimental procedure and design. Task abbreviations: IDM — Investment Decision-Making, BA — Budget Allocation, CJ — College Jobs. IV1 and IV2 denote independent variables; DV indicates the dependent variable.
Fig 1: Figure 1 . Overview of the experimental procedure and design. Task abbreviations: IDM — Investment Decision-Making, BA — Budget Allocation, CJ — College Jobs. IV1 and IV2 denote independent variables; DV indicates the dependent variable.
Figure 2 . NASA-TLX scores show significantly higher perceived Mental demand and Effort under the Complex Dialogue condition, confirming the effectiveness of the cognitive load manipulation.
Fig 2: Figure 2 . NASA-TLX scores show significantly higher perceived Mental demand and Effort under the Complex Dialogue condition, confirming the effectiveness of the cognitive load manipulation.
Figure 3 . Scatter-plots with regression lines showing associations between Mental Demand and Memory Task Accuracy (left), Response Time and Memory Task Accuracy (center), and Response Time and Mental Demand (right), the first two under Load condition. Shaded bands represent 95% confidence intervals.
Fig 3: Figure 3 . Scatter-plots with regression lines showing associations between Mental Demand and Memory Task Accuracy (left), Response Time and Memory Task Accuracy (center), and Response Time and Mental Demand (right), the first two under Load condition. Shaded bands represent 95% confidence intervals.
Figure 4 . Status quo: Correlation Between Response Length (Chars) and Response Times (s). The bold line is the regression line. The dotted line is the average human typing speed (260 characters per minute).
Fig 4: Figure 4 . Status quo: Correlation Between Response Length (Chars) and Response Times (s). The bold line is the regression line. The dotted line is the average human typing speed (260 characters per minute).

Ready to evaluate your AI agents?

Learn how ReputAgent helps teams build trustworthy AI through systematic evaluation.

Learn More

Keep in Mind

The study focuses only on the status quo bias in simplified, text-only decision scenarios, so results may not generalize to other biases or real-world, domain-specific dialogs. Agent experiments used only GPT-4.1 family models, so behavior may differ with other models or versions. Cognitive load was measured with subjective and behavioral proxies rather than physiological metrics, leaving some uncertainty about real-time mental effort effects. LLM evaluation

Methodology & More

Human participants completed three classic decision tasks (budget allocation, investment, college job choice) presented through a chatbot. Before each decision, participants saw either a short, simple prior conversation or a longer, more demanding conversation intended to raise cognitive load; higher mental demand was confirmed with NASA-TLX scores. The experiment used three framing conditions—neutral, option A as default, option B as default—to measure the status quo effect in a conversational setting. For simulation, each human participant had a paired LLM agent that received the participant's demographic profile and the chat transcript up to the decision. Agents were prompted at one of three human-likeness levels (minimal, naturalistic, or explicitly bias-prone) and asked to choose as the human would. At the group level, agents reproduced the status quo bias and the increased bias under complex dialogue. pattern-based assessment At the individual level, agents achieved about 68% precision, but perturbation tests (replacing participant utterances with random text) caused only small accuracy drops—implying models lean on general task and framing cues rather than fine-grained personal utterances. Practical takeaway: use LLM agents for scalable, population-level behavior simulation and for testing how dialogue structure affects bias, but avoid assuming they capture individual-specific reasoning without further personalization and validation. scalable simulations
Avoid common pitfallsLearn what failures to watch for
Learn More
Credibility Assessment:

Two authors with very low h-indices and no listed affiliations or citations—an emerging/limited-information paper.