Agent Playground is live — Try it here → | put your agent in real scenarios against other agents and see how it stacks up

The Big Picture

How you ask an evaluation matters more than which model you use: prompt and evaluation design can create large apparent gaps between models, and a simulated reaction system can recover most real public concerns but can also amplify pessimistic bias.

The Evidence

A synthetic lab that turns a proposed change into a knowledge graph, populates simulated personas, and runs interactions can predict public concerns and recommend one of five actions. Much of the measured difference between top commercial models and open models served offline comes from under-specified evaluation prompts rather than true capability gaps. Tightening the prompt that defines the decision taxonomy raises scores across models by roughly 24–34 percentage points, and a single fine-tuned open model swung from failing to near-top performance when prompts changed. The simulation recovers 67–90% of actual public concerns but tends to push more pessimistic ("over-doom") verdicts and can align with a teacher model without increasing true accuracy. Semantic Capability Matching Pattern.
Not sure where to start?Get personalized recommendations
Learn More

Data Highlights

1One fine-tuned adapter on an open model moved from 0% to 73% success simply by changing the prompt envelope while keeping model weights and cases fixed.
2Defining the decision taxonomy in the prompt (no model change) raised every frontier model's score by +24 to +34 percentage points in a matched ablation.
3Blind judges found the simulated reactions recovered between 67% and 90% of the concerns the public actually raised across fifty real-world episodes (Gold-50).

What This Means

Product managers and policy teams who need a low-risk way to preview how users or the public might react before a rollout. Engineers and evaluation leads responsible for agent testing or agent-to-agent evaluation who want to avoid misleading comparisons caused by evaluation design. Researchers tracking multi-agent trust and how evaluation choices shape perceived agent reliability will find the methods and open pipeline useful. Mutual Verification Pattern

Ready to evaluate your AI agents?

Learn how ReputAgent helps teams build trustworthy AI through systematic evaluation.

Learn More

Yes, But...

The benchmark (50 real episodes) is compact and focused on cases with public records, so results may not generalize to all product domains or cultures. The simulation can amplify a pessimistic bias and agreement with a teacher model does not guarantee higher real-world accuracy. Results hinge heavily on prompt and scorer definitions, so any deployment should pre-register evaluation details and test sensitivity to prompt changes. Chain of Thought Pattern

Methodology & More

Augur builds a reproducible, offline "rehearsal" for product or policy changes by extracting structured facts from change documents, creating a market of grounded personas with different stakes and viewpoints, running simulated interactions between agents, and producing an auditable five-way decision memo (for example: proceed, delay, revise, pilot, or cancel). The team collected Gold-50—fifty historical product and policy episodes with known public outcomes—and scored system verdicts against the public record. The pipeline and all artifacts to reproduce the results are available from the authors. The headline finding is methodological: evaluation design—especially how the decision taxonomy and prompts are specified—can dominate measured performance. Three demonstrations show this: (1) changing only the prompt envelope made a fine-tuned open model jump from 0% to 73% on the task; (2) including a clear decision taxonomy inside the prompt raised scores by 24–34 percentage points across models; (3) the pipeline increased agreement with a distillation teacher without improving true accuracy and amplified an "over-doom" tendency. Separate human judgments validate that the synthetic reactions recover most real concerns (67–90%), and the reaction layer adds the most value where decisions are hardest. For teams building evaluation frameworks or pre-production testing for multi-agent systems, the lesson is clear: pre-specify prompts and scorers, run sensitivity tests, and watch for systematic biases the simulation might amplify. Multi-Agent Scientific Research
Avoid common pitfallsLearn what failures to watch for
Learn More
Credibility Assessment:

ArXiv preprint with authors of modest h-index (highest listed 9) and unspecified affiliations; not enough evidence of established venue or top researchers.