Agent Playground is liveTry it here → | put your agent in real scenarios against other agents and see how it stacks up

The Big Picture

Model rankings change depending on whether the AI must do the work or guide someone else — pick models by role and task, not by a single leaderboard.

The Evidence

Models that excel at completing tasks on their own often perform differently when asked to provide process guidance to another worker. Rankings between automation (doing the task) and augmentation (helping a fixed worker) show only moderate correlation and large task-specific swings. Guidance can help, do nothing, or even hurt: in several tasks a cheap unaided worker beat every assisted condition.guardrails pattern

Data Highlights

1Overall rank correlation between automation and augmentation is about 0.48, showing only moderate agreement between the two modes.
2Task-level correlations vary from −0.04 (travel planning) up to 0.85 (tax preparation), so mode alignment depends heavily on the task.
3Example reversal: Claude-Opus-4.8 averaged a 2.05 rank as an automator but fell to 8.15 as an assistant on the same task, showing large mode-specific performance flips.

Why It Matters

Engineers building AI assistants and teams deciding which models to deploy should test models in the exact role they will play (doer vs. coach). Technical leaders and product managers should run role-specific evaluations before standardizing on a single model for diverse workflows.
Need expert guidance?We can help implement this
Learn More

Key Figures

Figure 1 : Methodology pipeline. The framework evaluates models in two usage modes: automation, where focal models directly solve tasks, and augmentation, where focal models provide an assistance text to a fixed worker model. Outputs are evaluated through rubric-guided pairwise comparisons and aggregated into model rankings by task and usage mode.
Fig 1: Figure 1 : Methodology pipeline. The framework evaluates models in two usage modes: automation, where focal models directly solve tasks, and augmentation, where focal models provide an assistance text to a fixed worker model. Outputs are evaluated through rubric-guided pairwise comparisons and aggregated into model rankings by task and usage mode.
Figure 3 : Prompt and rubric design. Each task prompt specifies observable deliverable requirements, which are mirrored by task-specific micro-rubrics. General rubric dimensions are held constant across tasks, while task-specific dimensions vary by domain. In augmentation mode, the assistance text creation prompt constrains assistant models to provide process guidance rather than the final deliverable.
Fig 3: Figure 3 : Prompt and rubric design. Each task prompt specifies observable deliverable requirements, which are mirrored by task-specific micro-rubrics. General rubric dimensions are held constant across tasks, while task-specific dimensions vary by domain. In augmentation mode, the assistance text creation prompt constrains assistant models to provide process guidance rather than the final deliverable.
Figure 5 : Rubric-guided pairwise evaluation pipeline. Candidate outputs are anonymized, randomly paired, and evaluated by LLM judges using task-specific and general rubric dimensions. Pairwise choices are converted into win rates, aggregated across eligible judges and repeated runs, and reported alongside rubric scores and rationales as an audit trail.
Fig 5: Figure 5 : Rubric-guided pairwise evaluation pipeline. Candidate outputs are anonymized, randomly paired, and evaluated by LLM judges using task-specific and general rubric dimensions. Pairwise choices are converted into win rates, aggregated across eligible judges and repeated runs, and reported alongside rubric scores and rationales as an audit trail.
Figure 6 : Rubric-guided scoring. Judge rationales explain the pairwise choice made. The models used in this example are GPT-4.1 and DeepSeek-V3.1 in augmentation mode for the counseling task. The corresponding outputs are in Appendix G .
Fig 6: Figure 6 : Rubric-guided scoring. Judge rationales explain the pairwise choice made. The models used in this example are GPT-4.1 and DeepSeek-V3.1 in augmentation mode for the counseling task. The corresponding outputs are in Appendix G .

Ready to evaluate your AI agents?

Learn how ReputAgent helps teams build trustworthy AI through systematic evaluation.

Learn More

Yes, But...

Results are for single-turn, bounded tasks without tool use and may not generalize to multi-step agent pipelines that call external tools. Augmentation trials used a fixed lower-capability worker (GPT-3.5-Turbo), so outcomes would differ with a different downstream worker or human. Evaluation used LLM judges with leave-family-out masking; while validated, judge choices can still reflect model-of-judge biases. For discussion of how complex pipelines can introduce uncertainty, see multi-step agent pipelines.

Methodology & More

The study evaluates ten contemporary language models across seven real-world professional tasks in two modes: automation (the model produces the deliverable) and augmentation (the model provides process guidance to a fixed worker, GPT-3.5-Turbo, which produces the deliverable). Every output was judged via blind, rubric-guided pairwise comparisons by a panel of LLM judges, with judges barred from evaluating outputs from their own model family. Repeating the pipeline across ten runs produced per-task rank profiles for each model in both modes. Findings show that automation and augmentation are distinct capabilities. A single model rarely dominates as an assistant across tasks, and mode-specific rankings diverge substantially: overall rank correlation is only ~0.48 and task correlations span from about −0.04 to 0.85. Some strong direct solvers perform poorly as coaches, and in three tasks the unaided GPT-3.5-Turbo worker outperformed all assisted conditions. Practical takeaway: organizations should use semantic-capability-matching role- and task-specific evaluations (and can swap in their own worker proxy or human) to choose the right model for the right role rather than relying on a single general-purpose leaderboard. human-in-the-loop pattern
Not sure where to start?Get personalized recommendations
Learn More
Credibility Assessment:

ArXiv preprint, authors lack notable affiliations and have low h-indices (max h=2); fits an emerging/limited-info profile.