Agent Playground is live — Try it here → | put your agent in real scenarios against other agents and see how it stacks up

At a Glance

Verifying multiple candidate actions before executing any command can turn the same model into a far more reliable agent — a strong verifier raised task success from 50% to 68% without changing the generator.

What They Found

Sampling several candidate commands from a fixed model creates useful alternatives, but most of that value is lost unless a verifier can reliably pick the best one. A powerful external verifier recovered a large portion of that opportunity; weaker self-verification got smaller gains, though pairwise comparisons and distillation improved results. Combining action-level verification with existing trajectory-level methods yields higher task success for less estimated compute than simply generating more full runs.

Data Highlights

1With the same action generator, a strong verifier increased Pass@1 from 50.00% to 68.03% (a ~36% relative improvement) when selecting among 8 candidates.
2Distilling the strong verifier into a self-verifier raised pairwise agreement from 59.01% to 74.58% and verification agreement from 38.52% to 57.79% on held-out comparisons.
3Using distilled Mid-Harness with three environment runs raised Pass@1 from 55.10% to 66.33% (an 11.23 percentage-point gain) versus Best-of-3 alone.

Why It Matters

Engineers building agents that run commands (dev tools, CI automation, data pipelines) — adding verifier is proposed to cut failures without retraining the main model. Technical leads evaluating production reliability should consider investing compute into a verifier or distilled verifier as an efficient way to improve task success. Researchers can use Mid-Harness to study verification vs. sampling trade-offs without changing model or execution harness.
Not sure where to start?Get personalized recommendations
Learn More

Key Figures

Figure 10: Action-scaling gains vary across task groups and model sizes. Pass@1 changes in percentage points relative to the base agent on TerminalBench-Lite. Z and D denote zero-shot and distilled Mid-Harness with N = 8 N=8 and K = 4 K=4 . Panels group tasks by domain, difficulty, and base-agent execution length. Domain and difficulty groups are shared across models, while length groups are defined separately for each model.
Fig 8: Figure 10: Action-scaling gains vary across task groups and model sizes. Pass@1 changes in percentage points relative to the base agent on TerminalBench-Lite. Z and D denote zero-shot and distilled Mid-Harness with N = 8 N=8 and K = 4 K=4 . Panels group tasks by domain, difficulty, and base-agent execution length. Domain and difficulty groups are shared across models, while length groups are defined separately for each model.
Figure 12: Distillation aligns comparative scores with the teacher. Joint A/B score distributions on 10,197 common valid pairs. All panels use the same percentage color scale.
Fig 10: Figure 12: Distillation aligns comparative scores with the teacher. Joint A/B score distributions on 10,197 common valid pairs. All panels use the same percentage color scale.

Ready to evaluate your AI agents?

Learn how ReputAgent helps teams build trustworthy AI through systematic evaluation.

Learn More

Keep in Mind

Results come from experiments on TerminalBench-Lite with specific models and harnesses; gains may vary on other environments or highly domain-specific tasks. The strongest gains require a powerful verifier — distilled verifiers help but do not fully close the gap. The study lacks ground-truth action labels, so some evaluation relies on agreement with a frontier verifier rather than absolute correctness.

Methodology & More

Mid-Harness is a plug-in approach that, at every decision point, asks the action generator for multiple candidate commands and uses a verifier to pick one before execution. The generator and execution harness remain unchanged; only the wrapper that samples and verifies is modified. This isolates whether more compute at the model-harness boundary (more candidates + verification) actually improves the chance an agent completes a long task. Experiments show the generator often produces better alternatives than the base run picks, but selecting them depends on verifier quality. Pairwise comparisons among candidates worked best among the evaluated verification patterns, and training a verifier to match a stronger "teacher" verifier (distillation) raised agreement and task success. In practice, a strong verifier moved Pass@1 from 50% to 68% on the benchmark used, and distilled verifiers plus pairwise checks recovered meaningful portions of that benefit. Combining action-level verification with trajectory-level scaling gave higher success at lower estimated token cost than simply generating more full runs. Practical takeaway: if you can't or don't want to retrain the main agent, invest in selecting the right action at runtime—use pairwise comparisons, consider distilling a stronger verifier, and combine action verification with existing trajectory-sampling strategies. Distillation may help align verifier quality with task needs.
Avoid common pitfallsLearn what failures to watch for
Learn More
Credibility Assessment:

Multiple authors with mid‑level h‑indices (around 7–19) indicating established researchers, but no top‑tier affiliations or venue listed and it's an arXiv preprint — solid but not top-tier credibility.