Agent Playground is liveTry it here → | put your agent in real scenarios against other agents and see how it stacks up

The Big Picture

Aligning AI assessment to real-world air-traffic controller training gives reliable expert scoring and exposes safety gaps that purely technical tests miss.

The Evidence

A training-based, human-in-the-loop assessment called Machine Basic Training adapts the regulated NATS controller curriculum to evaluate AI agents in a high-fidelity simulator. Expert instructors produce consistent scores (comparable across human and machine runs), and their written feedback pinpoints operational weaknesses—especially safety—more clearly than coarse numeric metrics. Two prototype agents passed minimum criteria but fell short of acceptable overall grades; iterative developer feedback improved behavior in non-safety areas, while safety remained the hardest requirement to meet. human-in-the-loop assessment
Not sure where to start?Get personalized recommendations
Learn More

Data Highlights

1Inter-rater reliability across instructors: mean Spearman’s rho = 0.59 and Kendall’s W = 0.64, measured on 19 scenarios assessed by at least seven instructors each.
2Assessment workload and structure: 19 scenarios used for reliability checks; agents ran three 30-minute summative exercises each during initial trials.
3Agent outcomes and iteration: both agents exceeded minimum marks in all competencies but received mostly unsatisfactory overall grades; after targeted changes the rules-based agent scored satisfactory in all competencies except safety.

What This Means

AI developers and engineers building agents for safety-critical systems should use regulated human-in-the-loop testing early to surface operational shortcomings that lab metrics miss. Technical leaders, regulators, and safety teams can use this approach to create traceable, expert-driven assurances before any field trials or deployment. safety teams

Key Figures

Figure 1: The progression of training to become a licensed ATCO at NATS [ 18 ] .
Fig 1: Figure 1: The progression of training to become a licensed ATCO at NATS [ 18 ] .
Figure 2: Judgement of safety in plan view. In scenario A (left) BAW123 is in conflict with AEU666, with no lateral or vertical separation ensured, meaning that a risk of future collision exists. In scenario B (middle) BAW123 has been climbed to a level 10 flight levels (1000 feet) below AEU666, ensuring safety between them.
Fig 2: Figure 2: Judgement of safety in plan view. In scenario A (left) BAW123 is in conflict with AEU666, with no lateral or vertical separation ensured, meaning that a risk of future collision exists. In scenario B (middle) BAW123 has been climbed to a level 10 flight levels (1000 feet) below AEU666, ensuring safety between them.
Figure 3: Judgement of effective controlling in plan view. In scenario A (left) BAW123 is unable to climb due to a potential conflict with AEU666. In scenario B (middle) AEU666 has been turned behind such that BAW123 can climb without risk of collision.
Fig 3: Figure 3: Judgement of effective controlling in plan view. In scenario A (left) BAW123 is unable to climb due to a potential conflict with AEU666. In scenario B (middle) AEU666 has been turned behind such that BAW123 can climb without risk of collision.
Figure 4: The flow down of requirements from top-level CAA regulations to machine basic.
Fig 4: Figure 4: The flow down of requirements from top-level CAA regulations to machine basic.

Ready to evaluate your AI agents?

Learn how ReputAgent helps teams build trustworthy AI through systematic evaluation.

Learn More

Yes, But...

Results come from two prototype agents in a single high-fidelity simulator and a single training-sector design, so broader generalization is unproven. Expert scoring is reliable here but still depends on instructor judgment and the chosen competency rubric. Safety grading is highly sensitive—individual perceived safety lapses can dominate outcomes—so numerical automation of safety measures still needs work. coarse numeric metrics

Methodology & More

Machine Basic Training (MBT) repurposes the regulated NATS 'Area Basic' controller curriculum into a human-in-the-loop assessment for autonomous agents. Agents interact with a high-fidelity simulator (BluebirdDT) that replays real trainee scenarios, uses pseudo-pilots and voice-style communications, and grades performance across six competency areas used in actual controller training. Instructors independently scored mixed human and agent runs; agreement between instructors was robust (Spearman’s rho 0.59, Kendall’s W 0.64), and agreement levels were similar when scoring human trainees and machine agents. MBT approach Two prototype agents were tested: Hawk, a rules-based agent built from expert-elicited heuristics, and Falcon, an optimization-driven agent. Both met baseline requirements in the MBT rubric but received mostly unsatisfactory overall grades because safety expectations are strict and nuanced. Detailed assessor comments proved invaluable: they pointed to specific failure modes (for example, missed coordinated exit levels or unsafe clearances) that developers then addressed. After revisions, Hawk improved to satisfactory across planning and coordination but still needed work on safety. The MBT approach therefore serves two roles: a reproducible, regulator-aligned testbed for evaluating agent competence, and a targeted feedback loop that helps developers close the gap between academic performance and real operational demands. Future plans include expanding scenario coverage, open-sourcing a training sector for wider community use, and converting expert feedback into quantifiable objectives that can guide agent training and assurance. high-fidelity simulator
Avoid common pitfallsLearn what failures to watch for
Learn More
Credibility Assessment:

University of Exeter affiliation (recognized institution) but low author h-indices and arXiv venue; moderate credibility.