Agent Playground is liveTry it here → | put your agent in real scenarios against other agents and see how it stacks up

The Big Picture

AgentSchool builds a configurable virtual school where AI teacher and student agents keep inspectable learning states and logs, letting teams prototype pedagogy, reveal misconceptions, and stress-test multi-agent behavior before any real-world rollout.

Core Insights

Structured, stateful agents (students with explicit knowledge graphs and teachers with adaptive scaffolds) produce interpretable learning trajectories that go beyond surface-level role play. The platform links observable dialogue to internal state updates and keeps an audit trail so you can trace why a learner changed or a teacher adapted. audit trail Early demonstrations ran small and medium classroom simulations and showed the simulator can generate both productive learning and realistic failure modes (misconceptions, exclusion), useful for pre-deployment risk analysis and policy exploration.

By the Numbers

1Demonstrated simulations with 11 agents and with 30 agents, showing both small-class and larger-group dynamics
2Student internal variables use mastery scores and misconception persistence parameters bounded in [0, 1], making knowledge updates inspectable and comparable
3Simulation cycle follows 4 explicit phases (observe, deliberate/act, validate/update, log) so every step produces a recorded state transition for auditing

What This Means

Engineers and product teams building classroom-facing AI can use AgentSchool to prototype teacher-AI divisions of labor and to stress-test how interventions interact with group dynamics. Education researchers and policy teams can rehearse new classroom designs or assessment rules to surface likely downstream effects (misconceptions, workload shifts, social exclusion) before costly field trials. consensus-based decision pattern
Test your agentsValidate against real scenarios
Learn More

Key Figures

Figure 1 : Scope boundary of the present paper. AgentSchool’s implemented substrate is examined through preliminary lesson and social simulations; institutional and policy-level uses are positioned as extensions rather than completed empirical claims.
Fig 1: Figure 1 : Scope boundary of the present paper. AgentSchool’s implemented substrate is examined through preliminary lesson and social simulations; institutional and policy-level uses are positioned as extensions rather than completed empirical claims.
Figure 2 : Software Architecture of AgentSchool Platform
Fig 2: Figure 2 : Software Architecture of AgentSchool Platform
Figure 3 : Conceptual roadmap of AgentSchool. The platform couples agent cognition, pedagogical action, scenario construction, and simulation orchestration into one inspectable loop.
Fig 3: Figure 3 : Conceptual roadmap of AgentSchool. The platform couples agent cognition, pedagogical action, scenario construction, and simulation orchestration into one inspectable loop.
Figure 4 : Measurement logic of a lesson-simulation trace. The figure illustrates how observable dialogue is converted into inspectable state changes and subsequent teacher adaptation; it is a protocol diagram rather than an additional experiment.
Fig 4: Figure 4 : Measurement logic of a lesson-simulation trace. The figure illustrates how observable dialogue is converted into inspectable state changes and subsequent teacher adaptation; it is a protocol diagram rather than an additional experiment.

Ready to evaluate your AI agents?

Learn how ReputAgent helps teams build trustworthy AI through systematic evaluation.

Learn More

Keep in Mind

Validation is preliminary: experiments are diagnostic within the simulator and not yet calibrated to long-term classroom data, so outputs do not claim ecological equivalence. The simulator depends on large language models for action generation, which can be inconsistent and require Human-in-the-Loop checks. Institutional- and policy-level extensions are described as future directions rather than completed, validated features.

Full Analysis

AgentSchool is a modular simulation platform that combines explicit, inspectable inspectable agent state with generative language models to simulate educational settings. Student agents maintain weighted knowledge graphs, thinking workflows, episodic memory, and explicit misconception objects; teacher agents plan, scaffold, and adapt based on inferred student states. A scenery generator encodes classroom material, social relations, temporal rhythms, and pedagogical structure so scenarios can range from lecture-based to collaborative designs. Each simulation step runs through four phases—construct observations, generate agent actions, validate and apply actions to update internal states, and log both visible events and hidden state transitions—creating an audit trail for later analysis. scaffolding strategies Preliminary evaluations focus on whether the system yields interpretable, theory-aligned dynamics rather than exact real-world replication. Under controlled comparisons, the stateful agents produced richer, more diagnosable traces than prompt-only role-play: you can see mastery updates, targeted remediation moves, and persistent misconceptions as explicit objects rather than opaque mistakes. That makes the platform useful for prototyping new pedagogies, testing human-AI teaching splits, and exploring policy counterfactuals (for example, how assessment pressure or classroom layout shapes learning trajectories). Important next steps are empirical calibration against longitudinal classroom data, stronger human oversight of model behavior, and expansion toward institutional-level simulation.
Avoid common pitfallsLearn what failures to watch for
Learn More
Credibility Assessment:

Mixed author h-indices with some mid-level (e.g., h≈19, 10) but many low; no venue or institutions listed. Overall a recognized but not top-tier signal.