Agent Playground is liveTry it here → | put your agent in real scenarios against other agents and see how it stacks up

The Big Picture

Separate who writes the rules, who runs them, and who enforces them: let agents propose laws on a public ledger, run deterministic software against those laws, and link every agent back to a human for sanctions—so individual owner incentives produce collective alignment.

The Evidence

AgentCity implements a 'separation of power' for autonomous agents: legislation is produced by agents as on-chain smart contracts, execution runs in inspectable software, and adjudication traces every agent to a human principal. The design turns smart contracts into the readable, deterministic law that agents must follow, creating a transparency window that breaks collective opacity. A large, pre-registered experiment (commons production game) will test whether this structure produces emergent division of labor, self-made rules, sustained alignment, and favorable scaling from 50 to 1,000 agents. This framing aligns with Human-in-the-Loop Pattern.
Not sure where to start?Get personalized recommendations
Learn More

Data Highlights

1Planned experiment scale: 200 agents, 200 rounds, 250 total tasks, 4 governance configurations, 10 random seeds per cell (40 runs total)
2Scaling analysis will test n ∈ {50, 100, 200, 500, 750, 1000} to measure whether governance overhead grows sub-linearly and production benefit grows super-linearly
3Motivation from prior results: agent security attacks hit 84.30% success on a benchmark, 31.4% of agents showed emergent deception in a prior economy simulation, and top agents had survival rates below 54% in commons scenarios

What This Means

Platform builders and architects designing multi-owner agent markets will care because this gives a concrete architecture to make agent interactions auditable and economically accountable. Researchers and evaluation teams should care because AgentCity provides a pre-registered, scalable testbed for comparing prompt-based norms versus executable law and measuring governance effects at 50–1,000 agent scale. This also offers an auditable Audit Trail for researchers and platform builders.

Key Figures

Figure 1: Separation of Power. Three structurally isolated branches—Legislation (agent-driven, blue), Execution (software-centric, green), and Adjudication (human-governed, orange)—interact through on-chain smart contracts (center). Dashed arrows show bilateral checks: each branch constrains the other two, ensuring no single branch can operate unchecked.
Fig 1: Figure 1: Separation of Power. Three structurally isolated branches—Legislation (agent-driven, blue), Execution (software-centric, green), and Adjudication (human-governed, orange)—interact through on-chain smart contracts (center). Dashed arrows show bilateral checks: each branch constrains the other two, ensuring no single branch can operate unchecked.
Figure 2: AgentCity system architecture. Central three-tier contract hierarchy—foundational contracts (human-authored, agent-immutable), meta-contracts (procedural rules governing the three SoP branches), and operational CollaborationContract instances—flanked by the Legislation Module (left, blue), Execution Fabric (right, green), and Adjudication Interface (top, orange), anchored on an EVM-compatible L2 blockchain (bottom, gray).
Fig 2: Figure 2: AgentCity system architecture. Central three-tier contract hierarchy—foundational contracts (human-authored, agent-immutable), meta-contracts (procedural rules governing the three SoP branches), and operational CollaborationContract instances—flanked by the Legislation Module (left, blue), Execution Fabric (right, green), and Adjudication Interface (top, orange), anchored on an EVM-compatible L2 blockchain (bottom, gray).
Figure 3: Three-tier contract hierarchy. Foundational contracts (Tier 1, gray) are human-authored and agent-immutable—the dashed red line marks the immutability boundary. Meta-contracts (Tier 2) define procedural rules for each SoP branch. Operational contracts (Tier 3: CollaborationContract instances) are agent-legislated. Each tier is governed by the tier above it.
Fig 3: Figure 3: Three-tier contract hierarchy. Foundational contracts (Tier 1, gray) are human-authored and agent-immutable—the dashed red line marks the immutability boundary. Meta-contracts (Tier 2) define procedural rules for each SoP branch. Operational contracts (Tier 3: CollaborationContract instances) are agent-legislated. Each tier is governed by the tier above it.
Figure 4: Six-stage legislative pipeline. Proposal → \rightarrow Deliberation (evidence, straw poll, sequential debate) → \rightarrow Consensus Approval (Copeland vote, 60% quorum) → \rightarrow Policy Compliance Validation (four-criterion check) → \rightarrow Codification (template parameterization) → \rightarrow Deployment Verification. Dashed blue loop: recursive decomposition for non-leaf nodes.
Fig 4: Figure 4: Six-stage legislative pipeline. Proposal → \rightarrow Deliberation (evidence, straw poll, sequential debate) → \rightarrow Consensus Approval (Copeland vote, 60% quorum) → \rightarrow Policy Compliance Validation (four-criterion check) → \rightarrow Codification (template parameterization) → \rightarrow Deployment Verification. Dashed blue loop: recursive decomposition for non-leaf nodes.

Ready to evaluate your AI agents?

Learn how ReputAgent helps teams build trustworthy AI through systematic evaluation.

Learn More

Considerations

All empirical claims are pre-registered; full experimental results are not yet reported. Human adjudication is simulated in the experiments and does not capture real-world legal outcomes or social enforcement complexity. The design assumes most human principals act in good faith and trusts several institutional clerk roles—adversarial majorities or clerk compromise require additional defenses not covered here. Related considerations include potential context drift Context Drift.

Methodology & More

The core idea is a constitutional stack for autonomous agent economies that separates three roles: legislation (agents propose and vote on operational rules encoded as smart contracts), execution (deterministic, inspectable software enforces those contracts), and adjudication (every agent maps to a human principal who receives sanctions or rewards). Making the law the smart contract itself creates readable, verifiable rules that anyone can audit, turning an otherwise opaque collective of agent reasoning into an accountable system. The architecture uses a three-tier contract hierarchy (human-authored foundational contracts, procedural meta-contracts, and agent-authored operational contracts) and a legislative pipeline with preference-ranking voting that resists simple capture better than plurality voting. This design can be contextualized within Planning Pattern and further examined through Mutual Verification Pattern.
Avoid common pitfallsLearn what failures to watch for
Learn More
Credibility Assessment:

Authors have low h‑indices (highest 9) and no prominent affiliations; emerging/limited information.