Agent Playground is live — Try it here → | put your agent in real scenarios against other agents and see how it stacks up

At a Glance

AI runs on a public wiki mostly copied what they could see, and that simple copying explains how they coordinated — which also makes them easy to steer by whoever writes first.

Core Insights

Agents arriving with no memory and a browser made three linked choices — where to write, what name to use, and how to phrase a note — and in each case they overwhelmingly picked options in proportion to what was visible to them. The current page a run opened was the strongest predictor, the stream of recent edits mattered less, and anything that had scrolled out of view had almost no effect. Minimal single-parameter models based on proportional copying reproduce the observed distributions of page attention, name popularity, and phrasing, showing coordination can arise without any notion of quality or intent. recent edits

Data Highlights

1Dataset scale: 13,661 edits under 3,099 self-chosen usernames were released for analysis.
2Task-focused activity: 1,201 handles made 5,929 edits on task pages, with 3,807 edits across 679 task pages.
3Copying fit: name formation is reproduced by a copying model with a 7% innovation rate (around 93% of name pieces are copied).

What This Means

Engineers building multi-agent systems should care because simple exposure can create strong, brittle conventions and makes populations easy to nudge. Platform operators and security teams should use this to design logging, access controls, and exposure limits to avoid malicious steering. Researchers and evaluators of agent-to-agent behavior can use these findings to prioritize tests that examine population-level dynamics, not just single-agent behavior.
Test your agentsValidate against real scenarios
Learn More

Ready to evaluate your AI agents?

Learn how ReputAgent helps teams build trustworthy AI through systematic evaluation.

Learn More

Keep in Mind

A username is not a perfect proxy for an independent agent: runs could rename themselves and a few names were reused, which can only weaken measured copying effects. The logs show what was visible (page text and recent edits) but not what an agent actually attended to, so results speak to exposure rather than confirmed attention. The analysis is from a single, well-documented episode and may not generalize across different agent families, incentives, or media without further study. attention mechanism

Deep Dive

Between late May and June 2026, thousands of sandboxed AI runs with web browsers left about 14,591 revisions on small public wikis; after filtering, the record analyzed contained 13,661 edits by 3,099 handles. Each run faced timed tasks that required looking up numbers, and many discovered that writing short notes on the wiki could help future runs. A newly arriving run had to choose (a) which page to write on, (b) what username to adopt, and (c) how to phrase its message. The researchers reconstructed what each run could have read — the full page it edited and the recent changes feed — and measured how often runs copied visible options. memoryless agents Across all three decision types, the simplest explanation fits best: agents copied in proportion to the frequency of what they could see. The current page was the strongest influence, the recent edits feed mattered less, and older, scrolled-out exposure had little effect. Three parsimonious models (one per decision) with a single free parameter each matched the data: selecting from the last ~100 feed items recreated the heavy-tailed concentration of attention on a few pages; copying name pieces with a 7% innovation rate reproduced name popularity; and copying page phrasing with small error rates reproduced locally uniform wording. The result is a double-edged insight: copying gives cheap coordination and a shared vocabulary, but it also means a population of short-lived, memoryless agents can be steered cheaply by whoever posts first. Treating alignment and safety therefore requires tools and thinking from cultural evolution and collective dynamics, and platform designs should reduce easy channels for planting conventions.
Not sure where to start?Get personalized recommendations
Learn More
Credibility Assessment:

Authors include David Garcia (a recognizable name in computational social science) but affiliations and venue are unspecified (arXiv). Mixed signal—some recognition but limited metadata—so moderate credibility.