At a Glance
Demand lifts business revenue sharply, but the cash stays with firms or in agent accounts rather than flowing to workers or customers; direct transfers are mostly saved, so short runs make the economy look frozen even though it slowly changes over longer horizons.
ON THIS PAGE
What They Found
In a fully accounted 100-agent simulation of a real town with 762 businesses, a tourist surge raised business revenue 4.62× but left wages and menu prices essentially unchanged. Direct cash transfers to households largely sat idle: about 96.7% of a grant stayed in recipients’ accounts many pulses later, implying a marginal propensity to consume near 3–4%. Wealth rankings look almost fixed over a common two-week test but drift measurably over 12–26 weeks. Which language model powers each agent changed outcomes substantially; removing agents’ memory did not. language model choices.
By the Numbers
1Business revenue rose 4.62× under high tourism, decomposed into 1.50× more businesses trading and 3.07× higher revenue per active business.
2A randomized transfer (NPR 5,000 to 20 of 100 agents) was 96.7% still held ~311 pulses later; marginal propensity to consume estimates were 3.3% and 4.0%.
3Wealth persistence (Spearman ρ) fell from 0.964 at 2 weeks to 0.832 at 12 weeks and 0.752 in a 26-week run, with the Gini index relaxing by week 16.
What This Means
Engineers building multi-agent systems should care because incentive and prompting design, not just model capacity, determine whether resources circulate. Product and platform leads evaluating agent deployments need longer test horizons and ledgered traces to avoid false stability claims. Researchers comparing agent economies should control for model choice and tool reliability when drawing distributional conclusions. Human-in-the-Loop
Not sure where to start?Get personalized recommendations
Key Figures

Fig 2: Figure 2. A run rendered over real Lakeside, Pokhara geography; markers are agents at their current place among 762 registered businesses. Map data © OpenStreetMap, via Mapbox. A screenshot of the simulation rendered over a street map of Lakeside, Pokhara, Nepal. Coloured circular markers scattered across the map show the current location of each of the 100 agents among the 762 registered businesses, which are labelled with real business names. The lake occupies the lower left of the frame. The map imagery is third-party material from OpenStreetMap via Mapbox.
Ready to evaluate your AI agents?
Learn how ReputAgent helps teams build trustworthy AI through systematic evaluation.
Learn MoreConsiderations
Long-horizon results come from few extended runs (four 12-week runs and one 26-week run), so sample size limits precision. Price stickiness is partly by design and set_price was rarely used (~18 calls per run), so lack of repricing may reflect agent prompting or the initial menu. Some failures (especially social coordination) arose from platform settings (a concurrency cap) and from model-serving quirks, so outcomes may vary with a different stack or prompt design. Reflection Pattern
Methodology & More
A closed, money-conserving economy of 100 memory-capable language-model agents was run on a realistic map of Lakeside, Pokhara, with 762 registered businesses and per-pulse ledger accounting. The authors logged 2.44M agent decisions and 21.5B tokens across 91 validated runs, including controlled tourism intensity sweeps and a within-run randomized cash transfer to 20 agents. Every rupee was tracked at pulse-level, enabling exact decomposition of where demand landed and how money moved (or did not).
Major findings: higher tourist demand raises firm revenue by 4.62×, split between more businesses trading and higher takings at active businesses, but wages and menu prices remain essentially flat. A direct cash grant is overwhelmingly saved rather than spent (MPC ≈ 3–4%). Wealth rankings look nearly fixed on the common two-week horizon but begin to relax on a roughly ten-week timescale. Which language model runs the agents substantially alters outcomes, while removing memory had no detectable effect in these tests. The simulation also exposed a near-universal failure of a social-invite tool caused largely by a platform concurrency cap. Practical implication: if you want agent economies that redistribute or generate second-round demand, fix the incentive and prompt levers that make firms, workers, and customers act differently; and evaluate over longer horizons and across model variants. The full validated dataset and analysis code are released so others can reproduce and probe intervention strategies. Emergence-Aware Monitoring Pattern Multi-Agent Government Services
Avoid common pitfallsLearn what failures to watch for
Credibility Assessment:
ArXiv preprint, zero citations, and authors/affiliations not specified or recognizable — low credibility signal per rubric.