Agent Playground is liveTry it here → | put your agent in real scenarios against other agents and see how it stacks up

The Big Picture

When firms let learning-based bidding agents control market offers, those agents can learn to sustain higher-than-competitive prices without any explicit agreement—especially when transmission limits matter.

The Evidence

Autonomous bidding agents, trained through repeated interaction, learned strategies that produced sustained supra-competitive prices in a stylized electricity market. Binding transmission constraints (the grid limits) amplified the effect: cases with network limits showed much larger markups than a copper-plate (no-grid) counterfactual. Behavioral tests confirmed collusion-like patterns: if one agent tried to undercut, others temporarily drove prices down (punishment) and then returned to the higher-price regime (forgiveness). This aligns with patterns identified in Sycophancy Amplification.
Not sure where to start?Get personalized recommendations
Learn More

Data Highlights

118 market environments evaluated (2 network settings × 3 demand levels × 3 cost heterogeneity levels).
2Only Grid setups produced large supra-competitive markups; 3 out of the 18 scenarios were flagged by screening for sustained collusive behavior.
3In the 3 flagged Grid cases, post-deviation dynamics showed a punish-then-forgive pattern: competitors cut markups immediately after a deviation and later restored higher prices.

What This Means

Grid operators and market regulators should care because learning agents can raise prices without any explicit cartel or communication, calling for monitoring and rules around automated bidding. Engineers and product teams building trading or bidding agents need to test for collusion-like dynamics during pre-production and include safeguards in agent design and evaluation. Evaluation-Driven Development (EDDOps)

Key Figures

(a)
Fig 1: (a)

Ready to evaluate your AI agents?

Learn how ReputAgent helps teams build trustworthy AI through systematic evaluation.

Learn More

Limitations

Results come from a stylized seven-bus test system with fixed demand, costs, and a particular learning setup, so real-world markets may behave differently. The study focused on one class of learning dynamics and did not explore every algorithm or market rule that could prevent or amplify collusion. These findings are an early warning, not a definitive proof that all deployed agents will collude. Defense in Depth Pattern

Methodology & More

Researchers set up a repeated electricity-bidding game where generators submit piecewise cost bids and the operator clears the market using a standard optimal-power-flow rule that enforces network limits and sets local prices. Agents controlled each generator and learned bidding policies through multi-agent reinforcement learning across many repeated rounds. The experimental grid had two variants: a realistic network with line limits (Grid) and a copper-plate version with no transmission constraints (NoGrid). Scenarios varied demand level and how different generator costs were. Across 18 tested environments, supra-competitive markups were modest in the NoGrid cases but substantially larger when network constraints were active. Screening picked three Grid scenarios for deeper behavioral checks. Those checks used three indicators of tacit collusion—whether deviators are punished, whether collusion collapses if agents are shortsighted, and whether short-term deviation is profitable—and found that agents learned punish-and-forgive patterns: a deviating agent that tried to be more competitive saw rivals temporarily lower prices (hurting the deviator) before reverting to the high-price regime. The implication is that automated agents can autonomously discover coordinated pricing behavior in settings where market structure (like transmission limits) enables market power. Practical response options include stricter monitoring of agent behavior, pre-deployment testing that tracks agent interactions, and market-design changes to reduce incentives or opportunities for such emergent coordination. Event-Driven Agent Pattern routing
Avoid common pitfallsLearn what failures to watch for
Learn More
Credibility Assessment:

No institutional affiliations listed and low author h-index; arXiv preprint with no citations. Limited credibility signals, so rated as emerging/limited info (2 stars).