At a Glance
Agents that keep the same partner learn to set higher prices; randomly rematching counterparties reduces those learned prices even when no punishment or messages are possible.
ON THIS PAGE
Key Findings
When two pricing agents repeatedly face the same rival, the prices they learn are higher than when they are re-paired each episode. That gap appears even if agents cannot see the rival’s last price and cannot send messages, so higher prices are not always driven by explicit retaliation. A standard forced-deviation test that looks for punishment can miss these high-price outcomes, but checking whether a deviation would have been profitable regret-style check flags most of the learned rest points.
Data Highlights
1Hybrid agent persistence raised normalized profits by about +0.204 (contrast between persistent vs rematched pairs), well above the smallest effect of interest (0.10).
2A plain tabular value learner showed an even larger persistence effect: CΔ = +0.363 with a 95% interval [+0.215, +0.520].
3After much longer training with the rival hidden, persistent pairs ended at Δ ≈ +0.77 while the rematched arm fell to Δ ≈ −0.09 (normalized profit metric).
Implications
Engineers and teams that delegate pricing to automated agents — because platform rules about who meets whom can change the prices those agents learn. Platform operators and auditors — because rematching counterparties is a practical lever to influence outcomes, and common enforcement checks may miss these non-punishment price increases. See how the Agent Registry Pattern informs how partners are managed and matched.
Explore evaluation patternsSee how to apply these findings
Ready to evaluate your AI agents?
Learn how ReputAgent helps teams build trustworthy AI through systematic evaluation.
Learn MoreConsiderations
Results come from a specific simulated two-firm market and a limited set of agent designs and prompts, so effects may differ in other markets or with larger models. Hiding the rival also reduces the agent’s observable state, so the hidden-rival effect mixes observability and state information. The study does not measure costs or welfare tradeoffs of rematching, nor does it claim legal collusion under law. The potential for context drift is relevant to interpreting these results Context Drift.
The Details
A randomized experiment compared two settings: agents that repeatedly face the same rival (persistent partners) and agents that are re-paired every episode (rematched). The design also varied whether the rival’s last price was visible and whether a free-text channel existed. Agents used a hybrid architecture where a tabular value module set prices, and the team probed settled policies by freezing them and running a forced-deviation test (forcing one agent to undercut by four price bins) to see how rivals respond. The main outcome is a normalized profit measure Δ that sits between the competitive benchmark and joint monopoly profit. Persistent pairing consistently raised learned prices across several agent families. The effect survives conditions where punishment is impossible (rival’s price hidden and no message channel), so higher prices cannot be explained solely by learned retaliation. A standard forced-deviation probe can overstate punishment because ordinary best-response behavior also looks like retaliation; however, audits that check whether a deviation would have been profitable (a regret-style check) flag most of the high-price equilibria. Practically, marketplaces that control partner matching can influence agent pricing, but the paper does not quantify the operational costs or welfare tradeoffs of rematching, nor does it generalize to many-firm markets or different real-world deployments. rematching
Need expert guidance?We can help implement this
Credibility Assessment:
Authors have very low/unknown h-indices (e.g., h≈1) and no listed affiliations; arXiv preprint with no citations — suggests emerging/limited-info credibility.