Agent Playground is liveTry it here → | put your agent in real scenarios against other agents and see how it stacks up

The Big Picture

Manifesto-grounded, behaviorally-steered language agents can negotiate multi-party coalitions while producing an auditable lineage that predicts which provisions actually appear in real-world agreements.

The Evidence

A pipeline that (1) embeds party commitments into model behavior, (2) grounds each agent in its party manifesto, and (3) runs a structured multi-agent negotiation produces realistic coalition agreements and a transparent audit trail. Provisions traced back to manifestos are far more likely to appear in the real 2019 Flemish coalition than model-invented provisions. A simple clause-scoring system then identifies which party influenced the final agreement — matching the real-world dominant party in the case study.

Data Highlights

157.4% of negotiated provisions were directly anchored in party manifestos (direct or diluted provenance).
245.8% ± 1.5% of provisions produced by the simulator found a real-world counterpart after four negotiation rounds.
3The Coalition Influence Score attributes 40.0% ± 3.3% of the attributable clause weight to the largest party (N-VA), identifying it as the clear winner.

What This Means

Engineers and teams building multi-agent systems can use manifesto-style grounding and behavioral steering to make agent decisions auditable and less prone to invented facts. Policy analysts and political researchers can explore counterfactual coalitions and trace which party language drives outcomes without running costly human negotiations. Human-in-the-Loop Pattern
Not sure where to start?Get personalized recommendations
Learn More

Key Figures

Figure 1: Overview of the individual party alignment model process.
Fig 1: Figure 1: Overview of the individual party alignment model process.
Figure 2: Overview of the hybrid chunking strategy.
Fig 2: Figure 2: Overview of the hybrid chunking strategy.

Ready to evaluate your AI agents?

Learn how ReputAgent helps teams build trustworthy AI through systematic evaluation.

Learn More

Keep in Mind

The simulation is closed: it omits media, public opinion, and cross-level bargaining, so findings are conditioned on the single 2019 Flemish referent. Clause weights are treated equally, so influence scores don’t reflect budgetary or political stakes. A notable share (~29%) of outputs were systemic errors tied to debate-phase hallucination and need human validation for the automated labels.

Methodology & More

The system combines three practical moves to produce negotiable, party-faithful agents: break each party manifesto into coherent, auditable chunks with a hybrid deterministic plus language-model chunker; apply two-stage behavioral alignment (supervised fine-tuning followed by direct preference optimization) so each agent adopts a stable partisan persona; and use a manifesto-based retrieval layer to keep arguments grounded in party text during a hub-and-spoke negotiation. Every clause generated carries a Multi-Layered Information Lineage Topology (MILT) label recording its manifesto origin, intermediate matches, and a grounding grade against the historical 2019 coalition agreement. Across three repeated simulations and multiple evaluator passes, provisions with verifiable manifesto provenance were reproduced far more often in the historical agreement than orphaned or pipeline-invented content. The authors introduce a Coalitional Influence Score that weights grounded clauses to quantify who shaped the agreement; in the Flemish 2019 case it identifies the largest party as the primary influencer, consistent across runs. The approach offers a practical, auditable signal for agent trustworthiness in multi-agent negotiations, while highlighting remaining failure modes (notably hallucination during debate) and the need to extend validation across contexts and with human evaluators.
Avoid common pitfallsLearn what failures to watch for
Learn More
Credibility Assessment:

One author with moderate h-index (Matthias Bogaert, h=12), unspecified affiliations and arXiv venue — recognized but not strongly established.