At a Glance
Adding a small Bayesian belief layer outside each language model gives a single per-agent knob that reliably controls how much agents change their minds and makes those changes auditable and recoverable after the AI speaks.
ON THIS PAGE
Core Insights
A compact belief layer stores each agent’s stance as a probability and updates it by a single Bayesian step when the agent hears something. One per-agent setting (the prior strength) smoothly controls behavior from full agreement to persistent disagreement to immovable anchors. After the AI generates language and another model reads it back, the prescribed ordering of stubbornness is perfectly recovered, and every belief change is logged so you can audit what was said versus what was believed.
Test your agentsValidate against real scenarios
Data Highlights
1Perfect rank recovery of stubbornness across four major models: Spearman ρ = 1.0 on every model tested.
2High language-channel fidelity: appraised speaker belief tracks latent speaker belief with Pearson r between 0.94 and 0.97.
3Classical-regime match: final beliefs align with Friedkin–Johnsen theory at R² = 0.93–0.99 across models.
What This Means
Engineers building multi-agent systems can use the belief layer to steer group behavior (from consensus to persistent disagreement) with a single, interpretable parameter. Technical leads and evaluators get audit logs and a measurable way to compare how faithfully agents communicate beliefs, improving agent-to-agent evaluation and trust signals.
Key Figures

Fig 2: Figure 2: One knob, three classical regimes ( gpt-5.4-mini ; one line per agent, averaged over seeds; all four models replicate these regimes, Tab. 2 ). (a) When everyone is pliable, opinions merge into a single consensus (the late drift is mild channel bias, App. C ). (b) Add stubborn agents and disagreement persists, settling onto the dotted theory-predicted levels. (c) Turn forgetting off and the same population collapses to consensus anyway. (d) A minority that never budges pulls the majority smoothly, with no tipping point.

Fig 3: Figure 3: Oracle γ \gamma sweep (belief mechanism only, no LLM calls): final cross-agent variance in the FJ configuration rises as soon as γ < 1 \gamma<1 and saturates near the chosen operating point, while the DeGroot configuration reaches consensus for every γ \gamma .

Fig 4: Figure 4: The belief layer as a measurement instrument. (a) Every model’s channel transfer function deviates from identity, each in its own way (GPTs: symmetric expansion; Llama: uniform downward shift; Claude: asymmetric negative-side amplification). (b) Pliable DeGroot populations compound asymmetric distortion into a confidently wrong consensus (final values labelled; reference 0.50 0.50 ). (c) Displacement from the theoretical reference: large for the pliable regime, small for the prior-anchored FJ regime on the same models—stubbornness absorbs channel bias.

Fig 5: Figure 5: Prescribe-then-recover: κ \kappa is recoverable through the language channel. Prescribed stubbornness (x) vs. value recovered from listeners’ belief movements (Eq. 3 ). Rank order is exact on all four models; magnitudes are uniformly attenuated by the channel (Tab. 3 ).
Ready to evaluate your AI agents?
Learn how ReputAgent helps teams build trustworthy AI through systematic evaluation.
Learn MoreKeep in Mind
Experiments use one synthetic topic, 20 agents on a fully connected graph, and binary stances, so results may not generalize to richer topics, larger populations, or realistic social networks. The appraiser that turns text into evidence is itself a language model calibrated on only 100 labeled examples, so directional distortions in reading or rendering can persist. The pipeline needs two language-model calls per utterance, which raises cost and may limit exact reproducibility without similar API access. language-model
Deep Dive
The approach inserts a lightweight belief layer between each agent’s persona and its language output. Each agent keeps a probability for a concept outside the language model; when an utterance is heard, a separate appraiser model converts the text into scalar evidence and the agent updates its probability via one Bayesian step. A single per-agent parameter (the prior strength) controls how much each update moves that agent—small values make agents pliable and lead to consensus, medium values reproduce persistent disagreement, and very large values create committed anchors that drag others. In experiments with 20 agents on a complete graph and four different large language models, the authors sweep the prior-strength knob and show three practical benefits: (1) prescribed stubbornness is recoverable in rank order after the full language round-trip (perfect Spearman rank recovery), (2) the per-agent knob reproduces classical opinion regimes closely (high R² against theoretical references), and (3) every belief update is logged so you can audit whether what an agent said matched what it actually believed. The language channel does introduce systematic distortions—magnitudes are attenuated and some models skew evidence asymmetrically—so the belief layer both enables control and surfaces model-specific communication bias. Next steps include multi-topic identities, correcting directional channel bias, and testing on richer network structures. three practical benefits
Not sure where to start?Get personalized recommendations
Credibility Assessment:
All authors have low h-index (3–4), no notable affiliations listed, and the paper is only an arXiv preprint with no citations — fits an emerging/limited-info profile.