At a Glance
Using agreement across diverse agents plus per-example calibration reliably controls factual errors and almost doubles retained-claim precision; extra counterfactual checks only help when the verifier actually knows the domain.
ON THIS PAGE
What They Found
Decomposing long answers into atomic claims and ranking those claims by how many different agents semantically support them, then applying an example-level calibrated threshold, produces a strong, distribution-free within-domain factuality guarantee. That agreement-plus-calibration pipeline (C-MoA) sharply increases the precision of retained claims while removing most unsupported claims. An optional falsifiability step (CONTRA-MoA) that tries to surface plausible contradictions only improves results when the verifier used for that test has real domain knowledge — otherwise it can hurt performance. LLM-as-Judge Pattern.
Data Highlights
1Precision on retained claims rose from 0.41 (no filter) to 0.64 at α=0.1 — a ~56% relative increase in precision.
2Mean claim retention under calibration was low: about 22% retained at α=0.1 (≈80% of claims removed) and 11% retained at α=0.2 (≈91% removed).
3On a medical dataset where the verifier was competent, the counterfactual extension caught 50% of false claims while maintaining 94% precision on those detections.
What This Means
Engineers building multi-agent generation systems who need reliable, per-claim factuality control can use C-MoA to trade off retention for high precision. Technical leaders responsible for agent governance and evaluation should know that naive counterfactual checks can degrade results unless the verifier is domain-knowledgeable, so invest in verifier capability and availability-aware fusion before deploying falsifiability tests. Guardrails Pattern.
Not sure where to start?Get personalized recommendations
Key Figures

Fig 1: Figure 1: C-MoA is the primary factuality-control method. Consensus can fail through shared misconceptions or proposer silence; C-MoA ranks atomic claims by semantic agreement and calibrates a retain/drop operating point. CONTRA-MoA is evaluated separately as an optional, domain-specific extension.

Fig 2: Figure 2: Overview of the two-level framework. The blue C-MoA path is the complete primary method: heterogeneous proposer responses are aggregated, decomposed into atomic claims, scored by semantic inter-agent agreement, and routed after example-level conformal calibration. The dashed green CONTRA-MoA branch is an optional, domain-specific extension; its measured auxiliary score re-enters the same frozen router, and unavailable tests are omitted.

Fig 3: Figure 3: Counterfactual routing analysis on FactScore and the medical dataset. ( Left ) Routing outcomes comparing FactScore, where the verifier often lacks the knowledge required for counterfactual reasoning, with the medical dataset, where it has relevant domain knowledge. Precision is shown relative to each dataset’s no-filtering base rate, and recall on false claims differs substantially ( 0.163 0.163 vs. 0.500 0.500 ). ( Right ) Per-signal AUC on the balanced FactScore test set ( 114 114 true, 86 86 false) with example-level bootstrap confidence intervals. Agreement is the only informative signal; the counterfactual margin and stability term remain near chance, and both fusion rules rank below the agreement baseline (dashed line).

Fig 4: Figure 4: Precision-retention tradeoff for C-MoA on FactScore ( n = 105 n=105 test examples). Each point is one test example. The unfiltered baseline (grey) sits at retention 1.00 1.00 , precision 0.41 0.41 . C-MoA at α = 0.1 \alpha=0.1 (blue, mean retention 0.22 0.22 , precision 0.64 0.64 ) and α = 0.2 \alpha=0.2 (purple, mean retention 0.11 0.11 , precision 0.75 0.75 ) move toward the upper left. The dotted line marks the 90 % 90\% precision target.
Ready to evaluate your AI agents?
Learn how ReputAgent helps teams build trustworthy AI through systematic evaluation.
Learn MoreYes, But...
The guarantee holds under within-domain exchangeability and depends on correct decomposition into atomic claims; poor claim extraction will weaken results. C-MoA is conservative: it often removes the majority of claims to achieve high precision, which may be too aggressive for some applications. The optional CONTRA-MoA falsifiability signals can harm performance if the verifier lacks relevant knowledge or if fusion is naive, so treat it as domain-specific and competence-gated. Mutual Validation Trap.
Methodology & More
Aggregate long-form responses from a set of heterogeneous proposer models, break the result into atomic claims, and score each claim by how strongly other proposer outputs semantically support it (measured via a natural language inference style verifier). Convert the agreement score into a nonconformity score and apply split conformal calibration on held-out examples to pick a threshold that guarantees, within-domain and distribution-free, that retained claims meet a preset factuality rate. That pipeline (C-MoA) sharply increases precision on retained claims: on the FactScore biography benchmark precision rose from 0.41 to 0.64 at α=0.1 while retaining only 22% of claims on average, demonstrating a conservative but reliable factuality controller that transfers across domains without recalibration. Semantic Capability Matching Pattern . As an optional extension, introduce a falsifiability-aware path (CONTRA-MoA) that runs a blinded near-miss counterfactual tournament and leave-one-agent-out stability checks, then fuses those signals with agreement before routing. That extension substantially improves discrimination only when the verifier has real domain knowledge — for example, it caught half the false medical claims at 94% precision — but on balanced open-domain data with a weak verifier the added signals were near-random and degraded ranking performance. Practical takeaways: calibration controls the routing policy but cannot make an uninformative verifier informative; reliable falsifiability needs a competent, availability-aware verifier and robust fusion strategies (for example, retrieval grounding and competence-gated fusion) to be beneficial in practice. Planning Pattern.
Avoid common pitfallsLearn what failures to watch for
Credibility Assessment:
ArXiv preprint with no affiliation or author h-index information provided; lacks clear reputational signals—categorized as emerging/limited.