In Brief
Retroactively deleting accounts overstates how much misinformation bans will help in real time; when the network can adapt, remaining users resharing dampens the gains from removing a few high-profile accounts.
ON THIS PAGE
Key Findings
A simulator calibrated on nearly a year of Italian COVID-vaccine Twitter data reproduces real activity patterns and the identity of influential spreaders well enough to stand in for the real platform. Four different methods for ranking misinformation spreaders produce similar removal effects on the simulated and empirical datasets, validating the simulator as a testbed. However, when account bans are applied during a running simulation (so users can adapt), the drop in low-quality reshares is smaller than what static, retroactive removal predicts, meaning static evaluation gives an optimistic upper bound. simulation-based evaluation framework
Not sure where to start?Get personalized recommendations
By the Numbers
1819,952 posts and reshares from 74,234 unique users used to calibrate the simulator.
250 independent simulation runs on 90,000-node networks reproduced temporal behavior; short-lag temporal autocorrelation matched the real data within 2% at lags 1–5.
330 paired network experiments showed a consistent static minus dynamic gap when banning the top-5 users (bans applied at day 186), with statistical testing (Wilcoxon signed-rank) confirming the difference across networks.
Why It Matters
Platform engineers and product teams designing moderation tools should care because static tests can overpromise the benefit of account removal—simulation-based, real-time testing gives a more realistic estimate. Event-Driven Agent Pattern Policy and trust-and-safety leaders can use calibrated simulators to explore adaptive responses before committing to disruptive actions like mass bans.
Key Figures

Fig 2: Figure 2: Targeted user removal: remaining fraction of LQ reshare weight as users are removed one by one in ranked order, for real (solid) and simulated (shaded CI) data.
Ready to evaluate your AI agents?
Learn how ReputAgent helps teams build trustworthy AI through systematic evaluation.
Learn MoreLimitations
The simulator deliberately omits external news shocks, so it reproduces baseline dynamics but not time-aligned event-driven spikes. Results come from one long-running, language-specific dataset (Italian COVID-vaccine discussions) and may not generalize to other languages, topics, or platform designs. The intervention studied is a single, one-shot account ban (top-5); results do not directly speak to content labels, demotion, staggered removals, or platform-wide policy mixes. context drift
Methodology & More
A calibrated agent-based simulator was built on a longitudinal dataset of Italian-language COVID-19 vaccine conversations collected over about a year (≈819k actions, 74k users). User activity rates were fitted to a flexible heavy-tailed distribution and content quality was modeled as a two-mode mixture to match the observed credibility scores. The simulator was validated across many dimensions—temporal patterns, activity and reshare ratios, and the set of influential spreaders—by running 50 independent simulations and comparing them to real data. Four established methods for ranking misinformation spreaders were applied to both the real data and the simulated runs; targeted removal tests (network dismantling) on simulated data produced effectiveness curves that closely match empirical results, supporting the simulator's intervention-level validity. Using 30 paired networks, the study compared two ways of measuring the effect of banning accounts: static retroactive removal from a completed dataset versus dynamic removal during a running simulation. Static removal consistently overstates the reduction in low-quality reshares because the remaining users adapt their resharing behavior, partially compensating for the lost accounts. The practical takeaway: validated simulations are a useful precautionary tool to forecast adaptive, real-time responses to moderation policies and to avoid overconfident conclusions drawn from retrospective experiments. The study's design aligns with the Model Context Protocol (MCP) Pattern. The approach also reflects Emergence-Aware Monitoring Pattern.
Avoid common pitfallsLearn what failures to watch for
Credibility Assessment:
One author (Luca Luceri) has a solid h-index (~20) but other authors have low h-indices and no affiliations listed; mixed signals suggest a moderate (recognized) credibility level.