The Big Picture
Autonomous AI agents can turn noisy social media data into clear, verifiable signals tied to a shared taxonomy—cutting investigation time and surfacing real leads (50% technique validation and 30+ new bot accounts found)—while still needing human judgement for final decisions.
ON THIS PAGE
The Evidence
A multi-agent pipeline automatically explored social media datasets, proposed technique-guided hypotheses, and converted complex findings into small, testable evidence units that map to a shared tactics-and-techniques taxonomy. Iterative rounds with full-history feedback and statistical checks produced structured outputs that experts found useful for triage and follow-up. The system validated roughly half of its technique-level findings and uncovered more than thirty previously undetected bot accounts in one dataset, demonstrating practical value as an analyst aid rather than a fully autonomous classifier. human review is required to avoid false positives and to resolve finer distinctions like bots versus coordinated humans.
Data Highlights
150% technique pass rate across 14 autonomous investigation rounds (technique-level validation)
2More than 30 previously undetected bot accounts surfaced in the Telegram dataset
384 atomic evidence claims were evaluated in total and 28 technique-level checks were performed; a single run can produce up to 42 evidence claims per dataset (15 iterations)
What This Means
Defense and intelligence analysts, open-source investigators, and engineering teams building agent-based tools should care: the pipeline standardizes outputs into a shared taxonomy, making results easier to share, prioritize, and act on. modular, language-driven workflow enables engineering leads to speed up triage and to add verifiable, testable evidence units into analyst workflows.
Not sure where to start?Get personalized recommendations
Key Figures

Fig 1: Figure 1: Workflow of the proposed agentic investigation and evidence assessment. The workflow consists of two sequential stages. In the investigation planning and execution stage (upper panel), planning proposes the hypotheses, which are subsequently examined using exploratory data analysis (EDA) and TTPs sampling over social media data and framework-based taxonomies. Planning, hypothesis-driven analysis, and evidence gathering are connected through an iterative loop, producing structured findings summaries across iterations. In the evidence assessment stage (lower panel), findings are decomposed into atomic evidence units, evaluated through automated machine verification (when labels are available), and followed by a final human evaluation.
Ready to evaluate your AI agents?
Learn how ReputAgent helps teams build trustworthy AI through systematic evaluation.
Learn MoreConsiderations
The approach focuses on behavioral signals, which are broadly applicable but lack the resolution to always separate distinct coordinated groups, human trolls, or multiple bot farms. Several metric thresholds required contextual tuning (for example message length or duplication), so out-of-the-box settings may produce false positives. The system is an analyst multiplier rather than a replacement: temporal modeling remains important for intent, long-term adaptation, and final classification.
Methodology & More
A multi-agent investigation pipeline turns the DISARM taxonomy (a standardized list of tactics and techniques) into an executable workflow. Agents perform an initial exploratory pass, then run 14 iterative rounds where each round proposes up to three small, verifiable evidence claims. Every finding is decomposed into atomic evidence units (small, testable claims) that are checked against labels or annotations using explicit statistical pass/fail criteria. Design choices include deferred anomaly detection (so findings are tied to techniques), full-history feedback across iterations, and natural-language task specifications so analysts can adapt workflows without coding.
Evaluation on two real-world datasets showed the pipeline is effective as a structured research aid: about half of technique-level proposals passed statistical checks, and the system discovered 30+ previously unknown bot accounts in one dataset. Outputs are standardized TTP mappings and reproducible evidence chains that analysts can inspect, which helps integrate results into operational processes and shared decision-making. Limitations include reliance on behavioral signals (which trade universality for detail), the need to tune thresholds to context, and the absence of a final weighted decision rule—so the pipeline is best used for prioritizing leads and supporting human-in-the-loop classification. Adding temporal modeling and content analysis would improve long-term detection and attribution.
Avoid common pitfallsLearn what failures to watch for
Credibility Assessment:
No clear institutional affiliations and low author h-indices; arXiv preprint with no citations.