Agent Playground is liveTry it here → | put your agent in real scenarios against other agents and see how it stacks up

The Big Picture

Using only a paper’s pre-publication references, single models rarely recover the original idea (~3–15%), but combining different models with cross-review and selection raises recovery to about 23–42%.

The Evidence

Given only the list of prior papers cited before publication, modern large language models seldom reconstruct the actual research idea on their own—best single-model performance sits in the low teens. Running multiple models, letting them critique each other, and selecting the strongest candidates lifts successful recovery substantially. The test uses a strict time cutoff and anonymized references so models cannot see the target paper; matches are judged by independent models with recusal rules to avoid self-evaluation.
Not sure where to start?Get personalized recommendations
Learn More

Data Highlights

1643 seed papers across six domains (machine learning plus five Nature-family journals) were evaluated.
2Best single-model average Match rate: 13.3% ± 2.3% (domain scores typically ~3–15%).
3Reference-only cross-model review plus Swiss selection achieved Match rates ≈ 23–42%, an observed ~2.4× lift over the best single-model baseline.

What This Means

Engineers building agent ensembles can use these results to justify multi-model review and selection when the goal is deep literature understanding. Technical leaders and product owners evaluating agent reliability should treat reconstruction as a stress test before trusting agents for idea generation or scholarly assistance. Researchers studying machine understanding and agent coordination can use the benchmark to compare approaches under strict anti-leakage controls.

Ready to evaluate your AI agents?

Learn how ReputAgent helps teams build trustworthy AI through systematic evaluation.

Learn More

Considerations

Models may partially recover papers due to memorized content from pretraining rather than true inference from the provided references, despite the temporal cutoff. Match scoring relies on model judges with leave-one-out and recusal rules; human agreement with the binary match rubric is not yet reported. The multi-agent setup uses more total compute and selects 5 winners from a larger candidate pool, so some of the gain may come from having more candidates rather than purely better collaboration mechanics.

Methodology & More

Reconstruction frames a strict challenge: given only a paper’s pre-publication bibliography (titles and abstracts, anonymized and time-cut), can an AI recover the idea that actually appeared in the held-out paper? For each seed paper, systems must produce five distinct hypotheses tied to supporting reference IDs. Independent model judges—who do not originate the hypothesis they evaluate—compare each hypothesis to the seed paper’s title and abstract and return a binary match label. The dataset covers 643 papers across six domains (ICML and five Nature-family journals), and the protocol enforces a hard temporal cutoff and other anti-leakage measures. Single-model baselines using several frontier models achieve low match rates (best mean ≈13%), showing that individual models struggle to infer an author's contribution from citations alone. A multi-agent pipeline—where multiple models draft hypotheses, cross-review each other, and a Swiss-style selection picks five champions from the combined pool—raises recovery to roughly 23–42%. The authors caution that some of this improvement may derive from picking among 20 candidates rather than collaboration alone, judges differ in permissiveness, and models could exploit pretraining exposure. Still, the benchmark timestamps an important point: heterogeneous model ensembles plus reference-grounded critique and selection substantially increase the chance of recovering a real research idea from only its references. The ultimate goal is to transfer these harness principles to forward-looking idea generation, not just reconstruction.
Avoid common pitfallsLearn what failures to watch for
Learn More
Credibility Assessment:

Some authors have moderate citation impact (e.g., Rahul Thapa h-index ~14) and at least one recognizable name, but overall low/unspecified affiliations and arXiv venue point to a solid but not top-tier credibility.