The Big Picture
Single-run results often misrepresent an idea: 25–44% of one-run winners reverse when the same idea is independently reimplemented, so automated research should test multiple realizations before crediting an idea.
ON THIS PAGE
The Evidence
Implementation choices (the ordinary coding decisions left open by a research idea) create far more outcome variation than rerunning the same saved artifact. In a controlled audit on 13 tabular tasks, independent reimplementations caused decision reversals in 25.6% to 43.6% of one-draw winners. Best-of-N search Tool Use Pattern can produce a high-scoring artifact, but that artifact’s score does not reliably prove the underlying idea unless it survives additional independent translations. Hallucination propagation Hallucination Propagation can also undermine reliability across translations.
Data Highlights
1One-draw winner reversal (leave-one-out) occurred in 25.6% and 43.6% of decisions across two implementation setups.
2Implementation choices explained 33% and 45% of within-task variation, versus only 6% and 4% for rerun noise.
3Estimated implementation variance was >5× rerun variance in the 'Bounded' setup and >10× in the 'Agentic' setup.
Why It Matters
Engineers and teams automating research or deploying agents should care because one high-scoring run can mislead branching, transfer, and what the system remembers. Technical leaders and benchmark designers should require multiple independent implementations before treating a run as evidence about a mechanism rather than just an artifact.
Need expert guidance?We can help implement this
Key Figures

Fig 2: Figure 2: Recommended prospective audit workflow. Portfolios should be frozen and undergo outcome-blind semantic validation before implementation. Our study applies the semantic gate after execution as a sensitivity (Section 5 ). Three fresh sessions implement each admitted card, fidelity is labeled outcome-blind, executable artifacts receive byte-identical reruns, and all attempts remain in ITT. The audit reports within-task idea ICC and LOO winner reversal.
Ready to evaluate your AI agents?
Learn how ReputAgent helps teams build trustworthy AI through systematic evaluation.
Learn MoreYes, But...
Results come from 13 medium-size tabular classification tasks, a fixed boosting baseline, three independent implementations per idea, and two implementation processes—so rates are conditional on this setup. The audit used outcome-blind fidelity checks but not human reviewers for all labels, and it did not measure large deep-learning workflows where implementation choices multiply. Best-of-N artifact search remains valid for delivering a working artifact; the concern is attributing that artifact’s score to a stable idea. State Inconsistency
Methodology & More
The study introduces an "idea reliability" perspective: an idea is a mechanism described at a semantic level, and different valid implementations can translate that idea into code and results. The proposed audit freezes mechanism-level descriptions (cards), then asks independent teams or sessions to implement each card three separate times. Each implementation is run under the same evaluation policy, fidelity is labeled without seeing outcomes, and saved artifacts are rerun byte-for-byte to separate rerun noise from implementation variation. Idea-level metrics reported are an intraclass correlation (how much ideas separate from implementation noise) and leave-one-out winner reversal (whether a one-draw winner survives the other implementations). Cascading Reliability Failures
Not sure where to start?Get personalized recommendations
Credibility Assessment:
All authors have very low h-indices and no affiliations listed; arXiv preprint with zero citations suggests limited established credibility (emerging/limited info).