Key Takeaway
Representing procedural knowledge as an explicit, editable graph gives language-model agents clear, situation-aware guidance and reliably improves performance across diverse tasks; the graph can also be refined automatically from experience.
ON THIS PAGE
What They Found
Organize steps, actions, and states as a directed graph where each edge carries when-to-act notes, how-to guidance, and common pitfalls. At runtime, a small guidance model reads the graph around the agent’s current state and turns it into next-step advice; offline, a refiner suggests structural edits based on success and failure traces. This approach boosts success rates across tasks and models and can build or repair useful procedures from minimal or flawed starting points.
Explore evaluation patternsSee how to apply these findings
Data Highlights
1Procedural Graph ranked first or tied for first in 21 of 24 model–benchmark settings.
2Against the strongest baseline per setting: 19 wins, 2 ties, 3 losses (one-sided binomial test p = 4.3×10⁻⁴).
3Largest gains included a +9.00 point jump (67.00% vs 58.00%) on a function-calling benchmark and several +6–7 point improvements on long-horizon professional tasks.
Why It Matters
Engineers building multi-step AI agents will get clearer, editable control over allowable actions and fewer repeated failures. Technical leads and researchers evaluating agent reliability can use these graphs to inspect, version, and automatically improve procedural knowledge without retraining models.
Key Figures

Fig 2: Figure 2: Overview of the Procedural Graph framework. Left: procedural triplets define 𝒢 \mathcal{G} . Middle: the framework localizes u t u_{t} and retrieves its 2 2 -hop neighborhood 𝒢 t \mathcal{G}_{t} (or the full graph if matching fails); a guidance LLM translates it into guidance g t g_{t} for the solver. Right: the refiner proposes edits from execution trajectories. Structurally valid candidates are committed when validation performance does not decrease; rejected candidates inform subsequent proposals through rejection memory.

Fig 3: Figure 3: Ensemble cash trajectories and Kaplan–Meier survival curves across four LLMs. Bold lines and shaded areas show means and 95 % 95\% confidence intervals. Vertical lines mark macroeconomic crises. Colors identify PG (blue), the baseline (red), and memory-based methods (green, orange).

Fig 4: Figure 4: Mean lifespan and capital raised across ten rounds of PG self-evolution. Gray dashed lines show training results; red lines show validation results; and green diamonds show test results for the baseline and accepted checkpoints.
Ready to evaluate your AI agents?
Learn how ReputAgent helps teams build trustworthy AI through systematic evaluation.
Learn MoreConsiderations
Generating guidance from the graph increases token use and may raise per-task cost even when it reduces planning steps. Transferability across different agent architectures and tool interfaces was not fully evaluated, so reuse between systems may require adaptation. Some benchmarks showed small or no gains, so expect variable improvements depending on task structure and available tools.
Deep Dive
Turn procedural graph whose nodes represent actions, reasoning steps, tools, or states, and whose edges encode allowed transitions plus three textual attributes: when the transition applies, how to proceed, and what to avoid. During solving, a guidance model finds the agent’s current node, reads a small neighborhood of the graph, and translates those edge attributes into situational next-step guidance. After batches of runs, an automated refiner contrasts successful and failed traces to propose topology and attribute edits; candidates are validated on held-out tasks and only accepted if performance does not drop, while rejected edits are recorded to avoid repeat mistakes.
Across seven benchmarks and four families of language models, the procedural graph consistently outperformed memory-style baselines and often beat hand-designed workflows. The self-evolution loop can produce effective graphs from scratch and fix flawed expert priors, meaning teams can start small and improve procedures from execution feedback without changing model weights. Practical trade-offs include higher token consumption and the need to validate transfer across different solvers and toolsets, but the approach offers a transparent, versionable way to make agent behavior more reliable and debuggable in production settings.
Avoid common pitfallsLearn what failures to watch for
Credibility Assessment:
One author, Sercan Ö. Arık, is a known researcher in industry (Google/ML). That affiliation/reputation increases credibility despite arXiv publication.