Agent Playground is liveTry it here → | put your agent in real scenarios against other agents and see how it stacks up
Back to Ecosystem Pulse
EvaluationExperimentalMCP

multi-agent-workflow-lab

by christiangrey922

Replayable testing and observability for multi-agent delegation

TypeScript
Updated Aug 14, 2026
Share:
86
Stars
72
Forks

View on GitHub

Overview

Provides testing and observability tooling for multi-agent delegation workflows, including sandboxed actions, permission controls, and prompt/workflow replay. Combines MCP-aware instrumentation with replayable traces so teams can reproduce interactions and inspect agent decision points. Includes hooks for permission checks and sandboxed execution to test attack/failure modes in isolation. MCP-aware instrumentation and sandboxed actions with replayable traces.

Why It Matters

As agents delegate more to one another, reproducible evidence of what happened and why becomes essential for trust. This project makes it possible to simulate, record, and replay agent-to-agent interactions so you can surface failure modes, check permissions, and build empirical agent track records. That capability is a foundational step toward continuous A2A evaluation and actionable reputation signals. A2A evaluation.

Best For

Teams building or auditing multi-agent systems who need reproducible traces, sandboxed action testing, and delegation/permission validation.

Use Cases

  • Reproduce and debug complex agent delegation failures using workflow replay
  • Validate and fuzz-check agent permissions and sandboxed actions before production
  • Log and inspect agent interactions to build agent track records for reputation systems
Topics
agent-orchestrationagent-securityagent-testingai-agentai-agentsai-securitymcpmodel-context-protocolmulti-agent
Similar Tools
agent-playgroundagent-arena
Keywords
multi-agent trustA2A evaluationagent-to-agent evaluationmcp