Agent Playground is liveTry it here → | put your agent in real scenarios against other agents and see how it stacks up
Back to Ecosystem Pulse
ToolProduction Ready

superset

by superset-sh

Run 100+ isolated coding agents in parallel for reproducible comparison

TypeScript
Updated Aug 17, 2026
Share:
13.0k
Stars
1.2k
Forks

View on GitHub

Summary

Runs hundreds of coding agents in parallel with isolated worktrees so each agent has its own filesystem and environment. Orchestrates different model backends (Claude Code, Codex, CLI agents) and captures outputs for side-by-side comparison and reproducible debugging. Distinctive features include per-agent isolated workspaces, BYO subscriptions, and tooling focused on large-scale, parallel code generation experiments.

Why It Matters

As teams evaluate and compare coding agents, being able to run many variants in identical, isolated conditions is essential for fair assessment and reproducibility. This repo makes large-scale side-by-side experiments practical, surfacing behavioral differences, failure modes, and agent track records. For multi-agent trust and A2A evaluation, it provides the controlled execution layer needed to collect comparable signals and build reputation over time.

Best For

Researchers and engineering teams who need to run large-scale, reproducible experiments comparing coding agents and model backends.

Applications

  • Compare output quality across model backends (Claude Code, Codex, etc.) under identical conditions
  • Stress-test and surface agent failure modes by running many agents in parallel
  • Collect reproducible traces and artifacts per-agent using isolated worktrees for debugging and reputation tracking
Works With
openaianthropiccodexclaudecli
Topics
adeagentagent-orchestrationai-agentsai-codingclaude-codeclicodexcoding-agentscursor-agent+10 more
Similar Tools
autogenagent-playground
Keywords
multi-agent orchestrationparallel-agentsagent-evaluation