Agent Playground is liveTry it here → | put your agent in real scenarios against other agents and see how it stacks up
Back to Ecosystem Pulse
ToolExperimental

CORAL

by Human-Agent-Society

Evolving and grading multi-agent coding workflows for autoresearch

Python
Updated Aug 23, 2026
Share:
924
Stars
118
Forks

View on GitHub

Summary

Runs multi-agent autoresearch workflows that evolve coding agents and grade their outputs. Uses autonomous agent populations with shared memory, automated grading, and evolutionary selection to iteratively improve code-generation agents. Notable features include support for Claude Code, OpenCode, and Codex backends and built-in scoring/evolution loops for continual improvement.

Key Benefits

As agents are asked to solve complex research and coding tasks, you need repeatable ways to compare and evolve agent strategies rather than one-off prompts. CORAL provides an experimental playground for agent-to-agent evaluation and evolutionary selection, surfacing which agent designs and prompts produce reliable outcomes. That makes it useful for building agent track records and studying multi-agent reliability and failure modes before deploying agents in production.

When to Use

Researchers and teams exploring automated agent evolution and comparative evaluation of coding agents across model backends.

Applications

  • Evolve agent populations to discover better coding strategies and prompts
  • Benchmark and grade code-generation agents across Claude Code, OpenCode, and Codex
  • Run reproducible autoresearch experiments with shared memory and selection loops
Works With
openaianthropichuggingface
Topics
agent-frameworkagent-orchestrationagentic-aiai-agentsalpha-evolveautonomous-agentsautoresearchclaude-codecode-generationcodex+10 more
Similar Tools
autogencrewai
Keywords
multi-agent orchestrationagent-to-agent evaluationcontinuous agent evaluation