Agent Playground is liveTry it here → | put your agent in real scenarios against other agents and see how it stacks up
Back to Ecosystem Pulse
EvaluationExperimental

agent-leaderboard

by rungalileo

Notebook-driven leaderboard for ranking LLMs on agentic tasks

Jupyter Notebook
Updated May 21, 2026
Share:
224
Stars
28
Forks

View on GitHub

Overview

Ranks language models on agentic tasks by running structured benchmarks and aggregating performance metrics. Uses Jupyter notebooks to define tasks, generate synthetic agent interactions, and compute leaderboard scores across multiple evaluation dimensions. Includes example prompts, scoring heuristics, and visualizations to compare agent behaviors, with structured benchmarks guiding evaluation patterns.

Key Benefits

As agents take on more autonomous, multi-step roles, simple LLM benchmarks miss interaction and delegation failure modes. This project provides a lightweight way to surface differences in agentic behavior and track record across tasks, which is a first step toward agent-to-agent evaluation and reputation. Having reproducible notebooks makes it easy to iterate on evaluation patterns and integrate new trust signals into your workflow.

Ideal For

Researchers and engineers prototyping evaluation suites to compare LLMs on multi-step, agent-like behaviors and generate reproducible leaderboards. The project also supports building reproducible leaderboards to track agent performance over time.

Real-World Examples

  • Evaluating LLM performance on multi-step agent tasks with reproducible notebooks
  • Comparing delegation and failure modes across models using synthetic agent interactions
  • Generating leaderboards and visualizations to inform model selection for agent systems
Topics
agent-evaluationaiai-agentsai-benchmarkai-evaluationevaluationllmssynthetic-data
Similar Tools
agent-playgroundagent-arena
Keywords
agent-evaluationA2A evaluationagent track recordmulti-agent trust