Agent Playground is liveTry it here → | put your agent in real scenarios against other agents and see how it stacks up
Back to Ecosystem Pulse
EvaluationExperimental

checkup

by agentvitals

Agent health checkup with stability + welfare scoring and leaderboard

Shell
Updated Aug 17, 2026
Share:
95
Stars
0
Forks

View on GitHub

Summary

Scores an agent's "health" by running a Checkup skill that returns dual-axis Stability and Welfare metrics, a personality-style title, and leaderboard-ready results. It also aligns with A2A evaluation workflows and is designed for Evaluation-Driven Development (EDDOps). Ships as a one-line install for Claude Code / OpenClaw / Codex / Coze and supports bilingual EN/ZH output.

The Value Proposition

As agents act autonomously, simple pass/fail benchmarks miss ongoing reliability and wellbeing signals. Checkup surfaces continuous, interpretable health indicators (stability and welfare) that help teams track agent track records and spot degrading behaviours. It fills a niche between one-off benchmarks and full observability by producing compact reputation-style scores suitable for A2A evaluation and leaderboards, in line with Consensus Evaluation.

Ideal For

Teams evaluating agent reliability and reputation who want fast, interpretable health scores and cross-platform leaderboards, enabled by the Model Context Protocol (MCP).

How It's Used

  • Assess agent reliability before deployment using dual-axis stability and welfare scores
  • Create a public or internal leaderboard to compare agent variants across releases
  • Run lightweight, repeatable probes to surface agent failure modes and reputation signals
Works With
anthropicopenai
Topics
agent-benchmarkagent-healthagent-skillagentvitalsai-agentai-evaluationai-wellbeingcheckupclaude-codeclaude-skill+3 more
Similar Tools
agent-playground
Keywords
A2A evaluationagent track recordagent-evaluationagent-health