every_eval_ever
by evaleval
Shared schema and crowdsourced database for standardized AI evaluation results
Overview
Defines a shared schema and crowdsourced database for AI evaluation results to standardize metadata across papers, leaderboards, and local runs. standardize metadata across papers, leaderboards, and local runs Normalizes how evals are described so results from different frameworks can be compared, reproduced, and merged. Provides a common format and import pathways to capture provenance, metrics, and dataset details.
Key Benefits
Ideal For
Researchers and engineers who need to aggregate, compare, and reproduce evaluation results across frameworks and papers to build reliable agent track records. agent track records This tool helps establish reproducible baselines and enables fair comparisons, supporting teams as they assemble robust evaluation datasets. For interoperability across systems, consider adopting the Model Context Protocol (MCP) to align evaluation metadata with wider agent contexts.
Real-World Examples
- Normalize disparate benchmark outputs into a single, comparable metadata format
- Aggregate leaderboard and paper results to build reproducible agent track records
- Store provenance and metric definitions for audits and continuous agent evaluation