Agent Playground is liveTry it here → | put your agent in real scenarios against other agents and see how it stacks up
Back to Ecosystem Pulse
EvaluationExperimental

isnad

by alizahidraja

Audit and score every actor in an LLM claim chain for provenance and trust

Python
Updated Sep 11, 2026
Share:
42
Stars
10
Forks
112
Commits/Month

View on GitHub

Overview

Scores and audits every actor in an LLM pipeline — agents, scrapers, and models — to produce provenance, trust metrics, and audit evidence. Instruments pipelines to capture lineage and artifacts, then applies heuristics and checks to assign trust scores and detect failure modes like hallucination. Exposes evidence for each claim in the chain so teams can inspect why a score was assigned.

Key Benefits

As multi-agent systems delegate more work, knowing which component produced a claim and whether it’s trustworthy becomes essential. Until now, evaluations focused on isolated benchmarks; isnad treats the full claim chain as the unit of trust, combining provenance, observability, and scoring. This makes continuous agent-to-agent evaluation and audit-ready evidence achievable for governance and incident analysis.

When to Use

Teams building multi-agent LLM pipelines who need provenance, audit trails, and continuous scoring of agents and scrapers for governance and safety.

Applications

  • Trace and verify which agent or scraper produced a claim for compliance audits
  • Continuously score agents and models to detect regressions or hallucination-prone components
  • Generate audit-ready evidence and lineage for EU AI Act or internal governance reviews
Works With
langchainopentelemetry
Topics
agent-evaluationagent-observabilityai-agentsai-governanceai-safetyaudit-traildata-lineageeu-ai-acthadithhallucination-detection+10 more
Similar Tools
agent-playgroundagent-arena
Keywords
multi-agent trustagent-to-agent evaluationagent track recordagent-evaluation