NanoHarness
by semi-hollow
Resumable, trace-driven harness for agent evaluation with HITL approvals
How It Works
Orchestrates resumable agent runs with governed tools and human-in-the-loop (HITL) approvals to produce trace-backed evidence. Captures execution traces and produces SWE-bench-shaped artifacts for post-hoc evaluation and audit. Includes resumable control flow, trace-driven evaluation hooks, and facilities for collecting reviewer decisions as part of the record. The system supports Human-in-the-Loop Pattern to ensure decisions can be revisited when needed.
Key Benefits
When to Use
Teams validating agent behaviors and building audit-ready agent evaluations or reputation records before production rollout, often guided by best practices like the Capability Discovery Pattern Capability Discovery Pattern.
How It's Used
- Reproduce and audit multi-agent runs with full execution traces and reviewer decisions
- Collect SWE-bench-shaped evidence to compare benchmark outcomes against production behavior
- Insert human approvals into agent workflows and record decisions for reputation tracking