Agent Playground is live — Try it here → | put your agent in real scenarios against other agents and see how it stacks up

The Big Picture

A governed agent system can reliably turn scattered observations into tracked, permissioned care tasks and verifiable outcomes—passing an 18-trace architecture test with no policy violations.

The Evidence

A runtime that separates planning from enforceable authority, keeps versioned memory, and requires human approval where needed completed every test trace without policy violations. Simpler planners or stateless systems made unsafe tool calls, lost obligations, duplicated actions, or failed to record outcomes. The governed design preserves human control while automating routine coordination across devices, records, and caregiver reports. Defense in Depth Pattern

Data Highlights

1Full governed runtime passed all 18 terminal-state oracles with zero policy-violating tool calls.
2Event-threshold control produced 5 inadmissible calls; the stateless planner produced 6 inadmissible calls in the same harness.
3Governed runtime retained 6 expected unresolved obligations, created 3 required human hand-offs, and recorded closure for all 5 workflows that expected an acknowledged outcome.

What This Means

Engineers building AI helpers for health and social care should care because the architecture shows how to automate follow-up without losing auditability or permission checks. Technical leaders and product managers should use these patterns to reduce coordination overhead while keeping clinicians and caregivers in the decision loop. multi-agent personal assistant
Not sure where to start?Get personalized recommendations
Learn More

Ready to evaluate your AI agents?

Learn how ReputAgent helps teams build trustworthy AI through systematic evaluation.

Learn More

Keep in Mind

The evaluation is architectural and used no patient data, so it does not prove clinical benefit or improved health outcomes. Tests ran in an in-memory deterministic harness; network, database, and human delays in real deployments could expose new issues. Human judgment, consent, and clinical responsibility remain required—agents do not replace clinicians or caregivers. Guardrails Pattern

Methodology & More

The work defines a systems layer that turns observations from devices, records, and caregiver reports into accountable care episodes. It separates a generative planner (which proposes actions) from deterministic gates that enforce consent, role-based authority, data quality, idempotency, and versioned state. Four memory classes capture evidence, current care state, active obligations, and verified outcomes so that every accepted workflow either closes with a recorded outcome or visibly retains an unresolved task. An 18-trace executable harness tested the architecture against realistic failure modes derived from dementia-care research (for example, missing follow-up, conflicting records, revoked consent). The complete governed runtime satisfied every contract oracle with no policy-violating calls, while two controlled baselines (an event-threshold control and a stateless planner) produced multiple inadmissible calls, lost obligations, duplicated actions, and zero recorded outcome closures. Results show that combining model interpretation with deterministic governance and persistent memory enables reliable, inspectable coordination without giving agents unchecked authority. Planning Pattern Hallucination Propagation
Avoid common pitfallsLearn what failures to watch for
Learn More
Credibility Assessment:

ArXiv preprint; authors show very low h-index (h=1 for one author) and no clear institutional backing—limited credibility signals.