Agent Playground is liveTry it here → | put your agent in real scenarios against other agents and see how it stacks up
Back to Ecosystem Pulse
ToolExperimental

polos

by polos-dev

Sandboxed runtime for durable, observable AI agent workflows

TypeScript
Updated Feb 24, 2026
Share:
36
Stars
4
Forks

View on GitHub

What It Does

Provides a sandboxed runtime for running AI agents with durable workflows, automatic retries, and prompt caching. Uses built-in tools, human-in-the-loop approvals, and Slack integration to keep long-running agent tasks observable and controllable. Distinctive features include sandboxed execution environments and durable workflows so agents can be retried or paused without losing context.

Why It Matters

As agents act autonomously, being able to safely run, observe, and intervene in their workflows is essential to building trust. Polos gives teams the infrastructure to contain risky actions, require human approvals, and persist execution state—so agent failures become diagnosable and repeatable. That makes it easier to collect agent track records and other trust signals needed for evaluation and reputation systems. Responsible AI

Best For

Teams building agentic applications that need safe execution, human approvals, and Agent Service Mesh Pattern in development or early production.

How It's Used

  • Run untrusted agent code safely in a sandboxed environment before production rollout
  • Add human approval gates to high-risk agent actions and pause/resume workflows
  • Persist long-running agent workflows with automatic retries and prompt caching for reproducible failure analysis
Works With
slacktypescriptpython
Topics
agent-orchestrationagentic-aiai-agentsai-observabilitydeveloper-toolsdurable-executionhuman-in-the-looppythonsandboxtypescript
Similar Tools
autogencrewai
Keywords
agent reliabilitydurable-executionagent-to-agent evaluationhuman-in-the-loop