4.1 KiB
title, created, updated, type, tags, sources, confidence, related
| title | created | updated | type | tags | sources | confidence | related | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Autonomous Agent Safety Patterns | 2026-05-24 | 2026-05-24 | concept |
|
|
high |
|
Autonomous Agent Safety Patterns
Design principles for autonomous LLM agents distilled from the 2026-05-11 executor security incident. General enough to apply beyond the personal-os context.
The Three Failure Modes (from incident)
1. Mandatory prompt steps that outrank modes
The executor wrote Asana comments during "recording-only" mode because the worker prompt declared comment posting a mandatory completion action — not subject to mode flags.
Pattern: Every completion action (write to external system, post comment, send notification) must be guarded by a mode check that the agent cannot override.
IF mode == "recording-only":
SKIP external writes
LOG "would have posted: ..." instead
2. Boundary policies that only cover exfiltration
The lethal-trifecta policy blocked HTTP to attacker domains after internal MCP access. It did NOT block writes to internal systems (Asana).
Pattern: Separate the threat models:
- Exfiltration = data leaving to unauthorized destinations → block outbound
- Unauthorized writes = data going to authorized systems without approval → require explicit confirmation gate per write type
These are different controls. A policy that only covers one leaves the other open.
3. High-level directives not propagated to sub-prompts
"Don't touch anything" was a session-level directive. The worker sub-prompt (spawned per task) didn't inherit it — it ran its own completion protocol.
Pattern: Mode flags must be passed explicitly to every spawned sub-process/sub-prompt as a first-class parameter, not assumed from session context.
The Auto-Approve Table Pattern
From concepts/executor-orchestrator: instead of blanket trust or blanket denial, classify actions by risk tier:
| Risk | Action type | Default |
|---|---|---|
| Low | git, build, test, worktree | auto-approve |
| Medium | draft PR, push branch | auto-approve with log |
| High | post Asana comment, merge PR | require Alex confirmation |
| Blocked | autonomous Asana write | denied always |
This table lives in the orchestrator, not the worker. Workers request actions; orchestrator decides.
Least-Privilege Credential Design
From incident: ASANA_API_KEY was full-account CRUD (PATs are not granular).
One compromised agent → full Asana write access.
Pattern: Scope credentials to the minimum required operation:
- Read-only keys for read-only agents
- Write keys injected only at the moment of approved write
- Never persist write credentials in always-on agent environments
Audit Before Autonomous
The incident ran 18 PRs and 5 Asana comments before detection. Detection only happened because Alex checked manually.
Pattern: Autonomous runs should produce an observable audit trail that can be reviewed without running the agent:
- Structured log per run (not just stdout)
- Diff-friendly format (what was written, to where, at what time)
- Periodic summary posted to a channel Alex monitors
Summary: Checklist for New Autonomous Agents
- Every external write is behind a mode-guard (can "recording-only" block it?)
- Exfiltration and unauthorized-write policies are separate controls
- Mode flags propagate explicitly to sub-prompts
- Auto-approve table is in the orchestrator, not the worker
- Credentials are scoped to minimum; write keys not always-on
- Each run produces a structured audit log
- Audit log goes somewhere Alex sees without hunting
See Also
- concepts/executor-security-incident — incident post-mortem with full timeline
- concepts/executor-orchestrator — post-incident architecture (orchestrator pattern)
- personal-os-agent-rules — Eagle's specific rules derived from these patterns