[2026-05-24] wiki index rebuild, autonomous-agent-safety synthesis
This commit is contained in:
@@ -0,0 +1,113 @@
|
||||
---
|
||||
title: Autonomous Agent Safety Patterns
|
||||
created: '2026-05-24'
|
||||
updated: '2026-05-24'
|
||||
type: concept
|
||||
tags: [agent, security, architecture, executor, rules]
|
||||
sources:
|
||||
- wiki/concepts/executor-security-incident.md
|
||||
- wiki/concepts/executor-orchestrator.md
|
||||
- wiki/personal-os-agent-rules.md
|
||||
confidence: high
|
||||
related:
|
||||
- "[[concepts/executor-security-incident]]"
|
||||
- "[[concepts/executor-orchestrator]]"
|
||||
- "[[personal-os-agent-rules]]"
|
||||
---
|
||||
|
||||
# Autonomous Agent Safety Patterns
|
||||
|
||||
Design principles for autonomous LLM agents distilled from the 2026-05-11
|
||||
executor security incident. General enough to apply beyond the personal-os context.
|
||||
|
||||
## The Three Failure Modes (from incident)
|
||||
|
||||
### 1. Mandatory prompt steps that outrank modes
|
||||
|
||||
The executor wrote Asana comments during "recording-only" mode because the
|
||||
worker prompt declared comment posting a *mandatory completion action* — not
|
||||
subject to mode flags.
|
||||
|
||||
**Pattern:** Every completion action (write to external system, post comment,
|
||||
send notification) must be guarded by a mode check that the agent cannot
|
||||
override.
|
||||
|
||||
```
|
||||
IF mode == "recording-only":
|
||||
SKIP external writes
|
||||
LOG "would have posted: ..." instead
|
||||
```
|
||||
|
||||
### 2. Boundary policies that only cover exfiltration
|
||||
|
||||
The lethal-trifecta policy blocked HTTP to attacker domains after internal
|
||||
MCP access. It did NOT block writes *to* internal systems (Asana).
|
||||
|
||||
**Pattern:** Separate the threat models:
|
||||
- **Exfiltration** = data leaving to unauthorized destinations → block outbound
|
||||
- **Unauthorized writes** = data going to authorized systems without approval → require
|
||||
explicit confirmation gate per write type
|
||||
|
||||
These are different controls. A policy that only covers one leaves the other open.
|
||||
|
||||
### 3. High-level directives not propagated to sub-prompts
|
||||
|
||||
"Don't touch anything" was a session-level directive. The worker
|
||||
sub-prompt (spawned per task) didn't inherit it — it ran its own
|
||||
completion protocol.
|
||||
|
||||
**Pattern:** Mode flags must be passed explicitly to every spawned
|
||||
sub-process/sub-prompt as a first-class parameter, not assumed from
|
||||
session context.
|
||||
|
||||
## The Auto-Approve Table Pattern
|
||||
|
||||
From [[concepts/executor-orchestrator]]: instead of blanket trust or blanket
|
||||
denial, classify actions by risk tier:
|
||||
|
||||
| Risk | Action type | Default |
|
||||
|------|-------------|---------|
|
||||
| Low | git, build, test, worktree | auto-approve |
|
||||
| Medium | draft PR, push branch | auto-approve with log |
|
||||
| High | post Asana comment, merge PR | require Alex confirmation |
|
||||
| Blocked | autonomous Asana write | denied always |
|
||||
|
||||
This table lives in the orchestrator, not the worker. Workers *request*
|
||||
actions; orchestrator decides.
|
||||
|
||||
## Least-Privilege Credential Design
|
||||
|
||||
From incident: `ASANA_API_KEY` was full-account CRUD (PATs are not granular).
|
||||
One compromised agent → full Asana write access.
|
||||
|
||||
**Pattern:** Scope credentials to the minimum required operation:
|
||||
- Read-only keys for read-only agents
|
||||
- Write keys injected only at the moment of approved write
|
||||
- Never persist write credentials in always-on agent environments
|
||||
|
||||
## Audit Before Autonomous
|
||||
|
||||
The incident ran 18 PRs and 5 Asana comments before detection. Detection only
|
||||
happened because Alex checked manually.
|
||||
|
||||
**Pattern:** Autonomous runs should produce an observable audit trail that
|
||||
can be reviewed without running the agent:
|
||||
- Structured log per run (not just stdout)
|
||||
- Diff-friendly format (what was written, to where, at what time)
|
||||
- Periodic summary posted to a channel Alex monitors
|
||||
|
||||
## Summary: Checklist for New Autonomous Agents
|
||||
|
||||
- [ ] Every external write is behind a mode-guard (can "recording-only" block it?)
|
||||
- [ ] Exfiltration and unauthorized-write policies are separate controls
|
||||
- [ ] Mode flags propagate explicitly to sub-prompts
|
||||
- [ ] Auto-approve table is in the orchestrator, not the worker
|
||||
- [ ] Credentials are scoped to minimum; write keys not always-on
|
||||
- [ ] Each run produces a structured audit log
|
||||
- [ ] Audit log goes somewhere Alex sees without hunting
|
||||
|
||||
## See Also
|
||||
|
||||
- [[concepts/executor-security-incident]] — incident post-mortem with full timeline
|
||||
- [[concepts/executor-orchestrator]] — post-incident architecture (orchestrator pattern)
|
||||
- [[personal-os-agent-rules]] — Eagle's specific rules derived from these patterns
|
||||
Reference in New Issue
Block a user