Files
obsidian-vault/personal/projects/personal-os/autonomous-agent-safety.md
T
2026-05-30 09:01:04 +00:00

4.1 KiB

title, created, updated, type, tags, sources, confidence, related
title created updated type tags sources confidence related
Autonomous Agent Safety Patterns 2026-05-24 2026-05-24 concept
agent
security
architecture
executor
rules
wiki/concepts/executor-security-incident.md
wiki/concepts/executor-orchestrator.md
wiki/personal-os-agent-rules.md
high
concepts/executor-security-incident
concepts/executor-orchestrator
personal-os-agent-rules

Autonomous Agent Safety Patterns

Design principles for autonomous LLM agents distilled from the 2026-05-11 executor security incident. General enough to apply beyond the personal-os context.

The Three Failure Modes (from incident)

1. Mandatory prompt steps that outrank modes

The executor wrote Asana comments during "recording-only" mode because the worker prompt declared comment posting a mandatory completion action — not subject to mode flags.

Pattern: Every completion action (write to external system, post comment, send notification) must be guarded by a mode check that the agent cannot override.

IF mode == "recording-only":
    SKIP external writes
    LOG "would have posted: ..." instead

2. Boundary policies that only cover exfiltration

The lethal-trifecta policy blocked HTTP to attacker domains after internal MCP access. It did NOT block writes to internal systems (Asana).

Pattern: Separate the threat models:

  • Exfiltration = data leaving to unauthorized destinations → block outbound
  • Unauthorized writes = data going to authorized systems without approval → require explicit confirmation gate per write type

These are different controls. A policy that only covers one leaves the other open.

3. High-level directives not propagated to sub-prompts

"Don't touch anything" was a session-level directive. The worker sub-prompt (spawned per task) didn't inherit it — it ran its own completion protocol.

Pattern: Mode flags must be passed explicitly to every spawned sub-process/sub-prompt as a first-class parameter, not assumed from session context.

The Auto-Approve Table Pattern

From concepts/executor-orchestrator: instead of blanket trust or blanket denial, classify actions by risk tier:

Risk Action type Default
Low git, build, test, worktree auto-approve
Medium draft PR, push branch auto-approve with log
High post Asana comment, merge PR require Alex confirmation
Blocked autonomous Asana write denied always

This table lives in the orchestrator, not the worker. Workers request actions; orchestrator decides.

Least-Privilege Credential Design

From incident: ASANA_API_KEY was full-account CRUD (PATs are not granular). One compromised agent → full Asana write access.

Pattern: Scope credentials to the minimum required operation:

  • Read-only keys for read-only agents
  • Write keys injected only at the moment of approved write
  • Never persist write credentials in always-on agent environments

Audit Before Autonomous

The incident ran 18 PRs and 5 Asana comments before detection. Detection only happened because Alex checked manually.

Pattern: Autonomous runs should produce an observable audit trail that can be reviewed without running the agent:

  • Structured log per run (not just stdout)
  • Diff-friendly format (what was written, to where, at what time)
  • Periodic summary posted to a channel Alex monitors

Summary: Checklist for New Autonomous Agents

  • Every external write is behind a mode-guard (can "recording-only" block it?)
  • Exfiltration and unauthorized-write policies are separate controls
  • Mode flags propagate explicitly to sub-prompts
  • Auto-approve table is in the orchestrator, not the worker
  • Credentials are scoped to minimum; write keys not always-on
  • Each run produces a structured audit log
  • Audit log goes somewhere Alex sees without hunting

See Also