← Field Notes
AUG 7 · Clipped · via X · @bcherny Background ExecutionDelegationPermissions

Claude Code's stacked defenses stop hidden-instruction attacks

If injection risk is structurally solved, approval gates we've shipped 'for safety' may just be friction. The design gravity shifts to delegation, observability, and recovery — not prevention.

Machine summary of the source

A Claude Code engineer reports that layering multiple prompt injection defenses achieves near-zero failure on unseen attacks — a significant security milestone that directly unlocks default autonomous operation. The claim is notable because it comes from a primary source inside Anthropic, and because it frames security not as a blocker but as a prerequisite design condition for full delegation.

For UX designers, this shifts the question. When injection failure rates approach zero, the conservative patterns we've leaned on — manual approval gates, tight permission scopes, frequent interruptions — can be relaxed without proportional risk increase. The safety net becomes structural rather than procedural, which changes what the surface-level experience needs to carry.

The practical implication is that background and autonomous execution modes become safer to ship as defaults rather than opt-ins. Designers who've been holding autonomous flows behind friction 'for safety' may find that friction is no longer load-bearing — and that the real design work shifts to delegation quality, observability, and recovery rather than prevention.

The summary above is generated; the note at the top is the editorial judgment. Primary source ↗