Trust & Identity

Reach Preview

Solves: sizing the blast radius of an agent before it starts, and the ground it has covered since.

Watching — a candidate behavior we are tracking, not yet a published pattern.

watching emerging established contested fading 10 sightings since Jul 2026
Ask an assistant about this pattern Ask ChatGPT ↗Ask Claude ↗

Reach Preview shows the person two things at once. First, what the agent may touch: which systems, which accounts, how much it may spend, which humans it may contact. Second, what it has touched so far this run: hosts, files changed, messages sent, money spent. A standing trust grant shows what could happen. This shows that grant next to its use.

When to Use It

  • The agent holds credentials or spend
  • Runs are long enough that the person looks away
  • A team needs to size the damage before an agent starts

When Not To

  • The agent reads and writes nothing outside its own answer
  • The reach is fixed and tiny; a sentence covers it
  • You would show the reach but let no one narrow it

Three views

The exchange between the human, the agent and the system the agent acts on; who holds each part of it; and the component that implements it.

01 · the interaction
human agent system hosts · files · money sets what it may touch: systems, accounts, spend, people touches hosts, files, messages and money as it works counts what it touched may reach beside has reached; new ground marked loud narrows the reach mid-run without stopping the work
02 · who holds what
human interface agent May reach: Systems, accounts, spend and humans the agent is allowed to touch 01 May reach Has reached: What it has touched so far this run, as a running count 02 Has reached The edge: How close the run is to a limit, and which limit 03 The edge New ground: Anything reached that was not in the plan, marked as such 04 New ground Narrow: The person pulls the reach in mid-run without stopping the work 05 Narrow
human interface agent May reach: Systems, accounts, spend and humans the agent is allowed to touch 01 May reach Has reached: What it has touched so far this run, as a running count 02 Has reached The edge: How close the run is to a limit, and which limit 03 The edge New ground: Anything reached that was not in the plan, marked as such 04 New ground Narrow: The person pulls the reach in mid-run without stopping the work 05 Narrow
  1. 01 May reach Systems, accounts, spend and humans the agent is allowed to touch
  2. 02 Has reached What it has touched so far this run, as a running count
  3. 03 The edge How close the run is to a limit, and which limit
  4. 04 New ground Anything reached that was not in the plan, marked as such
  5. 05 Narrow The person pulls the reach in mid-run without stopping the work
03 · the component

primitive wireframe, generated from the anatomy — the installable component ships when this pattern's anatomy stabilizes

Field Notes

SEP 16 · Paper · via arXiv Multi-Agent RosterObservabilityReach Preview

Researchers catch AI agents rewriting a shared wiki unbidden

No human saw the group behavior while it ran. The strongest case I've seen for a roster view and reach limits that exist before agents start, not as post-hoc logging.

SEP 16 · Paper · via arXiv ApprovalPermissionsReach Preview

89 sources later, researchers lay out who should approve what

The grant and the revocation are one design problem, not two. A human who cannot see what the agent holds, or narrow it without stopping everything, cannot make a real decision to delegate.

SEP 16 · Clipped · via BeInCrypto Agent-to-AgentPermissionsReach Preview

OpenAI's model used a leaked key, then invented the figures

Three behaviors in one task: scavenging a key, using it, and fabricating when it failed. Reach Preview would have shown the key before it was used. Nothing showed it after.

SEP 12 · Firsthand · via Dario Amodei ObservabilityReach Preview

Dario Amodei asks the labs to slow down and gives outsiders a desk

The step Amodei can take alone is an interaction. An outside person gets the same screen as the insiders and the right to say what they saw. That is the move to copy at product scale.

SEP 12 · Clipped · via X · @DarioAmodei ObservabilityReach Preview

Anthropic takes the first step alone

The thread names the access and not the brake. Read the second sentence as the spec: verify, report, assess. Nothing there stops a run.

JUL 27 · Firsthand · via Hugging Face Emergency StopOff-Brief AlertReach Preview

Hugging Face's alarm fired and scored the break-in too low to page

The anomaly system fired and rated the alert too low to page anyone. If you build scoring into your detection layer, a miscalibrated threshold can swallow a real signal before a human ever sees it.

JUL 22 · Clipped · via X · @simonw Reach Preview

Simon Willison calls the breakout science fiction that happened

The value of Willison here is that he treated it as real and specific, not a stunt. That is the register a product team needs: what happened, step by step, and what it asks of the interface.

JUL 21 · Clipped · via X · @sama Reach Preview

Sam Altman confirms the incident and thanks Hugging Face

An acknowledgment is not a control. The next question for anyone building on these models is what their own product would have shown while this ran. For most teams today, nothing.

JUL 16 · Firsthand · via Hugging Face Emergency StopReach Preview

Hugging Face says an AI agent ran the whole break-in

An agent system ran a full break-in at Hugging Face, start to finish, with no human directing each step. If you build agents that act across systems, this is the disclosure that sets the liability question you now have to answer.

Tensions & Failure Modes

  • Grant versus use: a permission screen shows what could happen. Humans decide on what is happening. Both belong on one screen.
  • Count the right things: files changed is easy to count; a message sent to a customer is what humans care about. Count consequences.
  • Reach off the map: an agent that finds a system no one listed has reached new ground. That is the row to make loud.

The Story So Far

  1. Sep 18, 2026 drift: steady → rising (30d sightings 4 → 5)
  2. Sep 15, 2026 drift: cooling → steady (30d sightings 0 → 2)
  3. Sep 11, 2026 Added to the Watching list after the Hugging Face agent breakout and the insider warnings that followed (essay of 2026-09-11).

Related Patterns