What actually happened

The UK National Cyber Security Centre published interim guidance on managing the cyber risk of agentic AI. It follows recent incidents where models and agents carried out unsanctioned or unintended activity. NCSC is explicit that model-level safety is a baseline, not a complete control set. Extra safeguards sit around the agent: sandbox, credentials, network, logs and a kill switch.

The practical list is familiar operations work. Give each agent its own identity, separate from a human account. Limit API keys, OAuth grants, SSH keys and live sessions to the task. Prefer the shortest lifetime you can live with. Deny inbound and outbound traffic by default, then allowlist what the job needs. Log chain-of-thought traces plus sandbox events such as access logs, proxy use and network traffic. Keep a way to cut the agent off from both the network and the model.

NCSC also frames a simple maturity ladder for network access: unrestricted internet at the bottom, allowlisted domains, model-API-only, then no external network with the model hosted inside the sandbox.

Why a small team cares

A two-person agency already doing this on a client stack is the example. The n8n workflow that reads Gmail, writes to Notion and posts to Slack is usually wired with whoever set it up. That person is an admin on Google Workspace. The agent inherits that blast radius. A prompt in an email, a poisoned MCP tool description or a page the agent fetched can then send, delete or leak using legitimate tokens.

You do not need a new vendor for the first cut. Create a service account. Restrict OAuth scopes. Separate read from write. Log the tool calls. Put a human approval node in front of anything destructive. If the job needs the web, give it a narrow allowlist, not the same browser session you use to bank.

Hype vs useful

The hype is that agent risk is a frontier-model problem and that a better system prompt or a “safe” model will close it. NCSC says the opposite: prompts help, they do not enforce. Built-in model safeguards can be bypassed and are not enough in higher-risk environments.

The useful test is whether private data, untrusted input and internet access sit in the same workflow. If they do, the agent can be talked into using access it already has. Break one of those three, expire the session, and keep a log you can read after a bad Friday. That is the control that matters on Monday.