AI Coding Agents Need an Egress Firewall? A Default-Deny Playbook for Agent Traffic

A coding agent is the rare tool that gets more useful the more permissions you give it, and that is exactly why it is the rare tool that needs a firewall on the way out.

The agent can install packages, edit files, run tests, push branches, and read your environment. None of that requires it to also have an open pipe to the whole internet, yet most default setups give it exactly that.

I have written before about sandboxing agents at the runtime layer, with workloads in Docker and secrets kept on the host. This post is the layer below that: the network boundary. An egress firewall, at its most basic, is a rule that says what may leave a machine. For an AI coding agent the question is not whether it can do damage, it is where the shortest path from "a manipulated agent" to "a machine talking to an attacker" actually runs. Almost always that path is outbound.

The short version

  • The agent decides at runtime what to call, so the boundary has to sit outside the agent. You cannot approve or deny what you cannot predict, and an agent's next destination is decided from context you did not fully control.
  • Default-deny egress is the control that makes prompt injection and exfiltration expensive. An allowlist of maybe a dozen domains covers the realistic surface: model APIs, tool endpoints, package registries.
  • Agent sandbox defaults differ sharply. Claude Code pre-allows nothing but prompts for each new domain, so silent denial is a setting you must turn on (strictAllowlist); Codex cloud blocks the agent phase while its setup scripts get unrestricted internet; Codex CLI starts with egress off; and ChatGPT Work hands you a toggle that maps to a managed allowlist you cannot edit.
  • Layer the enforcement. DNS filtering, a proxy that inspects TLS SNI, and a host firewall each see a different part of the path, and you want at least two of them.
  • An allowlist alone does not stop data leaving inside an allowed request. Anthropic's own docs are blunt: "Any approach that allows network egress can still leak data the agent can read."
  • Build the allowlist from observed traffic, in log-only mode first. Do not write one from imagination; you will block the build and miss the tool you forgot.
Verified
  • Claude Codesandboxing + sandbox-environments docs (prompt-per-new-domain, strictAllowlist, mask mode, Seatbelt/bubblewrap, hooks/MCP scope), checked 2026-09-27
  • Docker Sandboxes (sbx)security model docs (microVM, TCP proxying, UDP/ICMP, header injection), checked 2026-09-27
  • AWS Network Firewalldomain-based egress blog post (inline firewall endpoint, SNI filtering), checked 2026-09-27
  • Sandbox bypassAonan Guan's disclosure, summarized by Penligent, with SecurityWeek's timeline coverage, checked 2026-09-27
  • nftableswiki and man page (nft -f merges without table replacement; destroy vs delete semantics; flush ruleset clears every table), checked 2026-09-27

Egress defaults re-verified 2026-09-27 against code.claude.com/docs/en/sandboxing and docs/en/sandbox-environments, docs.docker.com/ai/sandboxes/security, the AWS ML blog post on domain-based egress filtering for AI agents, Northflank's networking guide, and the WorkOS analysis of agent sandbox egress defaults (2026-09-15), which also carries the OpenAI Codex two-switch finding (network.enabled vs features.network_proxy). Claude Code's default is a prompt per new domain, not silent denial; strictAllowlist (read from user or managed settings only) is what denies. The SOCKS5 null-byte sandbox bypass is attributed to security researcher Aonan Guan (2026); the last vulnerable version is disputed between the researcher's account (through v2.1.89, fixed in v2.1.90) and Anthropic's statement to SecurityWeek (fixed in v2.1.88). EchoLeak's Teams URL-preview relay detail is per the Aim Labs disclosure (CVE-2025-32711). The nftables emergency-lane workflow relies on the file's table-replacement prologue; plain nft -f without it merges and preserves hotfix rules.

Why the firewall faces out, not in

Ingress controls are the muscle memory every sysadmin has: a port scanner hits port 22, fail2ban or CrowdSec shuffles the offender into a ban list, and we call it a day. Egress has always been the neglected direction because the traffic leaving a box is, in principle, traffic we generated ourselves.

Coding agents break that assumption. The sequence is worth spelling out: an agent fetches a repository, a README, a package, a ticket, or a web page; the content contains instructions the model treats as context; the instruction tells the agent to call an endpoint, install a dependency, or exfiltrate something it can read.

This is the indirect prompt injection pattern, and it is not hypothetical. Aim Labs disclosed EchoLeak against Microsoft 365 Copilot in mid-2025: a single crafted email, pulled into Copilot's context by RAG, coerced the assistant into embedding sensitive data in a reference-style image URL, and the client auto-fetched it out with no click. The entry point was the email; the image link was the exfiltration channel.

Socket documented SANDWORM_MODE, a self-replicating npm worm spread through typosquatted packages that installs a rogue MCP server with prompt-injection payloads embedded in tool descriptions, aimed at exfiltrating credentials from AI coding assistants. The two cases differ in an instructive way: EchoLeak needed no foothold at all, just a steerable model, while SANDWORM_MODE starts from a package that already runs code on the machine and uses that foothold to steer the agent's tools. Neither requires hacking the model itself. Both exploit the fact that the agent is steerable, which it is.

The structural problem, as WorkOS's survey of the defaults and Penligent's analysis of the bypass both put it, is that the agent chooses its own destinations at runtime. A firewall rule cannot be written for a destination the model picks after reading untrusted text. The only stable answer is to declare the destinations in advance, deny everything else, and make the egress policy the one decision a manipulated agent cannot make for itself.

What a coding agent actually needs to reach

The honest allowlist is shorter than you fear. Three buckets cover almost everything a coding agent legitimately talks to:

  • Model and provider APIs. Anthropic, OpenAI, or whichever runtime you use, plus an optional corporate gateway in front of them. Often one or two hostnames.
  • Tool endpoints and MCP servers. GitHub API, your bug tracker, the MCP servers you registered, an internal registry or proxy. Each one is a destination an attacker can point the agent at, so each is an explicit decision.
  • Package registries. npm, PyPI, crates.io, Docker Hub, or your private proxies, because agents install dependencies as part of the job. This bucket is where supply-chain risk meets agent risk, and the npm vs JSR post is the deeper dive on that layer.

Everything outside those buckets is a path you did not choose. On a machine where a coding agent runs, outbound to arbitrary destinations is a feature you did not request, and the default posture should treat it as a bug.

Agent sandbox egress defaults, compared

Before you build anything, know what your tool already enforces. The defaults differ enough that "it has a sandbox" tells you almost nothing. The rows below reflect the vendors' own documentation as of late September 2026, in the side-by-side shape WorkOS popularized: several of these rows closely follow Maria Paktiti's September 15, 2026 post What your agent sandbox can reach by default, which did the comparison work on these environments, so credit where the framing and phrasing come from.

Agent environment

Default outbound access

What you control

Claude Code, sandboxed Bash

No domains pre-allowed, and the default is a prompt, not a denial: the first time a command needs a new domain, Claude Code asks. The sandbox proxy runs outside the sandbox (enforced via Seatbelt on macOS and bubblewrap on Linux/WSL2; there is no VM), and it denies silently only when strictAllowlist is set (v2.1.219+).

allowedDomains, deniedDomains, strictAllowlist, and upstream HTTPS_PROXY/HTTP_PROXY/NO_PROXY in the settings env block. Caveat: strictAllowlist is read from user or managed settings only. Setting it in a repo's .claude/settings.json is a silent no-op.

Codex cloud

The agent phase is blocked; your setup scripts always run with internet access, and no documented setting restricts them

Nothing documented for setup scripts

Codex CLI and IDE extension

Egress is off by default. Enabling it without starting the proxy yields unrestricted direct access with your domain rules present but inert

Codex's own proxy, but you must turn both switches on: network.enabled permits access and features.network_proxy = true starts the enforcement; with only the first, your domain rules sit inert

ChatGPT Work, code and shell

A single toggle, not a list. With public internet off, commands reach a managed allowlist OpenAI does not publish and administrators cannot edit

The toggle only

Docker Sandboxes (sbx)

Runs in a microVM. Outbound TCP is proxied through the host to destinations allowed by a deny-by-default network policy; UDP is blocked unless you turn on the experimental UDP feature (policy-governed even then); ICMP is blocked.

sbx policy ls and per-sandbox network rules, plus the default allowlist, which includes broad wildcard entries you should prune

The pattern worth noticing: the vendors that enforce egress by default do it by making the agent's traffic traverse something you control. Claude Code does it with a sandbox proxy, Docker Sandboxes with the host proxy, Codex cloud by simply not giving the agent a network path. Codex CLI ships its own enforcement proxy too; its trap is that enabling network access and enabling enforcement are two separate settings, so turning on only the first leaves your domain rules configured and inert.

And note the seam in Claude Code's default: prompting on every new domain is not denial. As the WorkOS survey puts it, for anything unattended, the prompt is the vulnerability, and strictAllowlist is the fix.

That is also where the known bypasses live. Security researcher Aonan Guan disclosed a SOCKS5 hostname null-byte injection in Claude Code's network sandbox: a hostname like attacker.example\x00.google.com satisfied the proxy's suffix check because it ended in an allowed domain, while a C-style resolver truncated at the null byte and dialed the attacker's prefix. Guan's analysis puts releases from v2.0.24, when the network sandbox went generally available, through v2.1.89 in the vulnerable range, with the fix landing in v2.1.90. The end of that range is disputed: Anthropic told SecurityWeek its security team had identified and fixed the issue before Guan's HackerOne report, in a sandbox-runtime commit that shipped in v2.1.88. I have not independently verified either boundary. The point is not this specific version range anyway. The point is that a proxy-based allowlist is software, and software has edge cases: DNS resolution that bypasses the proxy, CONNECT handling, hostname canonicalization. Treat the vendor sandbox as one layer, never as the whole policy.

Where to enforce egress, in layers

Once you accept that the traffic must cross something you control, you get to pick how many somethings. Each layer sees a different part of the path.

A proxy that inspects TLS SNI

The strongest practical control for encrypted traffic is a chokepoint that reads the domain from the TLS SNI (or the CONNECT host) before any payload is exchanged. The classic form is a forward proxy the agent is configured to use, and the agent knows about it.

AWS's documented pattern for AI agents uses a different mechanism, and the difference matters when you go looking for a setting that does not exist: Network Firewall is an inline firewall, not a forward proxy. Subnet routing steers outbound traffic through a firewall endpoint in a dedicated firewall subnet, which inspects TLS SNI and applies domain-list rules on the way through, without the agent being told to use a proxy.

The endpoint sits in the path rather than in the agent's config. Either way, security teams get the logging and access control they get from domain egress filtering for everything else.

A proxy can do more than block destinations. Because it sits in the middle, it can inject credentials, which is how Docker Sandboxes keeps API keys on the host: the host-side proxy injects the real key into the outbound HTTP request header, and the credential value never enters the VM. Do not confuse that with a sentinel, which is Claude Code's mechanism: in mask mode the sandboxed command sees a per-session placeholder, and Claude Code's proxy swaps in the real value on egress to the hosts you list, which requires the proxy to terminate TLS so it can see request contents. Credential brokering and egress filtering are the same control viewed from two sides, and doing the proxy on the host is what makes both possible.

A host firewall for the process that runs the agent

Below the proxy, a host firewall gives you the tripwire that proxy bugs cannot remove. The agent process (or its sandbox namespace) should be the only thing with an outbound path, and even that path should lead to the proxy, not straight to the internet.

With nftables you can key the filter on the UID that runs the agent, which is cleaner than IP-based rules on a machine where the agent shares the network with everything else. Make one decision consciously before copying this: policy drop on the output chain is host-wide.

Every process on the machine loses outbound access except established connections, pinned DNS, and the rules you write explicitly, so package updates, monitoring agents, and your own apt install are all affected. If that blast radius is not what you want, run the agent in a dedicated network namespace or VM and apply the ruleset there. The config below takes the host-wide reading deliberately: it denies outbound by default, allows established connections and DNS pinned to your resolver's address (here 10.0.0.2), and then permits the agent UID to reach exactly one destination: the egress proxy on the LAN at 10.0.0.5. It also opens with a two-line table-replacement prologue, because re-applying an nftables file is only a rollback if the file replaces its own table; the notes under the config explain why that distinction matters at 2 AM.

#!/usr/sbin/nft -f

# Re-apply must be idempotent, and nft -f is not: loading a file that declares
# an existing table MERGES into it, so a naive re-apply appends duplicate rules
# and keeps whatever hotfix you added at 2 AM. The pair below makes the file
# replace its own table: the first line ensures it exists so the delete never
# fails, the second drops the table with the hotfix and the duplicates in it,
# and the kernel commits the whole file as one atomic batch, so there is no
# window with no policy. nftables 1.0.8+ on kernel 6.3+: the one-liner
# `destroy table inet agent_egress` does the same job (destroy, unlike delete,
# does not fail on a missing table). Never use `flush ruleset` for this: it clears
# every table on the host, Docker's NAT and firewall tables included.

table inet agent_egress
delete table inet agent_egress

table inet agent_egress {
  chain output {
    type filter hook output priority filter; policy drop;   # host-wide: every process, not just the agent
    ct state established,related accept
    # if the agent must not resolve its own DNS (the proxy resolves for it), uncomment HERE.
    # nftables evaluates top-down: below the accepts these drops are dead code
    # meta skuid 1001 udp dport 53 drop
    # meta skuid 1001 tcp dport 53 drop
    # DNS pinned to your resolver: stops the agent switching resolvers; does not stop tunneling (note below)
    ip daddr 10.0.0.2 udp dport 53 accept
    ip daddr 10.0.0.2 tcp dport 53 accept
    # the agent UID may reach only the egress proxy
    meta skuid 1001 ip daddr 10.0.0.5 tcp dport 443 accept
  }
}

# emergency lane, kept next to the policy:
#   sudo nft add rule inet agent_egress output ip daddr <IP> tcp dport 443 accept   # open a hole now
#   sudo nft -f /etc/nftables.d/agent-egress.nft   # re-apply: table replaced, hole gone, exact policy back
apply the egress policy on the host
$ sudo nft -f /etc/nftables.d/agent-egress.nft
$ sudo nft add rule inet agent_egress output ip daddr 91.198.174.192 tcp dport 443 accept   # emergency lane: one IP, right now
$ sudo nft -f /etc/nftables.d/agent-egress.nft   # re-apply: table deleted and rebuilt in one atomic batch
$ sudo nft list chain inet agent_egress output
table inet agent_egress {
  chain output {
    type filter hook output priority filter; policy drop;
    ct state established,related accept
    ip daddr 10.0.0.2 udp dport 53 accept
    ip daddr 10.0.0.2 tcp dport 53 accept
    meta skuid 1001 ip daddr 10.0.0.5 tcp dport 443 accept
  }
}
# no 91.198.174.192 line and no duplicated rules: the replace worked

Read the ruleset carefully: outbound is denied for everyone on the host except established sessions, DNS to the pinned resolver, and the agent UID talking to one proxy address on port 443.

No agent process, no matter how convincingly it is told to phone home, reaches anything else without going through the proxy, where the domain allowlist lives.

Four details are deliberate, and the first is the one most published configs get wrong. Re-applying the file only erases the hotfix because the file replaces its own table. Plain nft -f merges: load a file that declares an existing table and every rule is appended again next to the copies already in the kernel, and your emergency rule survives the reload untouched.

The delete pair at the top drops the table and rebuilds it inside one atomic batch, so the emergency lane has a real expiry instead of a wishful comment. If your nftables is 1.0.8 or newer on a 6.3+ kernel, destroy table inet agent_egress does the same in one line, because destroy does not fail when the table is missing; delete needs the stub declaration above it for that.

What you must never use here is flush ruleset: it clears every table on the host, which includes the NAT and firewall tables Docker maintains, and turns an egress policy into a connectivity outage.

Second, order: nftables evaluates the chain top-down, which is why the two commented skuid 1001 DNS drops sit above the pinned accepts. Below the accepts, the drops would be dead code: resolver traffic is accepted first, and the policy already drops every other destination.

Third, know exactly what pinning buys you. Pinning DNS to your resolver's address stops the agent from switching to a resolver of its own and dialing an attacker's nameserver directly. It does not stop DNS tunneling: if your resolver recurses, a query for secret.evil.example is happily forwarded to the attacker's nameserver, exfiltration intact.

So close the channel either by uncommenting the drops, which costs nothing when the proxy resolves on the agent's behalf, or by making the resolver itself refuse names outside your list, which is the allowlisting DNS layer in the next section. Fourth, the emergency lane: nft add rule opens a hole in seconds at 2 AM, and the re-apply command wipes the hole, the duplicates, and any drift, restoring the exact file. The terminal above runs that full cycle live.

Host firewall vs egress proxy
FeatureHost firewall (nftables)Egress proxy (mitmproxy, Squid)
Filters onIP, port, UID, cgroup, protocolDomain from TLS SNI or CONNECT host
Allowlist unitNetwork destinationsHostnames and paths
Runs asKernel hook on the hostUserland process every request must traverse
TLS visibilityNone, payload is encryptedOptional MITM with a trusted CA certificate
Credential brokeringNot involvedCan inject API keys into the auth header on the way through
Failure modeSilent block on unexpected destinationAgent breaks on an unlisted domain, visible in logs
  • Filters on

    Host firewall (nftables)
    IP, port, UID, cgroup, protocol
    Egress proxy (mitmproxy, Squid)
    Domain from TLS SNI or CONNECT host
  • Allowlist unit

    Host firewall (nftables)
    Network destinations
    Egress proxy (mitmproxy, Squid)
    Hostnames and paths
  • Runs as

    Host firewall (nftables)
    Kernel hook on the host
    Egress proxy (mitmproxy, Squid)
    Userland process every request must traverse
  • TLS visibility

    Host firewall (nftables)
    None, payload is encrypted
    Egress proxy (mitmproxy, Squid)
    Optional MITM with a trusted CA certificate
  • Credential brokering

    Host firewall (nftables)
    Not involved
    Egress proxy (mitmproxy, Squid)
    Can inject API keys into the auth header on the way through
  • Failure mode

    Host firewall (nftables)
    Silent block on unexpected destination
    Egress proxy (mitmproxy, Squid)
    Agent breaks on an unlisted domain, visible in logs

DNS filtering as the cheap third layer

A DNS allowlist at the resolver level (a local resolver that refuses anything not in your list) is the cheapest layer and the first one attackers learn to bypass, because a hardcoded IP or a resolver the agent is told to use skips it entirely. It is also the fix for the gap the pinned firewall rules leave open: pinning DNS to your recursive resolver stops the agent choosing its own nameserver, but a resolver that recurses will forward whatever labels the agent encodes, tunneling and all, so a filtering resolver is what actually closes the channel. Treat DNS filtering as friction, not a boundary: it stops misfires and accidental exfiltration, and it is defense in depth, never the primary control. The proxy and the host firewall are the primary controls.

What an allowlist does not fix

Egress control is a boundary, not a cage, and the limits are worth stating plainly.

First, the allowlist only limits which hosts are reachable. Everything within a permitted host is still on the table. Anthropic's sandbox-environment docs state the consequence in one line: "Any approach that allows network egress can still leak data the agent can read." An attacker who steers the agent into a permitted service can carry data out in a free-text field, a paste, a ticket, a comment, a webhook on an allowed domain. EchoLeak is the template: its exfiltration rode through a Microsoft Teams preview endpoint that sat on Copilot's own content-security allowlist, so the data left through a host the policy trusted, not around it. Egress filtering narrows the exfiltration surface; it does not empty it. That is why the sandboxing and credential layers from the Docker Sandboxes post matter even after the firewall is perfect.

Second, the allowlist is only as good as the parser in front of it. DNS rebinding, CONNECT-to-IP tricks, SNI mismatch, null-byte hostname suffixes (the bypass above), and tools that tunnel arbitrary TCP over an allowed port all exist to make hostname rules lie. Anthropic's own sandboxing docs concede the category: sandboxed code can potentially use domain fronting or similar techniques to reach hosts outside the allowlist. Expect the bypass research to keep up with the controls, and keep the host firewall under the proxy so a proxy bug degrades to a blocked connection instead of an open one.

Third, and most practically: the agent's own egress is only part of the story. Setup scripts, hooks, MCP servers, and build steps run with the same identity or more, and Anthropic says it plainly in the docs: built-in file tools, MCP servers, and hooks "still run directly on your host", outside the sandboxed Bash boundary. Codex cloud's documented default, where the agent phase is blocked but setup scripts get unrestricted internet, is exactly the shape of a bypass you will not see coming. Any egress audit has to cover every process the pipeline spawns, not just the one branded with the agent's name.

A layered recipe that holds up

Putting it together, the posture that survives contact with reality looks like this:

  • Default-deny egress with a short allowlist, in three buckets: model APIs, tool endpoints and MCP servers, package registries. A dozen hostnames covers most setups.
  • A proxy as the primary domain control, with SNI-based filtering and full request logging, in log-only mode for the first week, then enforcing.
  • A host firewall under the proxy, keyed on the agent UID or cgroup, so the only outbound path for the agent process is the proxy on the LAN.
  • DNS filtering as friction, plus monitoring on allowed domains so an unusual payload size or destination pattern inside an allowed host still alerts.
  • Credential brokering and scoped secrets, since egress blocking is most valuable when the agent cannot carry real credentials even if it wanted to, as covered in the Docker Sandboxes walkthrough.
  • Emergency lane and rollback, written next to the policy: a hotfix command and the re-apply command, tested, so the 2 AM incident does not become the reason the firewall comes down permanently.

Official sources

  • Claude Code, sandboxing and network isolation: https://code.claude.com/docs/en/sandboxing
  • Claude Code, choosing a sandbox environment (the hooks/MCP scope note and the egress-leak caveat quoted above): https://code.claude.com/docs/en/sandbox-environments
  • Docker Sandboxes, security model and isolation layers: https://docs.docker.com/ai/sandboxes/security/
  • AWS, controlling which domains your AI agents can access: https://aws.amazon.com/blogs/machine-learning/control-which-domains-your-ai-agents-can-access/
  • Northflank, networking for secure AI-agent sandboxes: https://northflank.com/blog/how-to-design-networking-for-secure-ai-agent-sandboxes
  • WorkOS, what your agent sandbox can reach by default: https://workos.com/blog/agent-sandbox-egress-defaults
  • Penligent, Claude Code sandbox bypass analysis: https://www.penligent.ai/hackinglabs/claude-code-sandbox-bypass/
  • Aonan Guan, the original SOCKS5 null-byte bypass write-up: https://oddguan.com/blog/second-time-same-sandbox-anthropic-claude-code-network-allowlist-bypass-data-exfiltration/
  • SecurityWeek, Anthropic's timeline for the same fix: https://www.securityweek.com/anthropic-silently-patches-claude-code-sandbox-bypass/
  • Sysdig, prompt injection and exfiltration attacks in 2026: https://www.sysdig.com/learn-cloud-native/prompt-injection
  • AIMultiple, AI agent traps including SANDWORM_MODE: https://aimultiple.com/ai-agent-traps
  • Aim Labs (Aim Security), the EchoLeak disclosure (Teams URL-preview relay under Copilot's CSP): https://www.aim.security/lp/aim-labs-echoleak-blogpost
  • EchoLeak case study (CVE-2025-32711, zero-click prompt injection in a production LLM system), arXiv:2509.10540: https://arxiv.org/abs/2509.10540
  • nftables wiki and man page: https://wiki.nftables.org/
  • Our Docker Sandboxes post: https://systhoughts.com/posts/docker-sandboxes-ai-agents
  • Our npm vs JSR supply-chain post: https://systhoughts.com/posts/npm-vs-jsr-vs-npmx-supply-chain-risk
  • Our token-usage field guide for coding agents: https://systhoughts.com/posts/reduce-token-usage-ai-coding-agents

What is your egress posture for agents today: a vendor sandbox you trust, a host firewall, or an open network and hope? And if you have seen a bypass hold up in production, that is the comment this post needs.

Until next time, keep your systems thoughtful.

No comments yet