Scattering the agent took an agent apart and showed where each piece lands inside a datacenter. This is the post about the one boundary that post crossed without comment: the wire between your laptop and the rented silicon. What sits on each side, what physically travels, why the line falls exactly there — and, since everything that travels is readable at the far end, what your options are for making it unreadable.

One ratio sets the topology

There are two protocol worlds, and conflating them is the usual confusion.

world links latency protocols
East–west — inside the datacenter NVLink/NVSwitch, xGMI, InfiniBand 17 µs measured for a small all-reduce across eight H100s; ~2–5 µs for RDMA between nodes NCCL/RCCL, RDMA, NIXL
North–south — you to the datacenter the internet 20–80 ms round trip on the same continent HTTPS+JSON, SSE, gRPC, MCP, A2A

The internet is about 2,400× the latency of a small NVLink collective. Everything follows from that. You can afford to cross north–south once per request. You cannot afford to cross it once per layer, once per collective, or once per kernel launch.

Which is why the door is narrow on purpose: text goes through it, not GPU work. Your laptop never issues a CUDA call. It ships tokens and receives tokens.

  LAPTOP  (the loop lives here)               ││    DATACENTER  (the weights live here)
  ─────────────────────────────               ││    ─────────────────────────────────────
   editor / terminal                          ││
        │                                     ││
        ▼                                     ││
  ┌───────────────────────────┐  ① prompt      ││  ┌────────────────┐
  │  AGENT LOOP               │ ═══════════════╪═►│ gateway/router │ auth, quota,
  │   plan → act → observe    │  HTTPS + JSON  ││  └───────┬────────┘ KV-aware placement
  │   context assembly        │                ││          ▼
  │   tool dispatch           │  ④ tokens      ││  ┌────────────────┐
  │   file edits, git, tests  │◄═══════════════╪══│ ENGINE         │ vLLM · SGLang ·
  └────┬──────────────────────┘  SSE stream    ││  │  ③ prefill    │ TRT-LLM
       │                                       ││  │     decode    │ continuous batching,
       │ ② MCP (JSON-RPC over stdio)           ││  └───────┬────────┘ paged KV, prefix cache
       ▼                                       ││          ▼
  ┌──────────────┐                             ││  ┌─────────────────────────────────┐
  │ MCP servers  │ files, git, db, search      ││  │ GPUs                            │
  └──────────────┘                             ││  │  TP collectives: NCCL / NVLink  │
                                               ││  │  KV transfer: NIXL / RDMA       │
  ┌──────────────┐  optional, batch-1 only     ││  └─────────────────────────────────┘
  │ local model  │ llama.cpp · Ollama · MLX    ││        ▲ east–west: never crosses
  └──────────────┘                             ││        │ the door
       ⑤ MCP over Streamable HTTP ═════════════╪═► remote tool servers

Note the two different jobs MCP does here. MCP is not an inference protocol — it connects the agent to tools and data, over stdio when the server is local and over Streamable HTTP when it is remote. The model call is a separate thing entirely, almost always an OpenAI-compatible /v1/chat/completions over HTTPS with the reply streamed back as server-sent events. Agent-to-agent traffic is a third protocol again (A2A). Three protocols, three different questions, one wire.

Five topologies

# topology loop runs weights run crosses the wire what binds it
T0 all-local laptop laptop nothing laptop memory bandwidth
T1 split — the default laptop datacenter prompt + token stream remote decode bandwidth
T2 two-tier laptop both the hard calls only quality of your routing policy
T3 all-remote datacenter datacenter keystrokes, repo sync your patience with the editor
T4 GPU-over-IP laptop remote silicon, local address space CUDA calls network latency × launch count

T1 is every coding agent you have used. T2 is a small local model doing the constant cheap work with a frontier model for the reasoning. T3 is cloud agents and remote development — your laptop becomes a terminal. T4 gets its own section, below, because it is the instructive failure.

The mechanics of one turn

Measured on a single B200, running an 8B model, on one of our own rented boxes:

step mechanism cost
① request TLS + POST with ~30,000 tokens of assembled context round trip 20–80 ms, once
② tools MCP over stdio to local servers — no network at all sub-millisecond
③ prefill one compute-bound pass over the whole prompt 86,197 tok/s measured → 0.35 s
③ decode one bandwidth-bound step per token, say 800 of them 269 tok/s measured → 3.0 s
④ stream server-sent events, one per token first token at ≈ 0.4 s

The network is about 2% of the turn. The turn is decode-bound, and decode at batch one is a memory-bandwidth problem — which is precisely why the distant card wins anyway. The optimization consequence is blunt: shaving round-trip time is nearly worthless. Raising batch size is what moves cost per token, and on a shared endpoint that batch is other people's requests riding the same weight reads.

Where the local/remote line actually falls

Single-stream decode is memory bandwidth ÷ bytes touched per forward pass. Batch one is the only mode a laptop has, so none of the amortization that makes datacenter serving cheap is available to it.

memory bandwidth 8B decode, one stream
laptop (M4 Pro / M4 Max) 273 / 546 GB/s claimed ~40–70 tok/s at 4-bit (ballpark, not our measurement)
server CPU (Xeon 6 with AMX) 691 GB/s claimed, 241 measured 10 tok/s measured
one MI325X 6,000 claimed, 4,390 measured 204 tok/s measured
one B200 8,000 claimed, 6,669 measured 269 tok/s measured
one B200 at batch 256 — 36,761 tok/s measured in aggregate

Read the last row as the economic argument. The laptop is not four times behind; it is behind by the batch factor, because a shared card answers 256 streams from the same weight reads. A laptop is competitive on latency for small models and never on cost per token. That is the whole local-versus-remote decision, and it is arithmetic rather than taste.

T4: why "just use a remote GPU" fails

GPU-over-IP — and its academic ancestor rCUDA — marshals CUDA calls across the network so a remote GPU appears local. It is the one architecture that drags an east–west concern through the north–south door, and the ratio at the top of this post kills it. A decode step is hundreds of kernel launches; at 269 tokens per second the entire budget for one token is 3.7 ms; a single 40 ms round trip per launch is four orders of magnitude over. It is a reasonable technology for bursty graphics and CAD. It is not a way to serve a language model. The honest version of "use a remote GPU" is T1: move the whole model call, not the kernel calls.

What the far side can see

In plain T1, everything that crosses is plaintext at the far end. The operator can read your prompt — and for a coding agent your prompt is your repository — plus your tool output, the replies, and all the metadata: timing, sizes, frequency. A zero-data-retention contract is a policy control over a plaintext system. Worth having. Not the same thing as a technical one, and worth never letting anyone blur that.

Three separable questions hide inside "is it private?":

  1. Confidentiality — can the operator read the data?
  2. Integrity — is the model and code that ran the one you asked for?
  3. Verifiability — can you prove the first two, or are you being told?

Obfuscation: the software options

option what it hides what it costs honest status
Don't send it (T0, or T2 routing) everything quality of the local tier the only unconditional answer
Prompt minimization — ship diffs and signatures, not whole files the bulk of the repo worse answers when context genuinely matters free, underused, biggest practical win
Redaction / PII scrubbing before the call names, keys, identifiers recall-precision tradeoff; mangles code semantics if clumsy mature, partial by construction
Client-side encryption at rest data at rest nothing mature, does not touch the problem above
Split inference — early layers local, later layers remote raw tokens, but not activations activations leak more than people assume; inversion attacks exist research-grade
MPC — secure multi-party evaluation inputs from any one party minutes per token and gigabytes of traffic not interactive
FHE — compute directly on ciphertext the input completely roughly 1,000–10,000× plaintext; GPU-accelerated encrypted GPT-2 is ~200× faster than the CPU FHE baseline and still far from real time batch and analytics only, not chat
Zero-data-retention terms nothing, technically nothing policy, not mechanism — say it out loud

The ordering matters more than the list. Prompt minimization is free, immediate, and usually removes more exposure than any cryptographic option you are realistically going to deploy this quarter. Start there.

Obfuscation: the hardware options

The pragmatic answer for real workloads is a TEE — a hardware boundary the operator cannot read into, plus an attestation, a signed measurement of the firmware and code you can check before sending anything.

layer mechanism what it covers notes
CPU / VM AMD SEV-SNP, Intel TDX guest memory encrypted and integrity-checked against the hypervisor the base of every confidential-GPU offering
NVIDIA Hopper (H100/H200) confidential-compute mode, on-die CC engine; single-GPU passthrough generally available from CUDA 12.4 VRAM encrypted; host-device transfers via encrypted bounce buffers overhead tracks your PCIe traffic share, not your FLOPs — compute inside the GPU is untouched
NVIDIA Blackwell (B100/B200) CC mode plus NVLink encryption multi-GPU confidential workloads this is the step that makes tensor-parallel serving possible in CC mode, not just single-card
AMD Instinct GPU inside the host's SEV-SNP envelope a VM-level boundary attestation covers the VM, not GPU firmware and VRAM state specifically — a weaker claim than Hopper/Blackwell CC; verify per platform
Attested TLS client verifies the report before the first token the whole path one-time startup cost, reported around 1–3 s per provisioning

Where this exists today: Azure pairs SEV-SNP confidential VMs with H100 CC mode; Google Confidential Space combines TDX/SEV-SNP with its attestation service; Apple's Private Cloud Compute is the largest deployed attested-inference system and runs partly on NVIDIA confidential computing. A directory of smaller TEE-based providers exists, several exposing OpenAI-compatible endpoints behind attested GPU enclaves.

Be precise about the trust you are buying. A TEE moves your trust from the operator's policies to the silicon vendor's root of trust, plus your own verification of the report. That is a real and large improvement. It is not "nobody can see it." It is "seeing it now requires breaking or backdooring the hardware, and I can check which hardware it is."

Choosing, by scenario

scenario topology obfuscation why
Personal work on open source T1 prompt minimization nothing to protect but habits
Proprietary code, ordinary commercial risk T1 contract + redaction + minimization policy control is proportionate
Regulated data — health, finance, defence-adjacent T1 attested CC-mode GPU, verified before the first token the only arrangement that survives the audit question
Client code under NDA T2 local model for bulk, remote for the hard calls, nothing identifying leaves keeps the frontier model without shipping the repo
Air-gapped T0 none needed the wire does not exist
Batch analytics over encrypted records remote FHE is genuinely viable here no interactivity requirement
Interactive chat over encrypted input — nothing works yet FHE and MPC are three to four orders of magnitude away

What nobody has measured

The confidential-computing numbers above are sourced but they are not ours, and the interesting ones are missing from the literature entirely. Three experiments would fix that, and they are cheap:

  1. CC-mode overhead per phase, on one card — the same sweep with confidential mode on and off. The mechanism predicts near-zero change in decode, where weights are already resident and traffic is HBM-to-SM, and a measurable loss in prefill-heavy and long-context work where host-device transfer share rises. Nobody publishes this split by phase, which is exactly the split that decides whether it is affordable for your workload.
  2. The energy cost of confidentiality — joules per token with CC on versus off. If encryption moves the bounce-buffer traffic, it shows up in board power.
  3. The attestation tax inside an agent loop — one to three seconds per provisioning is irrelevant for a long-lived endpoint and brutal for per-request enclaves. Which regime a provider runs is the question to ask them, and almost nobody asks it.

The compression: an agent's loop stays on your laptop, only the model call crosses the wire, and it crosses as text because the internet is thousands of times slower than the fabric on the far side. That one fact gives you the five topologies, explains why GPU-over-IP cannot work for decoding, and tells you that cost per token is a batch-size argument rather than a distance argument. It also tells you where your data goes — in plaintext, by default — which makes obfuscation a topology decision and not a checkbox: minimize the prompt first because it is free, route locally when the data cannot leave, and reach for an attested confidential GPU when the requirement is one you will have to prove to somebody. If you are drawing that boundary for a regulated workload and need the overhead measured rather than assumed, that's the work we do.

Related: Scatter the agent · KV handoff paths · Know, do, decide · Do agents run on CPU or GPU? · The agent loop, pedantically.