Scattering the agent took an agent apart and showed where each piece lands inside a datacenter. This is the post about the one boundary that post crossed without comment: the wire between your laptop and the rented silicon. What sits on each side, what physically travels, why the line falls exactly there — and, since everything that travels is readable at the far end, what your options are for making it unreadable.
One ratio sets the topology
There are two protocol worlds, and conflating them is the usual confusion.
| world | links | latency | protocols |
|---|---|---|---|
| East–west — inside the datacenter | NVLink/NVSwitch, xGMI, InfiniBand | 17 µs measured for a small all-reduce across eight H100s; ~2–5 µs for RDMA between nodes | NCCL/RCCL, RDMA, NIXL |
| North–south — you to the datacenter | the internet | 20–80 ms round trip on the same continent | HTTPS+JSON, SSE, gRPC, MCP, A2A |
The internet is about 2,400× the latency of a small NVLink collective. Everything follows from that. You can afford to cross north–south once per request. You cannot afford to cross it once per layer, once per collective, or once per kernel launch.
Which is why the door is narrow on purpose: text goes through it, not GPU work. Your laptop never issues a CUDA call. It ships tokens and receives tokens.
LAPTOP (the loop lives here) ││ DATACENTER (the weights live here)
───────────────────────────── ││ ─────────────────────────────────────
editor / terminal ││
│ ││
▼ ││
┌───────────────────────────┐ ① prompt ││ ┌────────────────┐
│ AGENT LOOP │ ═══════════════╪═►│ gateway/router │ auth, quota,
│ plan → act → observe │ HTTPS + JSON ││ └───────┬────────┘ KV-aware placement
│ context assembly │ ││ ▼
│ tool dispatch │ ④ tokens ││ ┌────────────────┐
│ file edits, git, tests │◄═══════════════╪══│ ENGINE │ vLLM · SGLang ·
└────┬──────────────────────┘ SSE stream ││ │ ③ prefill │ TRT-LLM
│ ││ │ decode │ continuous batching,
│ ② MCP (JSON-RPC over stdio) ││ └───────┬────────┘ paged KV, prefix cache
▼ ││ ▼
┌──────────────┐ ││ ┌─────────────────────────────────┐
│ MCP servers │ files, git, db, search ││ │ GPUs │
└──────────────┘ ││ │ TP collectives: NCCL / NVLink │
││ │ KV transfer: NIXL / RDMA │
┌──────────────┐ optional, batch-1 only ││ └─────────────────────────────────┘
│ local model │ llama.cpp · Ollama · MLX ││ ▲ east–west: never crosses
└──────────────┘ ││ │ the door
⑤ MCP over Streamable HTTP ═════════════╪═► remote tool servers
Note the two different jobs MCP does here. MCP is not an inference protocol — it
connects the agent to tools and data, over stdio when the server is local and over
Streamable HTTP when it is remote. The model call is a separate thing entirely, almost
always an OpenAI-compatible /v1/chat/completions over HTTPS with the reply streamed back
as server-sent events. Agent-to-agent traffic is a third protocol again (A2A). Three
protocols, three different questions, one wire.
Five topologies
| # | topology | loop runs | weights run | crosses the wire | what binds it |
|---|---|---|---|---|---|
| T0 | all-local | laptop | laptop | nothing | laptop memory bandwidth |
| T1 | split — the default | laptop | datacenter | prompt + token stream | remote decode bandwidth |
| T2 | two-tier | laptop | both | the hard calls only | quality of your routing policy |
| T3 | all-remote | datacenter | datacenter | keystrokes, repo sync | your patience with the editor |
| T4 | GPU-over-IP | laptop | remote silicon, local address space | CUDA calls | network latency × launch count |
T1 is every coding agent you have used. T2 is a small local model doing the constant cheap work with a frontier model for the reasoning. T3 is cloud agents and remote development — your laptop becomes a terminal. T4 gets its own section, below, because it is the instructive failure.
The mechanics of one turn
Measured on a single B200, running an 8B model, on one of our own rented boxes:
| step | mechanism | cost |
|---|---|---|
| ① request | TLS + POST with ~30,000 tokens of assembled context | round trip 20–80 ms, once |
| ② tools | MCP over stdio to local servers — no network at all | sub-millisecond |
| ③ prefill | one compute-bound pass over the whole prompt | 86,197 tok/s measured → 0.35 s |
| ③ decode | one bandwidth-bound step per token, say 800 of them | 269 tok/s measured → 3.0 s |
| ④ stream | server-sent events, one per token | first token at ≈ 0.4 s |
The network is about 2% of the turn. The turn is decode-bound, and decode at batch one is a memory-bandwidth problem — which is precisely why the distant card wins anyway. The optimization consequence is blunt: shaving round-trip time is nearly worthless. Raising batch size is what moves cost per token, and on a shared endpoint that batch is other people's requests riding the same weight reads.
Where the local/remote line actually falls
Single-stream decode is memory bandwidth ÷ bytes touched per forward pass. Batch one is the only mode a laptop has, so none of the amortization that makes datacenter serving cheap is available to it.
| memory bandwidth | 8B decode, one stream | |
|---|---|---|
| laptop (M4 Pro / M4 Max) | 273 / 546 GB/s claimed | ~40–70 tok/s at 4-bit (ballpark, not our measurement) |
| server CPU (Xeon 6 with AMX) | 691 GB/s claimed, 241 measured | 10 tok/s measured |
| one MI325X | 6,000 claimed, 4,390 measured | 204 tok/s measured |
| one B200 | 8,000 claimed, 6,669 measured | 269 tok/s measured |
| one B200 at batch 256 | — | 36,761 tok/s measured in aggregate |
Read the last row as the economic argument. The laptop is not four times behind; it is behind by the batch factor, because a shared card answers 256 streams from the same weight reads. A laptop is competitive on latency for small models and never on cost per token. That is the whole local-versus-remote decision, and it is arithmetic rather than taste.
T4: why "just use a remote GPU" fails
GPU-over-IP — and its academic ancestor rCUDA — marshals CUDA calls across the network so a remote GPU appears local. It is the one architecture that drags an east–west concern through the north–south door, and the ratio at the top of this post kills it. A decode step is hundreds of kernel launches; at 269 tokens per second the entire budget for one token is 3.7 ms; a single 40 ms round trip per launch is four orders of magnitude over. It is a reasonable technology for bursty graphics and CAD. It is not a way to serve a language model. The honest version of "use a remote GPU" is T1: move the whole model call, not the kernel calls.
What the far side can see
In plain T1, everything that crosses is plaintext at the far end. The operator can read your prompt — and for a coding agent your prompt is your repository — plus your tool output, the replies, and all the metadata: timing, sizes, frequency. A zero-data-retention contract is a policy control over a plaintext system. Worth having. Not the same thing as a technical one, and worth never letting anyone blur that.
Three separable questions hide inside "is it private?":
- Confidentiality — can the operator read the data?
- Integrity — is the model and code that ran the one you asked for?
- Verifiability — can you prove the first two, or are you being told?
Obfuscation: the software options
| option | what it hides | what it costs | honest status |
|---|---|---|---|
| Don't send it (T0, or T2 routing) | everything | quality of the local tier | the only unconditional answer |
| Prompt minimization — ship diffs and signatures, not whole files | the bulk of the repo | worse answers when context genuinely matters | free, underused, biggest practical win |
| Redaction / PII scrubbing before the call | names, keys, identifiers | recall-precision tradeoff; mangles code semantics if clumsy | mature, partial by construction |
| Client-side encryption at rest | data at rest | nothing | mature, does not touch the problem above |
| Split inference — early layers local, later layers remote | raw tokens, but not activations | activations leak more than people assume; inversion attacks exist | research-grade |
| MPC — secure multi-party evaluation | inputs from any one party | minutes per token and gigabytes of traffic | not interactive |
| FHE — compute directly on ciphertext | the input completely | roughly 1,000–10,000× plaintext; GPU-accelerated encrypted GPT-2 is ~200× faster than the CPU FHE baseline and still far from real time | batch and analytics only, not chat |
| Zero-data-retention terms | nothing, technically | nothing | policy, not mechanism — say it out loud |
The ordering matters more than the list. Prompt minimization is free, immediate, and usually removes more exposure than any cryptographic option you are realistically going to deploy this quarter. Start there.
Obfuscation: the hardware options
The pragmatic answer for real workloads is a TEE — a hardware boundary the operator cannot read into, plus an attestation, a signed measurement of the firmware and code you can check before sending anything.
| layer | mechanism | what it covers | notes |
|---|---|---|---|
| CPU / VM | AMD SEV-SNP, Intel TDX | guest memory encrypted and integrity-checked against the hypervisor | the base of every confidential-GPU offering |
| NVIDIA Hopper (H100/H200) | confidential-compute mode, on-die CC engine; single-GPU passthrough generally available from CUDA 12.4 | VRAM encrypted; host-device transfers via encrypted bounce buffers | overhead tracks your PCIe traffic share, not your FLOPs — compute inside the GPU is untouched |
| NVIDIA Blackwell (B100/B200) | CC mode plus NVLink encryption | multi-GPU confidential workloads | this is the step that makes tensor-parallel serving possible in CC mode, not just single-card |
| AMD Instinct | GPU inside the host's SEV-SNP envelope | a VM-level boundary | attestation covers the VM, not GPU firmware and VRAM state specifically — a weaker claim than Hopper/Blackwell CC; verify per platform |
| Attested TLS | client verifies the report before the first token | the whole path | one-time startup cost, reported around 1–3 s per provisioning |
Where this exists today: Azure pairs SEV-SNP confidential VMs with H100 CC mode; Google Confidential Space combines TDX/SEV-SNP with its attestation service; Apple's Private Cloud Compute is the largest deployed attested-inference system and runs partly on NVIDIA confidential computing. A directory of smaller TEE-based providers exists, several exposing OpenAI-compatible endpoints behind attested GPU enclaves.
Be precise about the trust you are buying. A TEE moves your trust from the operator's policies to the silicon vendor's root of trust, plus your own verification of the report. That is a real and large improvement. It is not "nobody can see it." It is "seeing it now requires breaking or backdooring the hardware, and I can check which hardware it is."
Choosing, by scenario
| scenario | topology | obfuscation | why |
|---|---|---|---|
| Personal work on open source | T1 | prompt minimization | nothing to protect but habits |
| Proprietary code, ordinary commercial risk | T1 | contract + redaction + minimization | policy control is proportionate |
| Regulated data — health, finance, defence-adjacent | T1 | attested CC-mode GPU, verified before the first token | the only arrangement that survives the audit question |
| Client code under NDA | T2 | local model for bulk, remote for the hard calls, nothing identifying leaves | keeps the frontier model without shipping the repo |
| Air-gapped | T0 | none needed | the wire does not exist |
| Batch analytics over encrypted records | remote | FHE is genuinely viable here | no interactivity requirement |
| Interactive chat over encrypted input | — | nothing works yet | FHE and MPC are three to four orders of magnitude away |
What nobody has measured
The confidential-computing numbers above are sourced but they are not ours, and the interesting ones are missing from the literature entirely. Three experiments would fix that, and they are cheap:
- CC-mode overhead per phase, on one card — the same sweep with confidential mode on and off. The mechanism predicts near-zero change in decode, where weights are already resident and traffic is HBM-to-SM, and a measurable loss in prefill-heavy and long-context work where host-device transfer share rises. Nobody publishes this split by phase, which is exactly the split that decides whether it is affordable for your workload.
- The energy cost of confidentiality — joules per token with CC on versus off. If encryption moves the bounce-buffer traffic, it shows up in board power.
- The attestation tax inside an agent loop — one to three seconds per provisioning is irrelevant for a long-lived endpoint and brutal for per-request enclaves. Which regime a provider runs is the question to ask them, and almost nobody asks it.
The compression: an agent's loop stays on your laptop, only the model call crosses the wire, and it crosses as text because the internet is thousands of times slower than the fabric on the far side. That one fact gives you the five topologies, explains why GPU-over-IP cannot work for decoding, and tells you that cost per token is a batch-size argument rather than a distance argument. It also tells you where your data goes — in plaintext, by default — which makes obfuscation a topology decision and not a checkbox: minimize the prompt first because it is free, route locally when the data cannot leave, and reach for an attested confidential GPU when the requirement is one you will have to prove to somebody. If you are drawing that boundary for a regulated workload and need the overhead measured rather than assumed, that's the work we do.
Related: Scatter the agent · KV handoff paths · Know, do, decide · Do agents run on CPU or GPU? · The agent loop, pedantically.