The KV-handoff post followed one tensor — the KV cache — as it crosses a disaggregated server. This is the post one level up: what the whole agent is, part by part, and where each part physically runs. The two connect at a single point. An agent is not a new kind of model and not a hardware tier; it is a control loop plus a handful of parts that do not share a machine. Name the parts, and they visibly scatter across hardware — and the part called memory turns out to be exactly the KV cache the handoff post was chasing.

The parts — a decoder ring

An agent is assembled from six parts. None of them is "the AI"; the AI is one of them.

  • Control loop — plain code, the harness (Claude Code, Aider, Codex CLI). It assembles the context, ships it to the model, parses what comes back, runs the requested tool, appends the result, and goes around again. This is the whole agent in the sense that everything else is a service it calls. Runs on a CPU.
  • Model call, prefill — the forward pass over the whole prompt at once. It reads every prompt token and produces the K and V tensors for each. Compute-bound; runs on a GPU built for FLOPs.
  • Model call, decode — the same weights run one token at a time, appending one row of K and V per step and reading the entire cache back every step. Bandwidth-bound; runs on a GPU (or CPU-class hardware with fast memory) built for bytes, not FLOPs.
  • Memory — two faces of one thing. The face you see is the context: text, re-fed in full every lap, because that is the only memory an agent has. The face you pay for is the KV cache that text becomes — the K and V tensors living out on the decode workers. This is the exact object post 46 is about, and the binding constraint on how much history an agent can carry.
  • Tools — sandboxes, external services, code runners, web fetchers. The model emits a tool call as text; the harness executes it somewhere else entirely. Their own CPUs and I/O, and the place an agent spends most of its wall-clock waiting.
  • Router / orchestrator — the piece that decides which worker serves this turn. In a disaggregated stack it is KV-aware: it tries to send the turn to whatever already holds its KV, which is the whole game once memory scatters.

Where each part lands

control loop (CPU)the harness — post 40holds the context: the only memoryassemble, forward, parse, execute, loopthe laprouter / orchestratorKV-aware: route to the KVprefillcompute-bound GPUsprompt or delta -> K, Vposts 46, 10GEMMs, ~0.4 sdecode workersbandwidth-bound GPUs, token by tokenKV w0KV w1KV w2KV cache — scattered across workershalf the history's KV is born hereKV poolhost RAM / SSD, RDMA — Mooncake-styleK,Vtokenstoolssandboxes / API nodestheir own CPUs and I/Opost 41: waiting on the worldtests, builds, web, coderetrievalvector storeexternal memory, not KVcall / resultan agent's memory rides on this orange thread — post 46 asks whether it must moveCPU: loop + tools' logiccompute-bound GPU: prefillbandwidth-bound GPU: decodeKV cache: memory that scatterspooled store: KV pool / vector DBtool sandbox: its own node

Where each part lands. Read the diagram as a spine with limbs. The control loop (purple, CPU) is the spine — it holds the context and drives every other part. Through the router it fans out to the model, which is itself split: prefill (blue, compute-bound GPUs) turns the prompt into K and V, hands them to the decode workers (green, bandwidth-bound GPUs) — the orange K,V edge is the single handoff post 46 draws in full — and decode then generates token by token, appending its own K and V locally. The KV cache is not one object in one place: it is the orange shards w0/w1/w2 scattered across the decode workers, spilling to a KV pool when it will not fit. The loop's other limb reaches tools (their own sandboxes) and retrieval (a vector store) — a different kind of memory that never touches the KV path. The same agent, drawn honestly, is a small distributed system with parts on at least four kinds of hardware.

The scatter is the point

The reason to draw it this way is that the parts do not just run in different places — their costs live in different places, and they do not add up the way a single box would suggest.

Part Hardware Bound by What it costs
Control loop CPU (the harness) branch logic, I/O cheap, continuous
Prefill compute-bound GPU FLOPs a burst per turn; grows with prompt length
Decode bandwidth-bound GPU memory bandwidth the token-by-token bill, plus a full KV read every step
Memory — context CPU RAM, re-fed as text prompt length more tokens to prefill every lap
Memory — KV cache scattered on decode GPUs / KV pool GPU HBM, then the wire where it lives, and whether it must move
Tools their own CPUs / I/O the world usually the dominant wall-clock
Retrieval vector store index + network a lookup, unrelated to KV

The row that matters is memory — KV cache, because it is the one part whose cost is a location rather than an amount of work. And that location is precisely what post 46 is about: an agent turn attends to a whole history whose KV was made on the decode side last turn, so the turn's cost is dominated by whether that KV is already under the worker serving it (move nothing), sitting in a pool a few milliseconds away (move deltas), or has to be re-shipped or recomputed wholesale (move everything). Those are post 46's designs A / B / C — recompute-and-ship, shared pool (Mooncake-style), sticky routing to the holder (Dynamo-style) — and they are not a serving-team detail an agent is insulated from. They are the agent's memory subsystem. Change which one your stack uses and you change the per-turn cost of the same agent doing the same task.

Memory has two faces

what you re-feed: context, as text (CPU)turn 1turn 2turn 3turn 4grows every lap — posts 40, 45each new token writes K, Vthe text becomes tensorswhere its KV lives: decode GPUs (post 46)w0w1w2scattered, and growing — moving it is the costYou see the text the loop re-feeds; you pay for the K, V tensors it becomes on the decode workers. Same memory, two faces.

Memory has two faces. The left bars are the memory an agent appears to have — the context, re-fed in full and growing every lap, which is why prompt-prefix caching is an agent's best friend. The right box is the memory it actually has on the hardware: the K and V tensors that text turns into, sitting out on the decode workers, scattered across them and growing every turn. The agent-writer reasons about the left; the bill is set by the right. That gap — text you re-feed versus tensors that must not move — is the whole reason an agent's serving cost is a distributed-systems problem and not a prompt-length problem.


The compression: an agent is a CPU control loop that rents split GPUs by the lap, keeps its memory as re-fed text whose KV scatters across the decode workers, and borrows its hands from tools in their own sandboxes. Six parts, at least four kinds of hardware, and one of those parts — the KV cache — is the shared thread with the KV-handoff post: its location, not its size, is what an agent turn actually costs. If you are sizing an agent-serving stack and trying to decide whether your memory subsystem should pool the KV, route to it, or eat the recompute, that's the work we do.

Related: KV handoff paths · Know, do, decide · Do agents run on CPU or GPU? · The agent loop, pedantically · KV cache, the binding constraint.