The KV-handoff post followed one tensor — the KV cache — as it crosses a disaggregated server. This is the post one level up: what the whole agent is, part by part, and where each part physically runs. The two connect at a single point. An agent is not a new kind of model and not a hardware tier; it is a control loop plus a handful of parts that do not share a machine. Name the parts, and they visibly scatter across hardware — and the part called memory turns out to be exactly the KV cache the handoff post was chasing.
The parts — a decoder ring
An agent is assembled from six parts. None of them is "the AI"; the AI is one of them.
- Control loop — plain code, the harness (Claude Code, Aider, Codex CLI). It assembles the context, ships it to the model, parses what comes back, runs the requested tool, appends the result, and goes around again. This is the whole agent in the sense that everything else is a service it calls. Runs on a CPU.
- Model call, prefill — the forward pass over the whole prompt at once. It reads every prompt token and produces the K and V tensors for each. Compute-bound; runs on a GPU built for FLOPs.
- Model call, decode — the same weights run one token at a time, appending one row of K and V per step and reading the entire cache back every step. Bandwidth-bound; runs on a GPU (or CPU-class hardware with fast memory) built for bytes, not FLOPs.
- Memory — two faces of one thing. The face you see is the context: text, re-fed in full every lap, because that is the only memory an agent has. The face you pay for is the KV cache that text becomes — the K and V tensors living out on the decode workers. This is the exact object post 46 is about, and the binding constraint on how much history an agent can carry.
- Tools — sandboxes, external services, code runners, web fetchers. The model emits a tool call as text; the harness executes it somewhere else entirely. Their own CPUs and I/O, and the place an agent spends most of its wall-clock waiting.
- Router / orchestrator — the piece that decides which worker serves this turn. In a disaggregated stack it is KV-aware: it tries to send the turn to whatever already holds its KV, which is the whole game once memory scatters.
Where each part lands
Where each part lands. Read the diagram as a spine with limbs. The control loop
(purple, CPU) is the spine — it holds the context and drives every other part. Through
the router it fans out to the model, which is itself split: prefill (blue,
compute-bound GPUs) turns the prompt into K and V, hands them to the decode workers
(green, bandwidth-bound GPUs) — the orange K,V edge is the single handoff post 46
draws in full — and decode then generates token by token, appending its own K and V
locally. The KV cache is not one object in one place: it is the orange shards
w0/w1/w2 scattered across the decode workers, spilling to a KV pool when it
will not fit. The loop's other limb reaches tools (their own sandboxes) and
retrieval (a vector store) — a different kind of memory that never touches the KV
path. The same agent, drawn honestly, is a small distributed system with parts on at
least four kinds of hardware.
The scatter is the point
The reason to draw it this way is that the parts do not just run in different places — their costs live in different places, and they do not add up the way a single box would suggest.
| Part | Hardware | Bound by | What it costs |
|---|---|---|---|
| Control loop | CPU (the harness) | branch logic, I/O | cheap, continuous |
| Prefill | compute-bound GPU | FLOPs | a burst per turn; grows with prompt length |
| Decode | bandwidth-bound GPU | memory bandwidth | the token-by-token bill, plus a full KV read every step |
| Memory — context | CPU RAM, re-fed as text | prompt length | more tokens to prefill every lap |
| Memory — KV cache | scattered on decode GPUs / KV pool | GPU HBM, then the wire | where it lives, and whether it must move |
| Tools | their own CPUs / I/O | the world | usually the dominant wall-clock |
| Retrieval | vector store | index + network | a lookup, unrelated to KV |
The row that matters is memory — KV cache, because it is the one part whose cost is a location rather than an amount of work. And that location is precisely what post 46 is about: an agent turn attends to a whole history whose KV was made on the decode side last turn, so the turn's cost is dominated by whether that KV is already under the worker serving it (move nothing), sitting in a pool a few milliseconds away (move deltas), or has to be re-shipped or recomputed wholesale (move everything). Those are post 46's designs A / B / C — recompute-and-ship, shared pool (Mooncake-style), sticky routing to the holder (Dynamo-style) — and they are not a serving-team detail an agent is insulated from. They are the agent's memory subsystem. Change which one your stack uses and you change the per-turn cost of the same agent doing the same task.
Memory has two faces
Memory has two faces. The left bars are the memory an agent appears to have — the context, re-fed in full and growing every lap, which is why prompt-prefix caching is an agent's best friend. The right box is the memory it actually has on the hardware: the K and V tensors that text turns into, sitting out on the decode workers, scattered across them and growing every turn. The agent-writer reasons about the left; the bill is set by the right. That gap — text you re-feed versus tensors that must not move — is the whole reason an agent's serving cost is a distributed-systems problem and not a prompt-length problem.
The compression: an agent is a CPU control loop that rents split GPUs by the lap, keeps its memory as re-fed text whose KV scatters across the decode workers, and borrows its hands from tools in their own sandboxes. Six parts, at least four kinds of hardware, and one of those parts — the KV cache — is the shared thread with the KV-handoff post: its location, not its size, is what an agent turn actually costs. If you are sizing an agent-serving stack and trying to decide whether your memory subsystem should pool the KV, route to it, or eat the recompute, that's the work we do.
Related: KV handoff paths · Know, do, decide · Do agents run on CPU or GPU? · The agent loop, pedantically · KV cache, the binding constraint.