Disaggregated serving splits inference into a compute-bound prefill phase and a bandwidth-bound decode phase and puts them on separate pools of hardware (the idea behind Dynamo-style stacks; see also the KV cache as the binding constraint). The moment you do that, the KV cache has to move — and where it moves, and how much of it, is the whole engineering question. Sizes below are for Llama 3 70B, 320 KB of KV per token, over a 400 Gb/s link (50 GB/s).
One request, one handoff
One request, one handoff. The prompt's KV crosses the wire a single time and can stream layer by layer while prefill is still running. Decode appends its own KV locally, so nothing else crosses until the reply is done. The 52 ms is half a percent of the request — which is why disaggregation works on homogeneous clusters, and why the wire is not what would sink a cross-vendor split.
An agent's next turn, three ways
An agent's next turn, three ways. Turn N+1's new tokens attend to the whole history, and half of that history's KV was made on the decode side. Design A ships it all and recomputes; design B parks generated KV in a pool so only deltas move; design C sends the turn back to the worker that already holds the KV. The orange edges are the only bytes that matter, and their size is the difference between the designs.
The handoff as matrix operations
The handoff as matrix operations. Llama 3 70B: d = 8192, 64 query heads and 8
KV heads of 128, d_ff = 28672, 80 layers. Prefill runs every matrix multiply with
M = S rows; decode runs the same multiplies with M = 1 (or the batch). K and V
are the only outputs of the prefill pass that decode reuses, so they are the only
tensors that cross — everything else prefill computed is consumed inside the pass,
and the weights are already on both sides. Decode's bill is the weight sweep plus a
full read of the cache per step: the bandwidth wall that makes it the phase worth
disaggregating.
Why this settles it. The two columns are the same layer — same weights, same
matrix multiplies, in the same order. The only thing that differs is the row count
M: 8192 in prefill, 1 in decode. That one difference is the whole story: it is why
prefill is compute-bound and decode is bandwidth-bound, and it is why K and V — the
single product of the wide pass that every narrow pass reuses — are the only tensors
a handoff can possibly need to carry. Q, the scores, the FFN intermediates are all
recomputed per pass and thrown away; the weights already sit on both sides. So
everything the earlier diagrams drew as "the KV cache moving" is precisely this: the
K and V columns crossing, and nothing else. Once you see the split as one matrix
identity at two values of M, the design question stops being how do we move the
cache and becomes how do we avoid moving it — which is exactly what the pool and
the sticky router buy.
What a handoff weighs
| model | KV per token, BF16 | 8k prompt | 32k context | 32k over 50 GB/s |
|---|---|---|---|---|
| Llama 3.1 8B | 128 KB | 1.0 GB | 4.2 GB | 84 ms |
| Llama 3 70B | 320 KB | 2.6 GB | 10.5 GB | 210 ms |
| DeepSeek-V3 / R1, MLA | 70 KB | 0.6 GB | 2.3 GB | 46 ms |
KV per token is 2 × layers × KV_heads × head_dim × 2 bytes; MLA caches one
compressed vector per layer instead. A full-context move at 32k on a 70B model is a
time-to-first-token budget by itself — which is exactly the case that designs B and
C exist to avoid. Timings are transfer only (no encode, no handshake) and are
arithmetic, not measurements.
The takeaway is a shift of the question. For a single request the wire is a rounding error, so disaggregation is nearly free and the KV cache is not the thing to optimize. For an agent — many turns over a growing context, half the KV born on the decode side — the movement of the cache is the design axis, and the winning move is to not move it: pool it (Mooncake) or route to it (Dynamo). If you're sizing a disaggregated or agent-serving stack and trying to decide which of these three you actually need, that's the work we do.