Disaggregated serving splits inference into a compute-bound prefill phase and a bandwidth-bound decode phase and puts them on separate pools of hardware (the idea behind Dynamo-style stacks; see also the KV cache as the binding constraint). The moment you do that, the KV cache has to move — and where it moves, and how much of it, is the whole engineering question. Sizes below are for Llama 3 70B, 320 KB of KV per token, over a 400 Gb/s link (50 GB/s).

One request, one handoff

agentsends prompt, reads tokensprefill workerscompute-boundGEMMs over 8k prompt tokens~0.4 s, then donedecode workersbandwidth-boundone weight sweep per token500 tokens x 20 ms = 10 sprompt (text)prompt KV, once8k x 320 KB = 2.6 GB52 ms, streamed per layereach new token's KVwritten locallytokens back, ~2 bytes eachwall clock of the requestprefill 0.4 shandoff 0.05 s (0.5%)decode 10 s

One request, one handoff. The prompt's KV crosses the wire a single time and can stream layer by layer while prefill is still running. Decode appends its own KV locally, so nothing else crosses until the reply is done. The 52 ms is half a percent of the request — which is why disaggregation works on homogeneous clusters, and why the wire is not what would sink a cross-vendor split.

An agent's next turn, three ways

A recompute and ship allagent, turn N+1prefillre-prefills 32ktokens every turndecodeholds turn N KVunused next turnhistory + delta32k tokensfull KV10 GB, 200 mstokensmoves per turn:10 GBplus a 32k prefillTTFT pays both, every turnB shared KV poolagent, turn N+1prefillprefills the deltaonly (2k tokens)decodeany workerKV poolhost RAM / SSD, RDMA-reachabledelta onlydelta KV640 MB, 13 msturn N KV out160 MBhistory KV inwhat it lackstokensmoves per turn:~0.8 GBa few ms each hopMooncake-styleC sticky routing to the KV holderagent, turn N+1prefillidle for shortdeltasdecodethe worker thatholds turn N KVdelta, 2k tokensKV-aware routerprefills the deltaon the decode cardonly if the deltais long: delta KVtokensmoves per turn:~0delta prefill costs decode computeDynamo-style

An agent's next turn, three ways. Turn N+1's new tokens attend to the whole history, and half of that history's KV was made on the decode side. Design A ships it all and recomputes; design B parks generated KV in a pool so only deltas move; design C sends the turn back to the worker that already holds the KV. The orange edges are the only bytes that matter, and their size is the difference between the designs.

The handoff as matrix operations

PREFILL, one layer, S = 8192 prompt rowsDECODE, same layer, one step, 1 rowX [8192 x 8192]x Wq, Wk, WvQ [8192 x 8192]K [8192 x 1024]V [8192 x 1024]KV cache, layer L2 x [8192 x 1024] BF16= 32 MBscores = Q K^T [8192 x 8192], causal, per headP V -> [8192 x 8192]x Wo, + residualFFN: [8192 x 8192] x [8192 x 28672]gate and up, then down to [8192 x 8192]X' [8192 x 8192] -> next layerstays here, discarded after the pass: X, Q, scores,FFN intermediates. Weights: a copy on each side, never sent.GEMM shape [S x d] x [d x d_ff], M = S = 8192FLOPs/layer 2 x S x 855 M + attention = 16 TFLOPbytes/layer weights read once = 1.7 GBintensity ~ S = 8,000 FLOP/byte >> ridge 280compute-bound: pay for FLOPsx 80 layers = 2.6 GBthe only tensorthat crossesx [1 x 8192]x Wk, Wv, Wq: reads 0.3 GBk [1 x 1024]v [1 x 1024]q [1 x 8192]KV cache, layer L[8192+t x 1024] x 2append 1 row (4 KB), read all (32 MB)scores = q K^T [1 x 8192+t], per headP V -> [1 x 8192]x Wo, + residualFFN: [1 x 8192] x [8192 x 28672]reads 1.4 GB of weights to make 1 rowx' [1 x 8192] -> next layerthe cache prefill built is read in full at every step andgrows one row per layer per token; x, q, scores stay local.GEMM shape [B x d] x [d x d_ff], M = B = 1FLOPs/token 2 x 70 B = 0.14 TFLOPbytes/token 141 GB weights + 2.6 GB KV readintensity ~ B = 1 FLOP/byte << ridge 280bandwidth-bound: 18 ms/token at 8 TB/ssame weights on both sides, same layer math; the row count M is the whole difference,and K, V are the one product of the M = 8192 pass that every M = 1 pass needs.

The handoff as matrix operations. Llama 3 70B: d = 8192, 64 query heads and 8 KV heads of 128, d_ff = 28672, 80 layers. Prefill runs every matrix multiply with M = S rows; decode runs the same multiplies with M = 1 (or the batch). K and V are the only outputs of the prefill pass that decode reuses, so they are the only tensors that cross — everything else prefill computed is consumed inside the pass, and the weights are already on both sides. Decode's bill is the weight sweep plus a full read of the cache per step: the bandwidth wall that makes it the phase worth disaggregating.

Why this settles it. The two columns are the same layer — same weights, same matrix multiplies, in the same order. The only thing that differs is the row count M: 8192 in prefill, 1 in decode. That one difference is the whole story: it is why prefill is compute-bound and decode is bandwidth-bound, and it is why K and V — the single product of the wide pass that every narrow pass reuses — are the only tensors a handoff can possibly need to carry. Q, the scores, the FFN intermediates are all recomputed per pass and thrown away; the weights already sit on both sides. So everything the earlier diagrams drew as "the KV cache moving" is precisely this: the K and V columns crossing, and nothing else. Once you see the split as one matrix identity at two values of M, the design question stops being how do we move the cache and becomes how do we avoid moving it — which is exactly what the pool and the sticky router buy.

What a handoff weighs

model KV per token, BF16 8k prompt 32k context 32k over 50 GB/s
Llama 3.1 8B 128 KB 1.0 GB 4.2 GB 84 ms
Llama 3 70B 320 KB 2.6 GB 10.5 GB 210 ms
DeepSeek-V3 / R1, MLA 70 KB 0.6 GB 2.3 GB 46 ms

KV per token is 2 × layers × KV_heads × head_dim × 2 bytes; MLA caches one compressed vector per layer instead. A full-context move at 32k on a 70B model is a time-to-first-token budget by itself — which is exactly the case that designs B and C exist to avoid. Timings are transfer only (no encode, no handshake) and are arithmetic, not measurements.


The takeaway is a shift of the question. For a single request the wire is a rounding error, so disaggregation is nearly free and the KV cache is not the thing to optimize. For an agent — many turns over a growing context, half the KV born on the decode side — the movement of the cache is the design axis, and the winning move is to not move it: pool it (Mooncake) or route to it (Dynamo). If you're sizing a disaggregated or agent-serving stack and trying to decide which of these three you actually need, that's the work we do.