Two earlier posts settled the geography. Scatter the agent took an agent apart and showed where each piece lands; the wire between you and the GPU drew the boundary between your laptop and the rented silicon and showed what crosses it. Neither answers the question you hit the moment you start building: which software do you install, and which of it can your agent operate without you?
That second clause is the one that is never sorted for, and it matters more than the feature lists. An agent cannot use a product whose value is a graphical interface. It can use anything with a stable text surface. Sort the whole landscape on that and it collapses into something you can hold in your head.
The cut that sorts everything
Ask two questions of each piece of software.
First: what role does it play in the loop?
- Brain — the endpoint the loop calls in order to think.
- Hand — something the loop invokes as a tool: run this, fetch that, start a box.
- Floor — infrastructure underneath the agent, which the agent never perceives at all.
Second: is there a non-GUI surface the agent's act-step can reach? If yes it is integrable. If its value is the interface a human sits in front of, it is application-only — useful to you, invisible to your agent.
The interesting result is that the two questions are not independent. Brains are essentially all integrable and nearly interchangeable. Hands are where genuine product choice lives. Floors are neither — and one whole category of "remote GPU" software turns out to be a floor pretending to be a brain.
Brains: the endpoint that thinks
| software | surface | note |
|---|---|---|
vLLM (vllm serve) |
OpenAI-compatible HTTP + SSE | the de-facto default |
| SGLang | OpenAI-compatible | RadixAttention; identical client code |
TensorRT-LLM (trtllm-serve) |
OpenAI-compatible | compiled at build time; best per-card latency |
| TGI (Hugging Face) | HTTP, OpenAI-compatible mode | |
| Triton Inference Server | HTTP and gRPC, KServe v2 | the gRPC option when JSON framing costs you |
| NVIDIA NIM | OpenAI-compatible containers | packaged engines |
| NVIDIA Dynamo | a router above the engines | a brain to the client, a floor to the engines |
| Red Hat AI Inference Server | packaged vLLM | the supported-enterprise path |
| Ollama · llama.cpp server · LM Studio | OpenAI-compatible, on your machine | the local tier |
| Vendor APIs — Anthropic, OpenAI, Gemini, Mistral, Cohere | HTTPS + SDKs | |
| Serverless inference — Together, Fireworks, Groq, Cerebras, Baseten, Replicate, Hugging Face | API, per-token | no box to operate |
| OpenRouter · LiteLLM | one shape across many providers | LiteLLM also self-hosts as a proxy |
All integrable, and that is the point. Every row speaks some version of
/v1/chat/completions, so switching is a base-URL change and a key. The portability of
agents is a direct consequence of one API shape winning. It also means the software choice
at this tier barely reaches your agent's design — it reaches your bill and your latency,
which is a different conversation.
The nervous system: protocols and frameworks
| software | role | status |
|---|---|---|
| MCP | agent ↔ tools and data; JSON-RPC over stdio locally, Streamable HTTP remotely | the tool-access standard |
| A2A | agent ↔ agent: task negotiation, capability discovery, state handoff | Linux Foundation project; spec v1.0, 150+ organizations, SDKs in five languages, GA inside Copilot Studio, Azure AI Foundry and Bedrock AgentCore |
| ACP (AGNTCY) | the wire/framing layer for agent messages — REST plus the minimum | donated to the Linux Foundation, backed by Cisco, LangChain, LlamaIndex, Dell, Oracle, Red Hat |
| Claude Agent SDK · OpenAI Agents SDK / Responses API | the loop itself, vendor-supplied | |
| LangGraph · CrewAI · AutoGen · Pydantic AI · smolagents · DSPy | the loop itself, framework-supplied |
Integrable by construction — this is the integration layer. The division of labour worth memorizing: MCP for tools, A2A between agents, an SDK or framework for the loop. They are not competitors, and a design that reaches for A2A where MCP belongs has usually mistaken a tool for a peer.
Hands: renting a box, and running jobs on it
This is where real product choice lives, and where the application-only rows appear.
| software | surface | can an agent drive it? | role |
|---|---|---|---|
| SSH + tmux/nohup + rsync | the shell | yes — agents shell out to this constantly | hand |
Provider CLIs — doctl, runpodctl, vast, lambda, brev, cloud CLIs |
CLI, mostly with JSON output | yes | hand |
| SkyPilot | CLI and a Python API, across clouds | yes — the best fit for "the agent provisions its own GPU" | hand |
| Modal | Python decorators; serverless GPU functions | yes, SDK-first | hand |
| Runhouse | dispatch a Python function to a remote GPU | yes | hand |
| Ray (+ Ray Jobs, Ray Serve) | distributed Python | yes | hand / floor |
| dstack · Beam · Coiled | job and dev orchestration | yes | hand |
| Slurm · Kubernetes (+ Kueue, Volcano, KubeRay) · Run:ai | sbatch, kubectl, APIs |
yes, but rarely from inside a loop | floor |
| Jupyter · Colab | a notebook interface | no — application-only | — |
| VS Code Remote-SSH · JetBrains Gateway | an editor experience | no — application-only | — |
| GitHub Codespaces · Replit | a hosted development environment | effectively no — CLIs exist but are thin | — |
The application-only rows are the ones people accidentally design around. A notebook is a superb way for you to hold a GPU session open and a terrible substrate for an agent, because the thing it provides — an interactive human loop with visible state — is exactly what the agent already has and does not need to borrow. "Remote development" and "remote inference" are different products that get conflated because both are described as using a GPU in the cloud.
Floors: device remoting, and the trap
| software | what it does | classification |
|---|---|---|
| Juice Labs | GPU-over-IP | application-only, transparent to the app |
| rCUDA | CUDA API remoting; academic lineage | application-only |
| scuda | open-source CUDA remoting | application-only |
| Thunder Compute | GPU over TCP | application-only |
| VMware vSphere Bitfusion | GPU pooling over the network — dead: availability ended 5 May 2023, support ended 5 May 2025 | — |
| NVIDIA vGPU · MIG | partitioning one card, not remoting it | floor |
None of these is integrable in any useful sense, because they sit below the layer the agent can see. More importantly, for language-model decoding they are the wrong layer entirely: as the wire post works through, a decode step is hundreds of kernel launches and the whole budget for one token is a few milliseconds, so paying a network round trip per launch is orders of magnitude over. The honest version of "use a remote GPU" is to move the whole model call, not the kernel calls. That the best-known product in this category has been end-of-life for a year is not an accident of the market.
Hands, continued: where the agent runs its tools
Sandboxes are the other half of the hand story — not where the model runs, but where the agent's code execution lands.
| software | GPU inside the sandbox? |
|---|---|
| Modal Sandboxes | yes — the one option that runs a GPU in-sandbox, at roughly a 3× pricing multiplier |
| E2B | not in the managed service; GPU requires self-hosting the open-source build |
| Daytona | claims conflict — some comparisons list GPU support, others report its gVisor isolation layer blocking GPU passthrough. Verify against current docs before designing around it |
| Fly Machines · Cloudflare Containers · Northflank · remote Docker contexts | varies; treat as CPU unless proven otherwise |
All integrable. The pattern to notice: isolation strength and GPU access trade against each other, because the sandboxing techniques that give the strongest guarantees are the ones that make device passthrough hardest. If your agent needs to run GPU code rather than merely call a model, that tension is your actual constraint.
And the confidential variants
Attested endpoints — confidential-mode GPUs behind an attestation you check before the first token — are brains with one extra step, and they are covered as a topology decision in the wire post rather than repeated here. They integrate exactly like row A, which is the useful part: choosing confidentiality does not change your agent's code.
Three rules that fall out
- At the brain tier, do not agonize. One HTTP shape won; pick on price, latency and support, and keep the base URL in configuration so the decision stays cheap to revisit.
- At the hand tier, choose deliberately. These surfaces are heterogeneous and your agent will be writing against them directly. A Python API beats a CLI beats a GUI, in that order, for everything an agent must operate unattended.
- If it only has a GUI, it is for you, not for the agent. This single test retires a surprising number of candidate tools before you evaluate a feature.
The compression: remote-GPU software sorts into a brain your agent calls, hands it invokes, and floors it never sees — and only the surface, not the feature list, decides which is which. Brains have converged on one API shape and are therefore nearly interchangeable; hands are where your architecture is actually decided; notebooks and remote editors are human tools wearing cloud-GPU clothing; and CUDA remoting is a floor that cannot be promoted to a brain no matter how the marketing reads. If you are assembling this stack and want the integration surfaces mapped against what you are actually trying to build — rather than against a vendor grid — that's the work we do.
Related: The wire between you and the GPU · Scatter the agent · Know, do, decide · Do agents run on CPU or GPU? · The agent loop, pedantically · Parallel AI agents in practice.