Serverless LLMs · part 7 of 7
- Anatomy of an LLM Cold Start, Part 1: Where the Time Goes
- Anatomy of an LLM Cold Start, Part 2: Make It Predictable, Then Make It Fast
- Anatomy of an LLM Cold Start, Part 3: A Checkpoint Format Written for the Reader
- ServerlessLLM Paper Breakdown, Part 1: Your GPU Server Is Also a Storage Server
- ServerlessLLM Paper Breakdown, Part 2: Migrate Tokens, Not Gigabytes
- ServerlessLLM Paper Breakdown, Part 3: Scheduling for Startup Time
- Serverless Agents: When Every Step Is a Cold Start
The last six posts addressed the cold start for a single model: a checkpoint format built for readers, a store that turns idle server storage into a cache, migration that moves tokens instead of gigabytes and a scheduler that estimates startup time before committing. An agent, however, is usually a pipeline of models. This post looks at what happens when every node of the pipeline has a cold start and at what I think the system underneath should look like. Unlike the earlier posts, it presents an opinion rather than measurements and I point out where numbers are missing.
Agents from a systems perspective
From a systems point of view, an agentic workflow is a directed acyclic graph of functions, where some nodes call models. Production versions are rarely one big model behind an API. Inputs get sanitised and validated before they reach the generator, outputs get checked for safety and formatting after and there are often summarisers, rerankers, classifiers and retrieval steps in between. The nodes span modalities. A single workflow can involve several models of hundreds of gigabytes each, plus several small ones.
Two aspects make this a systems problem rather than a prompt-engineering one.
Heterogeneity. Some nodes call an API provider. Others call models the company hosts itself, because they are fine-tuned, or private, or cheaper at volume. The agent framework connects both and in practice the result is a mixed set of services that no single team fully owns.
How it gets built. Today each node tends to be a standalone microservice, hand-crafted per company, talking to its neighbours through message queues. Kafka is the default. Every team ends up maintaining its own bespoke inference platform and it gets harder to scale and change as the DAG grows and as more tools and data sources are added.
Why serverless suits agents
The requirements match what serverless platforms promise. Workloads are bursty and per-user. You want to pay per step, not per idle GPU. Developers want to write the DAG in Python and hand the infrastructure to someone else. The nodes are functions with inputs and outputs and the platform should own scheduling, scaling and failure.
Serverless Agent is the name I use for that idea: a cloud-native system where a user defines a multi-modal agentic workflow as a Python-first DAG and a container environment and the platform runs each node as a serverless function on shared GPU infrastructure. Existing agent frameworks give you API access to model providers. This is the equivalent for self-hosted models, with the infrastructure managed.
The problem is the one this series has been about. A serverless function that wraps a 70B model needs a 26-second download and a load of tens of seconds before it can run and an agent has several such functions.
Cold starts add up along the DAG
A modest agent: four nodes, three model sizes.
Take a modest DAG: an input guard on a small classifier, a 70B generator, a 7B safety model on the output, a 7B summariser. Run each node as a serverless function on a platform that fetches weights from object storage. From the paper posts we know what each of those cold starts costs on its own: tens of seconds for a 70B model over a cluster network, seconds for a 7B. In the naive design they are serialised along the DAG’s critical path, so the user’s latency is the sum of four cold starts plus four inferences and the inferences are a small part of it.
Without a cache the wall clock is mostly loading. Bar lengths are illustrative; measuring them is the experiment below.
For agents, the cold-start cost is multiplied by the number of nodes. The models are also heterogeneous, so the platform has to keep several different models close to the GPUs rather than a single one.
Two of the earlier designs apply directly. Put a ServerlessLLM store on every GPU server and the 7B nodes load in under a second from DRAM and the 70B in a few seconds from NVMe, instead of downloading. And a scheduler that knows startup times can route each node to the server where its weights already are.
Same DAG, same scale, weights already on the GPU servers.
The third idea is new and it is the part I find most interesting. The scheduler can see the DAG. When node N starts running, the scheduler already knows that node N+1 needs a 7B safety model and roughly when. It can prefetch that model into DRAM on some server, or migrate whatever is there, while node N is still generating. The ServerlessLLM scheduler answered “which server”. A DAG-aware scheduler would also answer “when”, which could remove the loading time from the critical path for every node except the first. I have not built this yet; it is only a sketch.
A sketch of the system
This is the architecture I have been drawing.

The DAG comes in from a Python client. Each node runs on a GPU server that has a ServerlessLLM store. State between nodes lives in a cache next to the compute, not in a message broker.
- A Python client defines the DAG: load and preprocess, input validator, LLM, output validator and so on.
- A control plane takes the DAG and turns it into deployments through a deployment service, with ServerlessLLM handling model placement and loading.
- A cluster manager runs the nodes as pods on GPU servers. Every server has a ServerlessLLM store, so checkpoints are cached across DRAM and SSD and loaded with the fast path. Object storage is only used when a model is not cached on any server.
- An inter-agent cache carries state between nodes. I have been prototyping with an in-process database such as LanceDB, so that intermediate results, embeddings and retrieved context live next to the compute and a node can read its predecessor’s output without a network hop to a broker.
- Tools hang off the same cache.
Building blocks I would not rebuild
As in the paper posts, I want to be clear about related work. Several pieces of this system already exist and I would use them rather than replace them.
- Modal is the closest thing to a serverless GPU platform built by people who understand the cold-start problem. Erik Bernhardsson’s writing on why he built it is worth reading before working on this area.
- DAG runners such as Metaflow, DAGWorks and the workflow layer of Stack AI already solve “define a pipeline in Python and run it somewhere”. What they lack is knowledge of where models are cached: none of them knows what a GPU server’s DRAM contains.
- SkyPilot handles the multi-cloud deployment layer.
- Langtrace and similar handle agent observability, which this design does not address.
- On communication, there is a growing literature on multi-agent protocols and on communication inside inference workloads, for example for multi-model image generation, that suggests the inter-node link deserves its own design rather than a general-purpose queue.
The piece that does not exist yet is the LLM-aware storage and scheduling layer underneath all of these. That is what the earlier posts built for a single model and what this one argues should be built for a DAG.
Open problems
Inter-node communication. Kafka is the common choice today and I think it is a poor fit for three reasons: payloads are large and structured rather than small events, the added latency per hop sits on the critical path of an interactive request and running a broker cluster is the kind of operational work a serverless platform is supposed to take away. The usual alternative, REST between services, has its own problems with large payloads and with backpressure. I do not have a measured answer yet, only an intuition that state should be passed through a cache placed next to the compute.
State between steps. Two DAG nodes that run the same base model with different adapters could share a KV cache prefix and a node that summarises a generator’s output has already got the tokens it needs. Nothing in today’s platforms lets a KV cache survive a node boundary. The migration idea from the paper, using tokens as the compact state, suggests one way to do it.
Scheduling with lookahead. The DAG-aware prefetch above. It is the ServerlessLLM scheduler with a time axis and a real evaluation would need traces of agent workloads that, as far as I know, do not exist publicly yet.
Multi-modal models. Image and audio models have different checkpoint shapes and much larger per-request payloads. The store works the same for any bytes, but the communication layer has to handle the larger payloads.
The experiment I want to run
The single most useful thing I could add to this post is a number and I do not have it yet. Below is the experiment, so that readers can check the argument instead of taking it on trust.
Build the four-node DAG above as serverless functions on Modal, or on a local FaaS stand-in. Run it cold and measure per-node cold start, end-to-end latency and cost. Then put a ServerlessLLM store under the platform, run it again and make a before-and-after table. The headline number of this post should come from that table. Then, only as a sketch, estimate what DAG-aware prefetch would take off the critical path given the measured load times.
As in the earlier posts, the scope is limited: no framework building, no multi-tenant fairness, no exactly-once semantics, one cloud provider.
Summary of the series
This series has seven posts. The first three built a loader on a laptop and found that once loading is fast the fixed costs dominate. The next three explained how a cluster avoids paying those fixed costs per request: cache the weights in storage it already owns, move tokens instead of gigabytes, estimate before scheduling. This last post argues that agents make the same problem larger, but also give the scheduler more information to work with.
What should I build from scratch next: the DAG-aware scheduler, or the agent-state store? I would like to hear which one you would rather read about.