ServerlessLLM InferenceSystems

Serverless platforms bill per second of inference, but loading a large model before the first token can take tens of seconds. This project follows that cold start end to end. It starts on a laptop, taking a 2 GiB checkpoint apart and measuring every stage until a loader-friendly checkpoint format logically follows from profile measurements in a bottom-up approach. Then it walks through ServerlessLLM (OSDI ‘24), which I co-authored: using the idle storage inside GPU servers as a checkpoint cache, live-migrating inference by moving tokens instead of gigabytes and scheduling for startup time. It ends with where this goes next, from agents where every step is a cold start to RL fine-tuning, where generation dominates the cost.

1 The Anatomy of an LLM Cold Start Taking a 2 GiB checkpoint apart on a laptop and timing every stage until a loader-friendly format logically follows from profling measurements in a bottom-up approach.
2 ServerlessLLM Paper Breakdown: Key Design Contributions The OSDI '24 paper, one contribution per post. Checkpoint caching on idle server storage, live migration of tokens and startup-time-aware scheduling.
3 Token Factories: A Design Proposal for Serverless LLM Agents Where this goes next. Agents where every step is a cold start and RL fine-tuning, where generation dominates the cost.
Distributed SystemsDeep Learning

Parallel SGD scales training out by exploding the batch size and synchronising every worker at every step and specialised networks only postpone the problem. KungFu, built with the Large Scale Data & Systems Group at Imperial College London and published at OSDI ‘20, lets the user declare how workers synchronise, monitor gradient and network statistics cheaply and change synchronisation strategy or parallelism at runtime. It started as my Master’s thesis, recognised as a Distinguished Project and was presented at the SOSP ‘19 AI Systems Workshop.

SecurityCompilers

Verification tools check the source program, not the binary the compiler produces, so code that is proven correct can still carry a back-door when a known compiler bug miscompiles it. In joint research with Cristian Cadar and Luís Pina at Imperial College London, I explored this attack vector: how it works, how to deploy it against the open-source programs Lighttpd and Vsftpd and how it could extend to cryptographic back-doors.