What is KungFu?

KungFu makes adaptation for Deep Learning easy and efficient. There is increasing focus on improving training time of Deep Neural Networks through adaptation. For example, Open AI tried to use gradient noise to optimise the batch size. However, adaptive training is not an easy task: monitoring gradients can be very expensive, most system parameters are not reconfigurable, adaptation needs to manage a lot of GPU servers which can hurt the hardware efficiency. KungFu is built to address these challenges.

Our approach is to provide KungFu as a distributed training library which allows the user to monitor training metrics efficiently. In return, users can dynamically optimise the SGD parameters at runtime, even including the synchronisation and the level of parallelism.

The results are promising and show that KungFu can implement and scale adaptive training.

KungFu at SOSP’19

The end of October 2019 brought the opportunity to attend the Symposium of Operating Systems Principles in Canada, in order to support the poster exhibition of KungFu, which has recently been released open source. I contributed to KungFu during my Master’s Thesis and continued to get involved after finishing because I believe that easy and efficient distributed training is one of the main future research & industry directions in a world were datasets and models dramatically increase in size.

Insights from SOSP’19

The conference was an immensely valuable opportunity for me to synchronise with the research community and to learn about cutting edge advances in a broad range of fields: AI Systems, data center regional balancing, networking, storage, operating systems and verification.

Some papers with substantial research contribution in AI Systems presented at SOSP’19

In the following, I will briefly present new cutting edge work in AI Systems. My summaries provide an overview of the ideas, with insights from discussions and presentations. For a more in-depth view, please see the full papers referenced.

PipeDream [1]: Hybrid distributed training for data and model parallelism. For model parallelism, it splits the operators onto multiple nodes using a dynamic programming algorithm. Inter-node communication is done using PyTorch communication library. Data parallelism is supported through multi-stage dataflow replication.

TASO [2]: Annotations instrumenting graph computations of common libraries used in ML (e.g., numpy) to leverage cache coherency and to fit cache sizes for computations on multi-dimensional data. One example optimization focuses on the implementation of Black-Scholes model for option pricing, where arrays holding computation data can be partitioned to better leverage caches.

Parity Models [3]: Inference acceleration is one of hot topics in AI Systems, given the number of applications where latency-critical predictions are crucial (e.g., stock price prediction, multi-object recognition in videos). This work proposes an efficient way of performing model inference based on erasure codes. It may seem less intuitive to use coding theory in Machine Learning, but this work shows a way to recover predictions from straggler model servers. It relies on duplicate requests issued by the front-end load balancer and an additional trained parity model, which is used to recover models from stragglers.

Identifying new research directions in AI Systems

SOSP gave me a fresh perspective of the hot topics in AI Systems that will define future work in general. These serve the requirements imposed by industry and by users, who want to be effective and efficient in their workflows. In the following, I emphasize topics people approach to meet these requirements:

  • Unified AI platforms: Work of an ML engineer often requires a complex pipeline requiring dataset manipulation, model definition, hyperparameter tuning, training and deployment. Recent efforts have been directed towards unification of these steps into a single platform which centralises the pipeline and enacts control through automation. Some existing systems already achieve this: MLFlow, Uber Michelangelo and TensorFlow extended.
  • Federated learning: A new topic that focuses on preserving privacy when training models with private data which does not leave the user’s device. This topic is gaining a lot of traction, as it opens up the question of feasibility and the ability of existing systems to support it in production.
  • Systems with Graph Neural Network support: Any type of entity interactions (e.g., a social network) can be represented as a graph. These graphs can become increasingly large (think of Facebook’s 2.38 billion user social network). A Deep Learning scientist must be able to manipulate such graphs (e.g., sample edges) efficiently, in order to create embeddings and train a model. This can only be achieved by a flexible system that provides a high-level API that is expressive enough to capture the complexity of graph operations for Deep Learning. How can such systems be built? Can we integrate multiple data sources to build the graphs: distributed file systems, streaming data, etc.? Building such system will lead to easy and effective usage of GNNs for modern applications.
  • Treat ML as a compilers problem: The problem tackled is in inefficient programming decisions made by ML engineers, leading to training bottlenecks and device compatibility (e.g., CPU, GPU, TPU). This can be achieved by a new layer of abstraction which allows the engineer to define highly configurable blocks grouping usual tensor operators. These can be customized to generate code for a specific architecture and leave no space for slippery programming mistakes in the dataflow.
  • SQL queries answered using ML algorithms: What are the top K images from the dataset that contain a bird, a sailboat and a jumping dolphin all in the same picture? What are the videos containing people running? What are the videos containing a car that runs a red light? What are the videos containing Roger Federer in a one-to-one interview? These questions are interesting however it becomes very inefficient to ask for a dataset containing millions of videos. It translates into a frame sampling problem coupled with very fast ML inference. How can an SQL-like API be defined for this problem?
  • Fast inference: With the recent shift towards latency-critical applications such as object detection & tracking in videos, buy/sell decisions for automated trading, video queries inference becomes crucial. Research efforts are directed towards accelerating inference to address such practical applications. Can we design systems that do better than Clipper and TensorFlow Serving?
  • Runtime scaling of distributed training: When system load increases, the demand for accurate ML predictions increases as well. As training is time consuming, how can resources be allocated to address fast training for most up to date models.
  • Fast training: Throw more hardware at the problem? Usually not the solution. Use smarter synchronisation? Use adaptive training? How do we treat new models for which set of best hyperparameters is not known? We have tried to address some of these questions in our project, KungFu.

Soft skills learnt at SOSP

  • How to do an elevator pitch of a problem I am trying to solve, motivating my work.

What is the problem you are trying to solve? One should be most comfortable answering this question. In a professional dialogue, your focus should be entirely on your listener and your words should enlighten. In order to lead your listeners into a world they may never dare enter alone, your world, you must make it obvious to them why your work is worthwhile by clearly stating the problem you are trying to solve through your efforts. “Ha why haven’t I thought about this?” should be the self-addressed question of your listener after you make the problem obvious.

In order to do an effective elevator pitch you must be careful to incorporate a high level of specificity in your problem statement. Your speech should follow according to the following questions, in order: What is the problem? Why is it a problem? How are you solving the problem? Result (e.g., numbers, direct impact) of your approach. Make your answers as compact and as specific as possible and keep your speech short (30 - 50 seconds).

  • How to sustain a productive research conversation when diving deep into problems. I admit that in order to achieve this, I required multiple iterations and refinements of my own introductory routine (e.g., my problem statement). Moreover, I exercised listening skills and shut down my thoughts while the other person was speaking. I was really listening. This allowed me to engage, leading to a great conversation.
  • My participation at SOSP made me understand that in systems, research and industry are complementary and they support each other, leading to great technological achievements.

References:

[1] Deepak Narayanan (Stanford University), Aaron Harlap (Carnegie Mellon University), Amar Phanishayee (Microsoft Research), Vivek Seshadri (Microsoft Research), Nikhil R. Devanur (Microsoft Research), Gregory R. Ganger (CMU), Phillip B. Gibbons (Carnegie Mellon University), Matei Zaharia (Stanford University), PipeDream: Generalized Pipeline Parallelism for DNN Training,2019. Proceedings of the 27th ACM Symposium on Operating Systems Principles. ACM, New York, NY, USA

[2] Zhihao Jia (Stanford University), Oded Padon (Stanford University), James Thomas (Stanford University), Todd Warszawski (Stanford University), Matei Zaharia (Stanford University), Alex Aiken (Stanford Univeristy), TASO: Optimizing Deep Learning Computation with Automated Generation of Graph Substitutions, 2019. Proceedings of the 27th ACM Symposium on Operating Systems Principles. ACM, New York, NY, USA

[3] Jack Kosaian (Carnegie Mellon University), K. V. Rashmi (Carnegie Mellon University), Shivaram Venkataraman (University of Wisconsin-Madison), Parity Models: Erasure-Coded Resilience for Prediction Serving Systems, 2019. Proceedings of the 27th ACM Symposium on Operating Systems Principles. ACM, New York, NY, USA