What I wanted to try
The clip above is not camera footage. Every frame is rendered from a set of 3D Gaussians that I fitted to 24 videos of the scene.
Spacetime Gaussians (Li et al., CVPR 2024) extends 3D Gaussian splatting to scenes that move. Each Gaussian gets a position and rotation that change over time and an opacity that fades in and out, so one model can render a person moving through a room from any viewpoint, in real time.
The paper’s results come from research datasets. For the Google Immersive dataset, the cameras are mounted on a rigid rig, synchronised and calibrated in advance and the dataset ships a models.json file with every camera’s exact pose and lens parameters. I did not have a rig. I had 24 phones, a mix of Samsung and Apple, recording 1080p at 30 fps, placed around a room, one person and a football. I wanted to find out how well the method works without the calibration.
Getting my videos into a pipeline built for a rig
I forked the official code and kept the model and training loop as they were. Everything I wrote sits in front of them.
1. Align the videos. The pipeline needs one video per camera, all aligned in time, so that the same frame index shows the same instant in every view.
2. Recover the camera poses with COLMAP. The rig dataset gives you poses; I had to estimate them. I took frame 0 from all 24 cameras and ran COLMAP’s automatic_reconstructor on those 24 images with a PINHOLE camera model. Taking a single time step turns the problem into ordinary structure-from-motion on a static scene and it gives each camera a rotation, a translation and a focal length.
3. Convert COLMAP’s output into the dataset’s format. This is the main piece of new code, extract_camera_info.py. It reads COLMAP’s binary cameras.bin and images.bin, sorts the cameras by name and writes two things: COLMAP text files with the poses fixed and a models.json in the Immersive schema so the existing preprocessing accepts my capture as if it came from the rig. Because my images come from COLMAP’s undistorted output, the lens distortion terms are written as zero.
4. Save distortion maps instead of undistorting up front. The “undistorted” preprocessing path used to warp every frame before training. I changed it to write frames as they are and save a per-camera distortion flow (camera_XXXX.npy) once, which the distortion-aware trainer (train_imdist.py) samples during training.
5. Make iteration cheap. The full method trains on 50 or 300 frames. I cut extraction to 10 frames and added checkpoints at iterations 10, 200, 500 and 1,000, so I could look at a render within minutes and check whether the poses were reasonable before starting a full run.
Setting up the GPU server
I trained on a rented GPU server. The code depends on custom CUDA rasterisers that compile against your exact CUDA and PyTorch versions and the machine came with CUDA 12 installed. The setup that compiled:
- CUDA 11.8 toolkit, installed by a script rather than the system package
- Python 3.10 in a fresh virtual environment
- PyTorch 2.0.0 built for cu118
CUDA_HOMEpointing at the 11.8 install, not the/usr/local/cudasymlink
What came out
The virtual camera sweeps across the scene while the person holds a squat. The cameras sat at a few fixed positions; everything between them is interpolated by the model.
For a first run, the result is good: the person is there, the motion is right and I can move the camera to positions no real camera occupied. The crouch and the kick in the first clip are reconstructed well.
As I expected with phones, it is clearly worse than the paper’s results. The background is streaky, there are floaters, one stretch of the full video shows a blue colour artefact and anything far from the person smears into radial blur. A handful of cameras around a room see the background from far fewer angles than a rig sees its scene and my phones were not synchronised by a shared clock the way a rig’s are, so frame 0 may not be quite the same instant in every video. Phones from different makers also differ in lenses, exposure and colour processing. A dedicated multi-camera rig with matched, genlocked cameras can cost upwards of $50,000 and it would give a noticeably better result.
What I would do next
- Synchronise the cameras by audio or a visual flash nd measure how far apart the “same” frame really is across the 24 videos.
- Report PSNR on a held-out camera rather than judging by eye.
After that, the next method I want to run on this capture is NVIDIA’s QUEEN (Girish et al., NeurIPS 2024). Spacetime Gaussians fits one model over a whole clip. QUEEN streams instead: it learns only the change in each Gaussian from one frame to the next, quantizes those changes and makes the position updates sparse. It also separates static from moving Gaussians, so it only trains the parts of the scene that change. The paper reports about 0.7 MB per frame, under 5 seconds of training per frame and about 350 FPS rendering, which starts to look like live volumetric video rather than an offline reconstruction.
The method in the paper is not mine. This project showed me how much a result like this depends on the calibration that the research datasets provide and how far a set of ordinary phones can get without it.