← Journal · · Deep dive

Speeding up image pulling: eStargz, Nydus, SOCI, and zstd

Hi everyone! Summer has finally arrived, and I want to spend more and more time outdoors, away from the laptop :) So posts in the channel have been infrequent lately, but I'm already looking for ways to fix that 😅

Right now I'd like to share a bit of research on optimizing image pulling and deployment :) As you remember, I earlier stumbled on eStargz, which had some really cool benchmarks. So I dug a little deeper into what else is interesting in containerd.

Honestly, this topic deserves an article on Habr, or even a talk, but for now I'll just share useful links I found :) In general it all comes down to using containerd with various add-ons :)

Nerdctl (specially for true nerds) is an analog of the Docker client that extends its functionality. This is the client I used for researching the technologies described below. Not too hard to install, you just need to download the binaries and put them in /usr/bin.

eStargz, we've already met it before: a modified gzip compression format, where the archive is not tar.gz but star.gz (which stands for seekable). Plus lazy loading: we pull layers on demand. I managed to test this technology, and here's what came out:

Nydus (always dreamed of sending you nudes 💪), another format, you can check benchmarks on the official website, honestly it looks like a fairy tale. I tried the Zran format, and per the benchmarks a Node18 image took 27 seconds versus 12. But the most interesting part is the image size: 1GB versus 15MB. Haven't found the reason yet, need to dig further. Couldn't convert my own ML image though.

SOCI 😂😂😂, probably the funniest name, but just as good as the rest. It's a fork of eStargz from Amazon. Tested it on a TensorFlow GPU image: 3m 39s versus 1m 45s, that's brilliant, roughly a 2x difference. But I also ran into a problem building my own images: you need a special registry with SOCI index support.

ZSTD, my crush 💋, because it's simply a new compression format. Gzip is an old format, and Facebook came up with algorithms to optimize all this. Build the image through buildx with zstd compression, and voila, the image is twice as light and extracts twice as fast. On my ML image it came out to 2m 30s versus 6m. This is my favorite!

Beyond that you could of course look at shared file systems like dragonfly and fight for every second of upload time, or check out the snapshotter from AliExpress. But right now it's not clear how much this is actually needed. For giants like Facebook, Alibaba, and Amazon, sure, it matters a lot for their workloads. But at smaller scale there's no such need yet. Still, if the need arises, I'll definitely come back to this and share the results with you :)

Original on Telegram ↗

↑↓ select · Enter open · Esc close