What are people doing for managing image build tim...
# general
s
What are people doing for managing image build times with Pants? I've tried various things like buildkit and using local cache. But it seems like docker limitations (like only wanting to keep the latest cache) is causing our build times to consistently go over an hour. We should probably adopt changed-since more aggressively but it's an awkward fit for us at the moment.
a
I use pants dependents to find all the docker images that need to be rebuilt and then pass them to my matrix job and that uses the docker/build-push-action@v6 action in our CI to build images. Could you elaborate what you mean by build times? there are many optimisations i.e multi stage, switching base images, caching but none pants related
w
What are people doing for managing image build times with Pants?
Can you be a bit more specific as to what you're doing in the docker image? This feels like a "every case is different" problem
s
Our CICD solution likes to evict our docker cache even though we self-host it. In main we also don't do anything like --changed-since and "deploy" our entire monorepo during releases. These two things together are very painful. I tried to solve this by relying on --changed-since and then using argocd and "promoting" builds from dev to prod just to reduce the amount of work that needs to be done. My MLEs don't love the idea though and so releasing everything is the current solution.
w
I'm still not quite getting the full problem. How big are these images? What’s in them? What backends are in use? Etc
s
There are about 30 images. A good number have Tensorflow / Torch / Cuda deps. My optimizations are the usual src and deps splits as well as splitting out big python deps into separate PEXs and transitively excluding from deps. We spend a lot of time in
RUN PEX_TOOLS=1 python cli-deps.pex venv --collisions-ok --compile venv
steps. Everything is multistage so ideally only srcs change and
RUN PEX_TOOLS=1 python cli-srcs.pex venv --collisions-ok --compile venv
is trivial.
I wish I could push more of this in Pants and just have my images be dumb wrappers. But we can’t deal with the startup latency from PEXs unzipping themselves.
w
So, in the world where you’re only changing 1st party sources, you’re still seeing long build times?
s
Yeah. Our cache appears to be best effort and is PVCs that get passed around.
w
Alright, so, “our cache appears to be best effort” is something that doesn’t necessarily make sense. Is it that you’re overrunning a cache limit and other stuff is getting wiped out? Or is there a bug in the cache-from/cache-to? I was working with podman integration earlier today, and as far as I can tell, the github actions cache, and the github registry cache seem to work pretty well. And I have multi-stage builds, and I make it a point to cache dep layers if I can.
I guess I’m wondering if you can dig into what the “best effort” part of that is., It’s a hard problem to solve unless you know exactly what’s going wrong with the cache, or pants, or whatever
s
I should clarify we don’t currently use cache-from/cache-to and Im iterating on trying to make that work
w
Ah… Yeah… if you’re not using those, I can see why you’re using hours to build
s
Yeah - I was hoping we’d just be able to keep local docker cache around on persisted shared disk. Local with an NFS was painful because it kept rewriting cache even if was a cache hit. Registry is currently painful because of cache transfer with a docker-container builder. Im hoping that using docker as the builder with snapshot whatever works. I sort of oscillate between machine learning work and devops/build support. Everyone likes to complain about the hour + docker build times. No likes to solve for the hour long build times.
Pants has sort of spoiled me. A lot of it works so well and I guess I revert to getting grumpy when I need to manage my own caching / build stuff for things like docker lol.
w
Are you seeing a lot of lingering containers that haven’t been pruned/removed?
s
That hasn’t been an issue for us
If anything our existing caches get wiped too frequently. For registry based caching I think I can set a TTL
w
Are they getting overwritten? Is pants wiping them? Or is the system?
s
We use https://github.com/buchgr/bazel-remote for Pants and it works wonderfully. Our cache for docker is basically. The issue is squarely on docker cache getting wiped
Currently, our cache lives on ReadWriteOnce PVCs and if our CI pods get scheduled across different nodes we’re out of luck and some pods just won’t get cache. Its frankly a dumb setup but I don’t want to get into the weeds of it (though an easier solution might live there). My initial attempt to fix this was just mount an NFS so all nodes can share cache and its just been iterating from there
A full docker build with everything cached only takes 10 minutes (this is good lol)! This is down from 50 minutes (a naive buildkit docker container setup is soo slow)
👍 1
c
We have use cache-to/from with S3. But I struggle to figure out from
#13 CACHED
lines and the buildx otel how well this cache layer is working. (I can't give you a simple summary of the hit rage for example)