We build `.pex` binaries with pants, and I noticed...
# general
s
We build
.pex
binaries with pants, and I noticed there are options to
pex_binary
targets that lets you include or exclude sources and dependencies. I'm wondering if there's a smooth way to reduce packaging time in those cases where the dependencies haven't changed. For instance, we have quite a few binaries that package
torch
, which takes a cool 800MB and a handful of minutes to package, and it adds up to a lot of unnecessary time and resources. Is there a nice way to incrementally build the pex, or something else? I'm open to trying out the different
layout
options too. Right now we just build them with mostly default options to
pex_binary
.
e
The main thing I use those for is for building docker images (multistage builds). I build a "dependencies pex" and a "sources pex", and then put those each into their own docker images, then I build a final docker image that `COPY`s the files from each the deps and srcs images into the same directory. The result is that the docker builds look exactly as if they were built from a single all-encompassing pex, but they can make use of caching (the "dependencies pex" and image rarely change) from both pants (the deps pex/image) and from docker (final image has cached layers covering the dependencies copy). I think this idea is roughly what you are trying to do, but on a pex directly, is that right?
You might be able to do something like: • build separate deps and srcs pexes (with layout=loose) • have a "final" pex that depends on those and then builds a new pex by copying
**/*
from the sandbox into the new pex.
At a glance, I think this would do the caching you are hoping for, but I'm not sure if it would work correctly to try to depend on a built target such as a pex, or how that might end up being pulled into a sandbox for a later build
FYI: https://www.pantsbuild.org/blog/2022/08/02/optimizing-python-docker-deploys-using-pants This is how I have this set up for docker builds, and I think the principles are the same, but I'm not sure how perfectly it will translate to building pexes directly. Would love to hear back if this does work out for you
s
I think this idea is roughly what you are trying to do, but on a pex directly, is that right?
Exactly šŸ™‚ We've migrated our dockerfiles to the pattern you're suggesting, but we currently still just build one pex. What I'm not quite clear on how to do is set up the dependency pex in a way that the change inference will only trigger its build if only the dependencies have changed, and not the source code. If I just change a function body in a python file, the dependency pex shouldn't be rebuilt. If I import a new library though, it should. But to successfully infer the correct requirements, I think I need to add the same entrypoint
python_source
as the full binary as a dependency. I'll experiment a little bit and get back to you!
Related question, what flags are necessary to disable caching completely for
pants package
, if I'm looking to benchmark the time it takes? Is it enough to set
Copy code
--no-local-cache
--no-remote-cache
--no-pantsd
?
h
Yes, that should do it. In this context,
cache
means on-disk caching of the results of running processes, while
pantsd
memoizes build graph state in memory
The only problem is that when you turn off
pantsd
you are also measuring daemon startup overhead, which can be significant
s
Awesome, thanks šŸ™‚