<#14354 Better caching/sharing of heavy, first par...
# github-notifications
q
#14354 Better caching/sharing of heavy, first party imports Issue created by engnatha Is your feature request related to a problem? Please describe. I have a suite of tests that, due to the architecture of the repo, all share a common, large dependency. None of the transitive dependencies do anything expensive on import (just normal class, function, and constant definitions). However, in their aggregate, they are expensive to "compile". This import time takes on the order of 1 to 2 seconds to complete. When running in a virtual environment setup, we benefit from pycache such that these expensive imports do not have to be totally rebuilt. Thus, if I have 100 tests depending on this library, I pay the expensive import price once and carry on. In this scenario, it's likely that test execution is going to dominate the time of what I'm doing so the expensive import time is not a concern. When I run the same scenario with Pants, we lose all benefit of pycache. Each test is spun up in its own processes and left to perform the common heavy compilation task. So, despite tests taking at most 0.5 seconds, all 100 tests now have to spend the ~1.5 seconds of setting everything up. With full concurrency this caused a huge load on the system and actually makes the import time much slower than one off cases. This makes process construction much greater than actual test time and is prohibitive to development. Once tests are run, Pants caching helps prevent rerunning them when not needed, but it doesn't completely resolve the issue. It's reasonable to ask that developers not architect code with such a large number of transitive dependencies. However, since one of the primary goals of Pants is to make it easy to adopt into existing repos (which very likely will use virtual environments), this could be valuable to address in source or offer alternatives in the documentation to mitigate the effects I have seen. Describe the solution you'd like The reason tests are so fast now is because they all have access to the same pycache. A solution proposed on slack that I like is to be able to amortize the import construction by batching multiple tests into a single Pants process. These tests could still be kept fast by internally running those in parallel with pytest-xdist. Without digging deeply by any means, it seems like xml generation wouldn't need many changes as xdist already plays well with that system. Describe alternatives you've considered Restructuring the source code Additional context Working with Pants v2.9.0 pantsbuild/pants