:wave: Hello! I'm trying to understand some pytho...
# general
f
šŸ‘‹ Hello! I'm trying to understand some python_test(s) target behavior I'm observing, and whether it's expected behavior, user error, or caused by a bug in the pants pytest_runner logic. An example file structure is pasted below. Basically, whenever I run
pants test src/python::
, I'm seeing that each individual
pytest
process builds a unique requirements PEX, with just the subset of requirements needed for that particular test process' run. So in this example scenario below, running
pants test src/python::
might produce two distinct requirements given the different 3rdparty requirements in
mod_a/a.py
vs
mod_b/b.py
. But that seems less ideal than building a single requirements pex with the requirements set intersection for all tests under the
src/python::
spec. Is this expected behavior to build multiple requirements.pex files for a test run, even if all the files fall under the same resolve?
Copy code
src/python
ā”œā”€ā”€ BUILD   # __defaults__(all=dict(resolve="my_resolve", interpreter_constraints=["CPython==3.12.*"]))
ā”œā”€ā”€ my-resolve-lockfile.json
ā”œā”€ā”€ mod_a
│   ā”œā”€ā”€ a.py  # uses anyio and pyyaml
│   ā”œā”€ā”€ test_a.py
│   └── BUILD   # contains a `python_tests` target
└── mod_b
    ā”œā”€ā”€ b.py  # uses pyyaml
    ā”œā”€ā”€ test_b.py
    └── BUILD   # contains a `python_tests` target
In this example, I see something like this in the
pants test
output:
Copy code
Completed: Building 2 requirements for requirements.pex from the src/python/lockfile.json resolve: anyio, pytest

Completed: Building 3 requirements for requirements.pex from the src/python/lockfile.json resolve: anyio, pytest, pyyaml
f
The main key here is the tracking of dependencies for caching results. You are correct that it would be more efficient to build one PEX with everything and run all the tests there. In fact, you can configure this by applying a batch_compatibility_tag to your tests. More efficient for one particular test run. And that's the big difference. Because the test run will be using all of the dependencies, pants won't be able to cache the results of individual test files. It will only be able to cache the batch against the full set of dependencies. So, making long term use of the caching means you will get greater optimization long term out of caching individual test files against their exact dependency set. You will have much more chance of being able to skip them on later test runs.
b
Also, see: https://www.pantsbuild.org/stable/reference/subsystems/python#run_against_entire_lockfile IOW: behavior is intentional - much of what @freezing-wall-19707 points out above is in the linked doc.
gratitude thank you 1
f
Ok great, thank you both for the helpful answers. I was aware of the batch_compatibility_tag, but I had misunderstood how it could factor into this until your explanation. And John thanks for pointing that out, that is very likely the lever I was originally looking for. Sounds like I may need to investigate what balance I can realistically strike between batch compatibility and cache effectiveness. The unfortunate thing in my situation is that the resolve in question is for running Airflow, which is notoriously requirement-heavy. That has very quickly led to
pants test ::
hitting OOM failures during the requirements.pex build phase in CI (we mitigated this for now by blanket applying the full set of requirements as a dependency to all tests) Anyways, thanks again for the helpful notes!
f
Limiting the number of workers to avoid running parallel builds is one option (although this obviously reduces parallelism by a lot.) Another question: Do you have some form of caching set up in your CI? If not, then by all means, run it all in a single batch because you can't benefit from the cache anyways
f
yeah we've got caching set up in CI, so we definitely want to strike a balance on the batching. Luckily, I think there might be a fairly natural way for us to segment the batch_compatability_tag values in our particular setup.
c
or running Airflow, which is notoriously requirement-heavy. That has very quickly led to pants test :: hitting OOM
Since airflow is open source, you may be able to create a public reproduction case for this. Pants should not run away with your memory.