Is it odd for my local pex cache to have 164 copie...
# general
p
Is it odd for my local pex cache to have 164 copies of pyspark-4.1.1 ?
h
Under what paths?
p
that sounds like a no 😂, but lemme put the paths together
somewhere in
.cache/pants/named_caches/pex_root/
codex seemed to think it was in
*.work
directories in the first pass
162 files: 81 in pip cache, 80 in .work build dirs, and 1 final built wheel.
Copy code
81  ~/.cache/pants/named_caches/pex_root/pip/1/24.2/pip_cache/wheels/*/*/*/*/pyspark-4.1.1-py2.py3-none-any.whl
  80  ~/.cache/pants/named_caches/pex_root/built_wheels/0/local_projects/pyspark-4.1.1/*/*.work/pyspark-4.1.1-py2.py3-none-any.whl
  1   ~/.cache/pants/named_caches/pex_root/built_wheels/0/local_projects/pyspark-4.1.1/*/cp312-cp312-manylinux_2_35_x86_64/pyspark-
  4.1.1-py2.py3-none-any.whl
  1   ~/.cache/pants/named_caches/pex_root/packed_wheels/0/*/pyspark-4.1.1-py2.py3-none-any.whl
  1   ~/.cache/pants/named_caches/pex_root/pip/1/24.2/pip_cache/.tmp/pip-wheel-828vyao9/pyspark-4.1.1-py2.py3-none-any.whl
h
Looks like
pip
is the culprit then?
p
Maybe? I am mostly just wondering if this is a known issue or if its worth root causing and filing a ticket for
Codex thinks this is a cleanup failure in pex:
Copy code
I cloned and inspected:

  - research-pants at release_2.29.0
  - research-pex at v2.55.2, because Pants 2.29.0 uses pex==2.55.2

  The 164 PySpark copies are coming from two related mechanisms.

  What's Happening
  Pants sets PEX_ROOT to the shared named cache:

  research-pants/src/python/pants/backend/python/util_rules/pex_environment.py:165

  ~/.cache/pants/named_caches/pex_root

  Your Espresso lockfile pins PySpark as a source tarball, not a wheel:

  pyspark-4.1.1.tar.gz
  sha256: 77f78984...

  So PEX has to build a wheel from source.

  The exact 164 pyspark-4.1.1 wheel files are:

  81  pip/1/24.2/pip_cache/wheels/...
  80  built_wheels/0/local_projects/pyspark-4.1.1/.../*.work/...
  1   built_wheels/0/local_projects/pyspark-4.1.1/.../cp312-cp312-manylinux_2_35_x86_64/...
  1   packed_wheels/0/...
  1   pip/1/24.2/pip_cache/.tmp/...

  Why The Pip Cache Has 81 Copies
  PEX extracts the PySpark sdist into per-run temp dirs like:

  /tmp/pants-sandbox-.../.tmp/.../project/pyspark-4.1.1

  Then pip caches the built wheel using that local file://... origin. I checked the origin.json files: the cache contains 80 unique
  temp file:///tmp/pants-sandbox-... origins for the same PySpark wheel.

  Relevant PEX code:

  - PEX extracts sdists into temp dirs in research-pex/pex/resolver.py:388
  - PEX invokes pip with a persistent cache dir in research-pex/pex/pip/tool.py:455

  So pip is not deduping by package/version/hash here. It is effectively caching by changing local source path.

  Why The .work Dirs Piled Up
  PEX builds the wheel into an AtomicDirectory work dir:

  research-pex/pex/resolver.py:427

  But that path uses an unlocked/random .work dir. If multiple Pants/PEX processes try to build PySpark around the same time, they can
  all build into separate work dirs before the final cache dir exists.

  research-pex/pex/atomic_directory.py:150

  if self.is_finalized():
      return

  That early return means if another process already finalized the target, this process returns without cleaning its own .work dir. The
  docstring says losing the race should remove the work dir, but this path does not.

  That explains the orphaned .work dirs. On your machine there are currently:

  83 work dirs total for pyspark-4.1.1
  80 contain the pyspark wheel
  1 final built wheel dir

  Bottom Line
  This is not Pants intentionally wanting 164 copies. It is a bad interaction between:

  - PySpark being a huge source tarball in the lockfile
  - Pants running PEX in temp sandboxes
  - pip caching wheels by changing local file:// source paths
  - PEX 2.55.2 leaving losing/concurrent .work dirs behind

  The .work dirs look like the clearest stale/buggy part. The pip-cache duplicates are expected from pip's perspective, but wasteful in
  this PEX/Pants flow.
b
The 80 .work are failed wheel builds that neglected to cleanup the work dir. Does this match experience? (Did you encounter wheel build failures). Or not? (Maybe this is Pants killing Pex as part of it's speculation architecture).
And @powerful-scooter-95162 thanks for the PR, but no AI please. If you don't have time, just a good bug is better. I guess I need to erect an AI policy doc over in Pex.
p
Feel free to ignore the PR, I mostly sent it because it looked very short
b
It also looks very wrong
But, if you can explain how you know it works, that would be great
p
I had it include the repro in the issue: https://github.com/pex-tool/pex/issues/3187 I have no real confidence in it being a good fix architecturally, but it hopefully points to the root of the problem
b
I knew the root just from .work above
👍 1
p
I am not very familiar with the code, but I will say the AI had a slightly different thesis, which is that the finalizer bypasses cleanup
b
But my question is still useful to answer. Did you observe build failures or was this Pants preemption
p
I didn't notice any adverse effects on any builds, I was just looking at why my cache was so huge. The repro in the bug seems to not include failed builds and does not include pants in the repro
b
Ok, well when I get to it I'll try repro with modern Pex 1st, 2.55.2 is > 9 months old.
p
ah, I can try repro it with a newer one
The robot thinks the finalizer is still missing some cleanup, but can no longer repro the issue with 2.96.0
b
No robots please.
👍 1