<#23526 Concurrent PEX builds can lose files from ...
# github-notifications
q
#23526 Concurrent PEX builds can lose files from the shared cache after cancellation Issue created by mjlbach Describe the bug Pants can nondeterministically fail while concurrently building PEX environments against its shared
pex_root
named cache. The observed sequence includes Pants cancelling and restarting one requirements PEX build while other PEX builds are active, followed intermittently by a missing file beneath the shared download cache:
Copy code
[INFO] Canceled: Building requirements.pex ...

ProcessExecutionFailure: Process 'Building requirements.pex ...' failed with exit code 1.
stderr:
[Errno 2] No such file or directory:
  .../named_caches/pex_root/downloads/1/<hash>.lck.work/<distribution>.whl
The failure is timing-dependent. Retrying the same CI workload may pass. Root cause (revised)
Note: this issue originally attributed the failure to locked-work-directory cleanup ownership in PEX. Deeper analysis of both the PEX and Pants code showed that theory doesn't hold up, and pex-tool/pex#3204 has been re-pointed at the actual defects. Summary of the revised analysis:
1. Download metadata v2 → v3 mass-invalidation. PEX bumped its downloaded-artifact metadata version from 2 to 3 in pex-tool/pex@826ef8e3 (pex-tool/pex#3112, first released in PEX 2.91.0). A
pex_root
cache populated by an older PEX (we upgraded from a Pants version bundling ~PEX 2.81 with the named cache restored across the upgrade) fails every cached-artifact load with a
LoadError
on first use under the new PEX. 2. Unlocked deletion of shared cache entries. On that
LoadError
,
DownloadManager.store
deletes the shared, finalized download directory with
safe_rmtree
while holding no lock, then re-downloads. Readers are lock-free by design (
atomic_directory
only locks when its target does not exist). Under Pants's concurrent PEX builds, every process that observed the stale v2 metadata queues up a deletion of the same directory — so one process deletes the entry a sibling process just repaired or is actively reading, and a wheel disappears mid-resolve. 3. Additionally, the fingerprint-addressed download cache had writers in two independent file-lock domains (POSIX
lockf
vs BSD
flock
, which do not exclude one another on Linux), so a lock-creation operation running concurrently with a lock resolve could put two writers inside the same critical section. The cancellation visible in the log is normal scheduler behavior (dynamic-concurrency preemption; Pants SIGKILLs the whole process group, so the cancelled build leaves no surviving writers and runs no cleanup). Its role here is that restarts increase concurrent contention on the same artifacts, which raises the probability of the deletion race in (2) — it is a trigger amplifier, not the defect. Fix • pex-tool/pex#3204 (revised): reads v2 metadata in place instead of invalidating it, removes invalid cache entries only under the artifact lock with a re-check for concurrent repair, and unifies all download-cache writers on BSD locks. This issue tracks updating Pants to a PEX release containing that fix. Because this affects the current Pants 2.32 stable line, please also consider a 2.32.x backport after the PEX release is available. Related report with a similar missing-cache-file symptom: • #21973 Workaround for anyone hitting this today: drop the restored
pex_root
named cache when crossing a Pants upgrade that moves PEX across the 2.91.0 boundary (the cache repopulates cleanly), or pin CI to not restore
named_caches/pex_root
across Pants version changes. Pants version 2.32.1 OS Linux (Ubuntu 24.04 CI runners) Additional info Pants 2.32.1 bundles PEX 2.95.1. pantsbuild/pants