<#23657 Pants 2.33.0 occasionally broken venvs in ...
# github-notifications
q
#23657 Pants 2.33.0 occasionally broken venvs in pants sandboxes with uv resolver Issue created by eguiraud-pf At Proxima Fusion we observe occasionally broken venvs in pants sandboxes in CI, running pants + uv. The symptom: in a CI run with a cold
PEX_ROOT
and 16-way
pants test
parallelism, about a third of a few hundred test targets failed, all identically, at `import pytest`: ​
Copy code
File ".../site-packages/pygments/formatters/__init__.py", line 18, in <module>
    from pygments.plugin import find_plugin_formatters
ModuleNotFoundError: No module named 'pygments.plugin'
​ i.e.
pygments
was present,
pygments/plugin.py
was not. This is speculation fed by pointing Claude at the logs, but this might be a leftover race after #23461 was fixed. See below for Claude's investigation of the issue: it would seem that
uv sync
is not synchronized with the PEX command that reads the same venv, so sometimes PEX reads half-formed venvs. I have not verified that claim. Pants version Pants 2.33.0,
[python] resolver = "uv"
, PEX 2.100.5 and 2.101.0, Linux, CPython 3.12. OS Linux Additional info Apologies for the AI-generated report, I tried to make it as useful as possible! ​ The lock asymmetry ​ Pants keeps one long-lived, mutable venv per resolve and hands it to PEX as
--venv-repository
. Writes to it are locked. Reads are not. ​ Where the lock is taken — `pants/backend/python/util_rules/uv.py:271-291`: ​ command = dedent( f"""\ cache_root="$({realpath_binary.path} {shlex.quote(VenvRepository.cache_dir)})" # :273 project_env="${{cache_root}}/{venv_path_suffix}" # :274 lock_path="${{project_env}}.lock" # :275 mkdir -p "$(dirname "${{lock_path}}")" # :276 ( if [ -x "{flock}" ]; then {flock} 200 || exit 1 # :279 elif [ -x "{pants_lock}" ]; then {pants_lock} 200 || exit 1 # :281 ... UV_PROJECT_ENVIRONMENT="${{project_env}}" {uv_cmd} # :288 ) 200>"${{lock_path}}" || exit $? # :289 """ ) ​ The lock lives on fd 200, which is opened by the
( ... ) 200>"$lock_path"
subshell at
:289
. The subshell contains exactly one thing:
uv sync
at
:288
. When the subshell exits, fd 200 closes and the lock is gone. ​ The venv path is shared across all consumers of a resolve — `uv.py:234-235`: ​ buildroot_entropy = hashlib.sha256(buildroot.path.encode()).hexdigest() venv_path_suffix = os.path.join(buildroot_entropy, metadata.resolve, request.python.fingerprint) ​ Where the lock is not taken. The PEX invocation is a separate subprocess of the same composite process — `pants/backend/python/util_rules/pex.py:938-941`: ​ composite_process = CompositeProcess.from_process(pex_process).prepend_subprocesses( [venv_repo.creation_subprocess] ) pex_process = await composite_process_to_process(composite_process, **implicitly()) ​ and the two are concatenated with a newline — `pants/engine/composite_process.py:173`: ​ command = "\n".join(subproc.get_command() for subproc in subprocs) ​ So the process that actually runs is: ​ ( flock 200; UV_PROJECT_ENVIRONMENT=$project_env uv sync --frozen ... ) 200>"$lock_path" pex ... --venv-repository=$project_env --no-transitive --no-pre-install-wheels # pex.py:834 ​ Line 2 reads the venv with no lock held.
flock
is never mentioned outside the subshell on line 1 — grepping
uv.py
for
flock
returns only
:279
and the
:289
redirect. ​ Consequence: N concurrent PEX builds for one resolve serialise their
uv sync
calls against each other, but any one of them can be reading the venv while another is syncing it. The comment at
uv.py:262-270
explains the lock exists because
uv sync
is not actually safe to run concurrently against one venv; the same reasoning applies to reading a venv while
uv sync
runs against it, and readers were left out. ​ What we have not established ​ We have not identified which write produced the missing file. Specifically we have not shown that
uv
ever leaves a package with its
RECORD
present and some of its files absent, which is the state our reproducer starts from. If
uv
writes
.dist-info
last, a mid-install reader would instead see no distribution at all — which is what the coarser error above looks like. So there are at least two candidate initiators we cannot currently distinguish: ​ 1. a concurrent
uv sync
writing while an unlocked PEX reads, and 2. an interrupted
uv sync
leaving a package that a later sync considers already installed — which would need no race at read time at all. ​ If (2) is what happens, the reader lock is not the trigger, and this is a
uv
durability issue plus the PEX caching issue. Either way the two code-level facts stand on their own: readers of a shared mutable venv take no lock, and PEX turns any partial read of one into a permanently cached artifact under a valid-looking key. We would rather report both than guess at the initiator. • Consider making the missing-file case at
pep_427.py:1014
an error when the source is a
--venv-repository
, keeping the warning for genuinely malformed third-party wheels. ​ Diagnosability — the
PEXWarning
at
pep_427.py:1016
is the only signal, and it is emitted once, by the build that poisons the cache. If that build is a cache hit in the run being debugged, the warning never appears and all you have is an
ImportError
. Running with
--pex-verbosity=1
and grepping for
bad RECORD
is what closed this out for us. pantsbuild/pants