refined-lifeguard-66639
08/29/2026, 5:47 PMbrief-scientist-13682
08/29/2026, 5:50 PMbrief-scientist-13682
08/29/2026, 5:51 PMrefined-lifeguard-66639
08/29/2026, 5:51 PMrefined-lifeguard-66639
08/29/2026, 5:52 PMbrief-scientist-13682
08/29/2026, 5:52 PM--layout packed and not loose or zipapp, and getting to the bottom of the corruption is the preferred route.refined-lifeguard-66639
08/29/2026, 5:52 PMbrief-scientist-13682
08/29/2026, 5:53 PMrefined-lifeguard-66639
08/29/2026, 5:54 PMbrief-scientist-13682
08/29/2026, 5:56 PMbrief-scientist-13682
08/29/2026, 5:56 PMbrief-scientist-13682
08/29/2026, 5:58 PMbrief-scientist-13682
08/29/2026, 6:00 PM[pex-cli]
version = "v2.101.1"
known_versions = [
"v2.101.1|macos_x86_64|1fc9c277152e450ab5de2728eac6130b2433d8e94c491d9eaeeae63433a25c15|5314994",
"v2.101.1|macos_arm64|1fc9c277152e450ab5de2728eac6130b2433d8e94c491d9eaeeae63433a25c15|5314994",
"v2.101.1|linux_x86_64|1fc9c277152e450ab5de2728eac6130b2433d8e94c491d9eaeeae63433a25c15|5314994",
"v2.101.1|linux_arm64|1fc9c277152e450ab5de2728eac6130b2433d8e94c491d9eaeeae63433a25c15|5314994"
]
(That's in pants.toml)refined-lifeguard-66639
08/29/2026, 6:01 PMbrief-scientist-13682
08/29/2026, 6:01 PMrefined-lifeguard-66639
08/29/2026, 6:02 PMbrief-scientist-13682
08/29/2026, 6:02 PMrefined-lifeguard-66639
08/29/2026, 6:02 PMbrief-scientist-13682
08/29/2026, 6:03 PMbrief-scientist-13682
08/29/2026, 6:04 PMrefined-lifeguard-66639
08/29/2026, 6:05 PMbrief-scientist-13682
08/29/2026, 6:08 PMpip_version do you set in pants.toml if any?
+ What interpreter_constraints are in-play?
+ Do you enable resolves (locks)?
Basically, the more of pants.toml I can see the better.refined-lifeguard-66639
08/29/2026, 6:09 PMrefined-lifeguard-66639
08/29/2026, 6:14 PMpant.toml [pex-cli]
[pex-cli]
known_versions = [
"v2.97.3|macos_x86_64|61bc0afdb879b791c30211a809fddc542e36f6b4533240c85fb56753f6d65c05|5038874",
"v2.97.3|macos_arm64|61bc0afdb879b791c30211a809fddc542e36f6b4533240c85fb56753f6d65c05|5038874",
"v2.97.3|linux_x86_64|61bc0afdb879b791c30211a809fddc542e36f6b4533240c85fb56753f6d65c05|5038874",
"v2.97.3|linux_arm64|61bc0afdb879b791c30211a809fddc542e36f6b4533240c85fb56753f6d65c05|5038874",
]
version = "v2.97.3"
global_args = ["--tmpdir=/tmp"]
pants.ci.toml [pex-cli]
The --tmpdir stuff is because of the same issue George-Cristian reported with sockets on macos.
[pex-cli]
global_args = []brief-scientist-13682
08/29/2026, 6:14 PMrefined-lifeguard-66639
08/29/2026, 6:15 PMbrief-scientist-13682
08/29/2026, 6:15 PMrefined-lifeguard-66639
08/29/2026, 6:15 PMpants.toml [python]
[python]
enable_resolves = true
resolves_generate_lockfiles = true
interpreter_constraints = [">=3.10,<3.13"]
tailor_requirements_targets = falserefined-lifeguard-66639
08/29/2026, 6:16 PMbrief-scientist-13682
08/29/2026, 6:16 PMpip_version the version used depends on the Python version selected by Pex from the range.refined-lifeguard-66639
08/29/2026, 6:17 PMbrief-scientist-13682
08/29/2026, 6:18 PMbrief-scientist-13682
08/29/2026, 6:19 PMrefined-lifeguard-66639
08/29/2026, 6:19 PMrefined-lifeguard-66639
08/29/2026, 6:19 PMbrief-scientist-13682
08/29/2026, 6:19 PMrefined-lifeguard-66639
08/29/2026, 6:19 PMbrief-scientist-13682
08/29/2026, 6:20 PMrefined-lifeguard-66639
08/29/2026, 6:20 PMpants-sandbox dirs around. Which is why i initially said yesbrief-scientist-13682
08/29/2026, 6:21 PMrefined-lifeguard-66639
08/29/2026, 6:22 PMrefined-lifeguard-66639
08/29/2026, 6:23 PMbrief-scientist-13682
08/29/2026, 6:23 PMbrief-scientist-13682
08/29/2026, 6:23 PMrefined-lifeguard-66639
08/29/2026, 6:24 PMrefined-lifeguard-66639
08/29/2026, 6:36 PM2.97.1 from 2.81.0 to prepare for the upgrade of pants
• On aug 18th I moved pants' cache dirs on our build nodes to be outside the dockerized build environment so that future builds could potentially cache-hit on local artifacts.
• On the 19th I fixed our bazel-remote-cache servers to be properly provisioned / have a functioning loadbalancer
• On aug 20th the same teammate then bumped pants to 2.33.0 from 2.31.0 and raised pex to 2.97.3
• On the 26th that --tmpdir change for pex was merged but without the override in the pants.ci.toml. We thought that caused some of these issues that cropped up so we reverted & re-implemented with the pants.ci.toml global args override of []
• In various attempts to try to limit the scope of the build issue we see with these malformed artifacts we've moved most build state out of shared directories with the host and back into the container. None of that worked.
• Last night we finally were able to find a build that produced such a malformed artifact and persisted it into the remote-cache rather than consuming one. This helped us understand that we're still seeing the issue and prompted in part Aiden's proposed patch.refined-lifeguard-66639
08/29/2026, 6:39 PMrefined-lifeguard-66639
08/29/2026, 6:40 PMbrief-scientist-13682
08/29/2026, 6:43 PMLast night we finally were able to find a build that produced such a malformed artifact and persisted it into the remote-cache rather than consuming one.Good datapoint. How do you know the build produced the artifact? This was a build that had 0 local pants caches?
refined-lifeguard-66639
08/29/2026, 6:44 PMbrief-scientist-13682
08/29/2026, 6:44 PMbrief-scientist-13682
08/29/2026, 6:44 PMrefined-lifeguard-66639
08/29/2026, 6:45 PMrefined-lifeguard-66639
08/29/2026, 6:45 PMnamed-caches is independent of local cachebrief-scientist-13682
08/29/2026, 6:45 PMYeah we've disabled local caches entirelyWhat is the config that does that?
brief-scientist-13682
08/29/2026, 6:45 PMrefined-lifeguard-66639
08/29/2026, 6:46 PMrefined-lifeguard-66639
08/29/2026, 6:46 PMbrief-scientist-13682
08/29/2026, 6:46 PMbrief-scientist-13682
08/29/2026, 6:47 PMrefined-lifeguard-66639
08/29/2026, 6:47 PMrefined-lifeguard-66639
08/29/2026, 6:48 PMbuildroot/.pants.d/named-caches is not in fact ephemeral like we thought it was.brief-scientist-13682
08/29/2026, 6:48 PMrefined-lifeguard-66639
08/29/2026, 6:49 PMbrief-scientist-13682
08/29/2026, 6:49 PMrefined-lifeguard-66639
08/29/2026, 6:49 PMbrief-scientist-13682
08/29/2026, 6:50 PMbrief-scientist-13682
08/29/2026, 6:50 PMbrief-scientist-13682
08/29/2026, 6:51 PMrefined-lifeguard-66639
08/29/2026, 6:51 PMbrief-scientist-13682
08/29/2026, 6:51 PMrefined-lifeguard-66639
08/29/2026, 6:51 PMbrief-scientist-13682
08/29/2026, 6:52 PMbrief-scientist-13682
08/29/2026, 6:52 PMbrief-scientist-13682
08/29/2026, 6:52 PMrefined-lifeguard-66639
08/29/2026, 6:52 PMbrief-scientist-13682
08/29/2026, 6:53 PMbrief-scientist-13682
08/29/2026, 6:53 PMrefined-lifeguard-66639
08/29/2026, 6:53 PMbrief-scientist-13682
08/29/2026, 6:53 PMbrief-scientist-13682
08/29/2026, 6:53 PMrefined-lifeguard-66639
08/29/2026, 6:56 PMOk, but what you're relaying is what it claimed?Yes, it had something that it managed to repro with. Once I get you the docker fs question answer I'll actually review what it has.
brief-scientist-13682
08/29/2026, 6:56 PMrefined-lifeguard-66639
08/29/2026, 6:57 PMm8gd.4xlarge and m8id.4xlarge hosts in AWS.
All the disk at play here should be disk from the dedicated NVME side of those nodes (ofc they still have a root EBS as well)brief-scientist-13682
08/29/2026, 6:59 PMbrief-scientist-13682
08/29/2026, 7:01 PMrefined-lifeguard-66639
08/29/2026, 7:05 PMbrief-scientist-13682
08/29/2026, 7:06 PMrefined-lifeguard-66639
08/29/2026, 7:07 PMbrief-scientist-13682
08/29/2026, 7:08 PMrefined-lifeguard-66639
08/29/2026, 7:10 PMrefined-lifeguard-66639
08/29/2026, 7:10 PMbrief-scientist-13682
08/29/2026, 7:11 PMrefined-lifeguard-66639
08/29/2026, 7:17 PMbrief-scientist-13682
08/29/2026, 10:27 PM#!/usr/bin/env bash
set -euo pipefail
COUNT=30
MOD=3
export PEX_ROOT="${PEX_ROOT:-/tmp/pex-root}"
pex3 lock create --pip-version latest-compatible ansible --indent 2 -o /tmp/lock.json
rm -rf "${PEX_ROOT}"
declare -a kill_pex_pids
declare -a wait_pex_pids
for ((x=1; x<COUNT; x++)); do
pex ansible -c ansible --lock /tmp/lock.json --venv prepend --layout packed -- --version &
pid=$!
if (( x % MOD == 0 )); then
kill_pex_pids+=($pid)
else
wait_pex_pids+=($pid)
fi
done
for pid in "${kill_pex_pids[@]}"; do
sleep 2
kill ${pid}
done
for pid in "${wait_pex_pids[@]}"; do
wait ${pid}
done
I get things like this with exit code 0:
/home/jsirois/bin/tools.venv/lib/python3.14/site-packages/pex/atomic_directory.py:268: PEXWarning: [pid:60339, tid:131625590040256, cwd:/home/jsirois/support/pex/slack/8-29-2026]: After obtaining an exclusive lock on /tmp/pex-root/downloads/2/.fb06b66c8da04172d9e72a21d7d06186d8919e32ae5ab5cdf5b9d920be805ac2.atomic_directory.lck, failed to establish a work directory at /tmp/pex-root/downloads/2/fb06b66c8da04172d9e72a21d7d06186d8919e32ae5ab5cdf5b9d920be805ac2.lck.work due to: [Errno 17] File exists: '/tmp/pex-root/downloads/2/fb06b66c8da04172d9e72a21d7d06186d8919e32ae5ab5cdf5b9d920be805ac2.lck.work'
pex_warnings.warn(
/home/jsirois/bin/tools.venv/lib/python3.14/site-packages/pex/atomic_directory.py:268: PEXWarning: [pid:60347, tid:126699643143872, cwd:/home/jsirois/support/pex/slack/8-29-2026]: After obtaining an exclusive lock on /tmp/pex-root/downloads/2/.9e7dd367f7dc5d5e9fc5ae1baf8af9c4edc09e916a73a40108a3f32e3ad93f10.atomic_directory.lck, failed to establish a work directory at /tmp/pex-root/downloads/2/9e7dd367f7dc5d5e9fc5ae1baf8af9c4edc09e916a73a40108a3f32e3ad93f10.lck.work due to: [Errno 17] File exists: '/tmp/pex-root/downloads/2/9e7dd367f7dc5d5e9fc5ae1baf8af9c4edc09e916a73a40108a3f32e3ad93f10.lck.work'
pex_warnings.warn(
/home/jsirois/bin/tools.venv/lib/python3.14/site-packages/pex/atomic_directory.py:279: PEXWarning: [pid:60339, tid:131625590040256, cwd:/home/jsirois/support/pex/slack/8-29-2026]: Continuing to forcibly re-create the work directory at /tmp/pex-root/downloads/2/fb06b66c8da04172d9e72a21d7d06186d8919e32ae5ab5cdf5b9d920be805ac2.lck.work.
pex_warnings.warn(
/home/jsirois/bin/tools.venv/lib/python3.14/site-packages/pex/atomic_directory.py:279: PEXWarning: [pid:60347, tid:126699643143872, cwd:/home/jsirois/support/pex/slack/8-29-2026]: Continuing to forcibly re-create the work directory at /tmp/pex-root/downloads/2/9e7dd367f7dc5d5e9fc5ae1baf8af9c4edc09e916a73a40108a3f32e3ad93f10.lck.work.
pex_warnings.warn(
/home/jsirois/.pyenv/versions/3.14.7/lib/python3.14/multiprocessing/resource_tracker.py:475: UserWarning: resource_tracker: There appear to be 2 leaked semaphore objects to clean up at shutdown: {'/mp-eleeewhf', '/mp-inn89lbx'}
warnings.warn(
/home/jsirois/.pyenv/versions/3.14.7/lib/python3.14/multiprocessing/resource_tracker.py:475: UserWarning: resource_tracker: There appear to be 2 leaked semaphore objects to clean up at shutdown: {'/mp-r2k9mixn', '/mp-fv5ir0r1'}
warnings.warn(
/home/jsirois/.pyenv/versions/3.14.7/lib/python3.14/multiprocessing/resource_tracker.py:475: UserWarning: resource_tracker: There appear to be 2 leaked semaphore objects to clean up at shutdown: {'/mp-qg5m83q7', '/mp-wtniolvk'}
warnings.warn(
/home/jsirois/.pyenv/versions/3.14.7/lib/python3.14/multiprocessing/resource_tracker.py:475: UserWarning: resource_tracker: There appear to be 2 leaked semaphore objects to clean up at shutdown: {'/mp-7oaiujg4', '/mp-yawzw6h2'}
warnings.warn(
/home/jsirois/.pyenv/versions/3.14.7/lib/python3.14/multiprocessing/resource_tracker.py:475: UserWarning: resource_tracker: There appear to be 6 leaked semaphore objects to clean up at shutdown: {'/mp-3elo32t7', '/mp-jb8j0soh', '/mp-__qceg7m', '/mp-19byohln', '/mp-7xca8xo8', '/mp-u8yvl5h8'}
warnings.warn(
Process ForkServerPoolWorker-1:
Traceback (most recent call last):
File "/home/jsirois/.pyenv/versions/3.14.7/lib/python3.14/multiprocessing/pool.py", line 131, in worker
put((job, i, result))
~~~^^^^^^^^^^^^^^^^^^
File "/home/jsirois/.pyenv/versions/3.14.7/lib/python3.14/multiprocessing/queues.py", line 397, in put
self._writer.send_bytes(obj)
~~~~~~~~~~~~~~~~~~~~~~~^^^^^
File "/home/jsirois/.pyenv/versions/3.14.7/lib/python3.14/multiprocessing/connection.py", line 210, in send_bytes
self._send_bytes(m[offset:offset + size])
~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^
File "/home/jsirois/.pyenv/versions/3.14.7/lib/python3.14/multiprocessing/connection.py", line 441, in _send_bytes
self._send(header)
~~~~~~~~~~^^^^^^^^
File "/home/jsirois/.pyenv/versions/3.14.7/lib/python3.14/multiprocessing/connection.py", line 404, in _send
n = write(self._handle, buf)
BrokenPipeError: [Errno 32] Broken pipe
...
ansible [core 2.21.3]
config file = None
configured module search path = ['/home/jsirois/.ansible/plugins/modules', '/usr/share/ansible/plugins/modules']
ansible python module location = /tmp/pex-root/venvs/3/0006b71e2cfba8fe25d1d711b7e930f1aec28f6f/c40de391655dede9f50dced114f8a1dc8f21dd91/lib/python3.14/site-packages/ansible
ansible collection location = /home/jsirois/.ansible/collections:/usr/share/ansible/collections
executable location = /tmp/pex-root/venvs/3/0006b71e2cfba8fe25d1d711b7e930f1aec28f6f/c40de391655dede9f50dced114f8a1dc8f21dd91/pex
python version = 3.14.7 (main, Aug 5 2026, 17:40:48) [GCC 15.2.0] (/tmp/pex-root/venvs/3/0006b71e2cfba8fe25d1d711b7e930f1aec28f6f/c40de391655dede9f50dced114f8a1dc8f21dd91/bin/python)
jinja version = 3.1.6
pyyaml version = 6.0.3 (with libyaml v0.2.5)ansible [core 2.21.3]
config file = None
configured module search path = ['/home/jsirois/.ansible/plugins/modules', '/usr/share/ansible/plugins/modules']
ansible python module location = /tmp/pex-root/venvs/3/0006b71e2cfba8fe25d1d711b7e930f1aec28f6f/c40de391655dede9f50dced114f8a1dc8f21dd91/lib/python3.14/site-packages/ansible
ansible collection location = /home/jsirois/.ansible/collections:/usr/share/ansible/collections
executable location = /tmp/pex-root/venvs/3/0006b71e2cfba8fe25d1d711b7e930f1aec28f6f/c40de391655dede9f50dced114f8a1dc8f21dd91/pex
python version = 3.14.7 (main, Aug 5 2026, 17:40:48) [GCC 15.2.0] (/tmp/pex-root/venvs/3/0006b71e2cfba8fe25d1d711b7e930f1aec28f6f/c40de391655dede9f50dced114f8a1dc8f21dd91/bin/python)
jinja version = 3.1.6
pyyaml version = 6.0.3 (with libyaml v0.2.5)
ansible [core 2.21.3]
config file = None
configured module search path = ['/home/jsirois/.ansible/plugins/modules', '/usr/share/ansible/plugins/modules']
ansible python module location = /tmp/pex-root/venvs/3/0006b71e2cfba8fe25d1d711b7e930f1aec28f6f/c40de391655dede9f50dced114f8a1dc8f21dd91/lib/python3.14/site-packages/ansible
ansible collection location = /home/jsirois/.ansible/collections:/usr/share/ansible/collections
executable location = /tmp/pex-root/venvs/3/0006b71e2cfba8fe25d1d711b7e930f1aec28f6f/c40de391655dede9f50dced114f8a1dc8f21dd91/pex
python version = 3.14.7 (main, Aug 5 2026, 17:40:48) [GCC 15.2.0] (/tmp/pex-root/venvs/3/0006b71e2cfba8fe25d1d711b7e930f1aec28f6f/c40de391655dede9f50dced114f8a1dc8f21dd91/bin/python)
jinja version = 3.1.6
pyyaml version = 6.0.3 (with libyaml v0.2.5)
...
N.B.: This has the failed to establish a work directory at ... Continuing to forcibly re-create the work directory at logging @echoing-psychiatrist-92346 reported in the PR due to the kills as well as multiprocessing complaining about kills, but all the non-killed Pex processes run to successful completion.brief-scientist-13682
08/29/2026, 10:29 PMrefined-lifeguard-66639
08/29/2026, 11:43 PMbrief-scientist-13682
08/29/2026, 11:45 PM#!/usr/bin/env bash
set -euo pipefail
CONTAINER_CONCURRENCY=5
docker build -t pex-cache-write-fail-repro .
mkdir -p $PWD/pex-root
sudo find $PWD/pex-root/ -type f -name "*.lck" | while read lock_file; do
# pex-root/downloads/2/.b727414169a36b7d524c1c3e31839a521725078d7b2ff038656844266160a992.atomic_directory.lck
base_dir="$(dirname "${lock_file}")"
lock_file_name="$(basename "${lock_file}")"
target_dir_name="$(echo "${lock_file_name}" | cut -d. -f2)"
sudo rm -rf "${base_dir}/${target_dir_name}"
done
declare -a pids
for ((x=0; x<CONTAINER_CONCURRENCY; x++)); do
docker run --rm -v $PWD/pex-root:/tmp/pex-root pex-cache-write-fail-repro:latest repro.sh &
pids+=($!)
done
for pid in "${pids[@]}"; do
wait $pid
donerefined-lifeguard-66639
08/29/2026, 11:46 PMbrief-scientist-13682
08/29/2026, 11:48 PMrefined-lifeguard-66639
08/29/2026, 11:48 PMrefined-lifeguard-66639
08/29/2026, 11:49 PMbrief-scientist-13682
08/29/2026, 11:50 PMrefined-lifeguard-66639
08/29/2026, 11:50 PMrefined-lifeguard-66639
08/29/2026, 11:50 PMbrief-scientist-13682
08/29/2026, 11:51 PMrefined-lifeguard-66639
08/30/2026, 1:31 AMbrief-scientist-13682
08/30/2026, 1:35 AMrefined-lifeguard-66639
08/30/2026, 1:36 AMrefined-lifeguard-66639
08/30/2026, 1:37 AM[
"/pants-named-caches/python_build_standalone/aacb154e49a3600820e795704db17f7177cf27c4e08fc7b159882144f27f022d/bin/python3",
"./pex",
"--tmpdir",
".tmp",
"--jobs",
"{pants_concurrency}",
"--no-emit-warnings",
"--pip-version",
"24.2",
"--python-path",
"/usr/local/bin/python3",
"--output-file",
"python.klaviyo.redacted.server/app-deps.pex",
"--emit-warnings",
"--check=warn",
"--venv",
"prepend",
"--include-tools",
"--runtime-pex-root=/tmp/pex-redacted-server-app-deps",
"--requirements-pex",
"local_dists.pex",
"--interpreter-constraint",
"CPython<3.13,==3.12.*,>=3.10",
"--python-path",
"/usr/local/bin/python3.12",
"--entry-point",
"redacted.server.main:main",
"--sources-directory=source_files",
"REDACTED many deps, left grpcio cause that's the one that failed",
"grpcio",
"--lock",
"3rdparty/python/deps_lock.txt",
"--no-pypi",
"--index=<https://pypi.org/simple/>",
"--index=<https://REDACTED>",
"--manylinux",
"manylinux2014",
"--layout",
"packed"
]
I think relevantly it looks like pants is explicitly setting pip 24.2refined-lifeguard-66639
08/30/2026, 1:37 AMbrief-scientist-13682
08/30/2026, 1:39 AMrefined-lifeguard-66639
08/30/2026, 1:41 AMbrief-scientist-13682
08/30/2026, 1:42 AMbrief-scientist-13682
08/30/2026, 1:43 AMbrief-scientist-13682
08/30/2026, 1:43 AMbrief-scientist-13682
08/30/2026, 1:44 AMbrief-scientist-13682
08/30/2026, 1:44 AMrefined-lifeguard-66639
08/30/2026, 1:52 AMbrief-scientist-13682
08/30/2026, 1:53 AMbrief-scientist-13682
08/30/2026, 1:54 AMrefined-lifeguard-66639
08/30/2026, 1:54 AMrefined-lifeguard-66639
08/30/2026, 1:54 AMbrief-scientist-13682
08/30/2026, 1:56 AMrefined-lifeguard-66639
08/30/2026, 1:59 AMbrief-scientist-13682
08/30/2026, 2:00 AMbrief-scientist-13682
08/30/2026, 2:00 AMbrief-scientist-13682
08/30/2026, 2:01 AMrefined-lifeguard-66639
08/30/2026, 2:02 AMbrief-scientist-13682
08/30/2026, 2:02 AMrefined-lifeguard-66639
08/30/2026, 2:02 AMrefined-lifeguard-66639
08/30/2026, 2:03 AMbrief-scientist-13682
08/30/2026, 2:03 AMrefined-lifeguard-66639
08/30/2026, 2:04 AMrefined-lifeguard-66639
08/30/2026, 2:04 AMrefined-lifeguard-66639
08/30/2026, 2:05 AMbrief-scientist-13682
08/30/2026, 2:06 AMbrief-scientist-13682
08/30/2026, 10:07 AMNope EBS and instance-store only
@refined-lifeguard-66639 just checking: not EBS Multi-Attach I presume https://docs.aws.amazon.com/ebs/latest/userguide/ebs-volumes-multi.html#considerations
echoing-xylophone-39473
08/30/2026, 12:48 PM"Iops": 6000,
"VolumeType": "gp3",
"MultiAttachEnabled": false,brief-scientist-13682
08/30/2026, 3:27 PMechoing-xylophone-39473
08/31/2026, 3:15 AMv2.101.2 Pex update, we're testing this out and will update by tomorrowbrief-scientist-13682
08/31/2026, 6:43 PMrefined-lifeguard-66639
08/31/2026, 7:11 PMLAST_ACCESS_FILE write since we run the app in a container with a RO filesystem, so had to roll back / re-roll out. I believe the re-rollout is happening ~now unsure if we've had enough builds go through to hit a repro yet.brief-scientist-13682
08/31/2026, 7:13 PMLAST_ACCESS_FILE is old news for Pex; so how was this something new you hit?refined-lifeguard-66639
08/31/2026, 7:15 PMbrief-scientist-13682
08/31/2026, 7:16 PM:; git grep LAST_ACCESS_FILE
pex/cache/access.py:LAST_ACCESS_FILE = ".last-access"
pex/cache/access.py: return os.path.join(pex_dir.path, LAST_ACCESS_FILE)
pex/venv/installer.py: cache_access.LAST_ACCESS_FILE,
:; git blame pex/cache/access.py | grep LAST_ACCESS_FILE
991883c71 (John Sirois 2024-11-04 16:03:58 -0800 101) LAST_ACCESS_FILE = ".last-access"
991883c71 (John Sirois 2024-11-04 16:03:58 -0800 106) return os.path.join(pex_dir.path, LAST_ACCESS_FILE)
:; git blame -- pex/venv/installer.py | grep LAST_ACCESS_FILE
991883c71 pex/venv/installer.py (John Sirois 2024-11-04 16:03:58 -0800 582)echoing-psychiatrist-92346
08/31/2026, 7:21 PMFile "/bin/app/pex", line 336, in boot
with open(os.path.join(os.path.dirname(__file__), ".last-access"), "a") as fp:
OSError: [Errno 30] Read-only file system: '/bin/app/.last-access'
which I linked to pex 2.98.3 which added an unconditional last-access write to the generated --venv boot script with no try/except:
github.com/pex-tool/pex/pull/3226
That collides with two things on our side:
• we materialize the venv at build time into /bin/app, i.e. on the container root filesystem, not under PEX_ROOT
• appfile-cli (custom tool which generates our kube manifests) sets ReadOnlyRootFilesystem: true unconditionally for every rendered container, with no appfile knob to opt out
So the boot script tries to create /bin/app/.last-access on a read-only filesystem and dies before reaching the entrypoint.
Worth noting the write buys us nothing anyway: it exists to feed pex3 cache prune, which only ever manages venvs inside PEX_ROOT. A build-time venv at /bin/app is never pruned. So I set _PEX_CACHE_SET_LAST_ACCESS=0 in all container env which skips the write and app boots clean. I was also able to reproduced the crash and the fix locally.echoing-psychiatrist-92346
08/31/2026, 7:24 PMbrief-scientist-13682
08/31/2026, 7:24 PMwe materialize the venv at build time intoVia what?/bin/app
PEX_TOOLS=1 /bin/app venv ... If so then the .last-access touch is a bug.brief-scientist-13682
08/31/2026, 7:27 PMbrief-scientist-13682
08/31/2026, 7:45 PMbrief-scientist-13682
08/31/2026, 8:52 PMWaiting for the actual cache change from 2.101.2 to propagate to check if it fixed the issues we've been seeing the last couple days or not
I think it's clear, but in case not: I don't think this will fix anything unless there are wonky perms / ownership which you all have ruled out. I only expect more diagnostics in the logs.
echoing-psychiatrist-92346
08/31/2026, 9:01 PMbrief-scientist-13682
08/31/2026, 10:42 PM.last-access fix is available here: https://github.com/pex-tool/pex/releases/tag/v2.101.3brief-scientist-13682
09/01/2026, 5:47 PMrefined-lifeguard-66639
09/01/2026, 5:48 PMechoing-psychiatrist-92346
09/01/2026, 7:24 PMEEXIST on .lck.work, then "Continuing to forcibly re-create the work directory" all across several different downloads/ cache keys. In every case that's the only pair that appears; neither the "Failed to forcibly re-create" nor "Using new random workdir instead" follow-ups fired, meaning rmtree + mkdir succeeded cleanly on the first retry. Self-heal worked as intended in the runs that ultimately passed. Where the run still failed overall despite this pattern, the log tail cuts off before showing the actual terminal error, so I can't yet say whether that failure is downstream of this same event or unrelated
Exact warnings, representative of this group:
[WARN] .../pex_root/installed_wheels/.../pex-2.101.2-py2.py3-none-any.whl/pex/atomic_directory.py:290:
PEXWarning: [pid:<pid>, tid:<tid>, cwd:/tmp/pants-sandbox-<x>]: After obtaining an exclusive lock on
.../pex_root/downloads/2/.<hash-a>.atomic_directory.lck,
failed to establish a work directory at
.../pex_root/downloads/2/<hash-a>.lck.work due to:
[Errno 17] File exists: '.../downloads/2/<hash-a>.lck.work'
.../pex/atomic_directory.py:301: PEXWarning: [pid:<pid>, tid:<tid>, cwd:/tmp/pants-sandbox-<x>]:
Continuing to forcibly re-create the work directory at
.../pex_root/downloads/2/<hash-a>.lck.work.
.../pex/atomic_directory.py:290: PEXWarning: [pid:<pid>, tid:<tid2>, cwd:/tmp/pants-sandbox-<x>]: After
obtaining an exclusive lock on
.../pex_root/downloads/2/.<hash-b>.atomic_directory.lck,
failed to establish a work directory at
.../pex_root/downloads/2/<hash-b>.lck.work due to:
[Errno 17] File exists: '.../downloads/2/<hash-b>.lck.work'
.../pex/atomic_directory.py:301: PEXWarning: [pid:<pid>, tid:<tid2>, cwd:/tmp/pants-sandbox-<x>]:
Continuing to forcibly re-create the work directory at
.../pex_root/downloads/2/<hash-b>.lck.work.
Same pid across both keys within a run, different tids which is consistent with your abnormal-termination read rather than a live lock race, same pattern you'd already established. Every run in this group shows the identical two-line shape, just different pids and cache keys, nothing new, just repeated independent occurrences.
The other failed runs in this batch show a different pattern entirely: the exact symptom class from your earlier thread which was ModuleNotFoundError: No module named 'pex.version' on a .bootstrap zip during venv --scope=deps inside the Docker build but grepping the entire job log (and its sibling arch build in the same run) for `PEXWarning`/`atomic_directory.py` turns up zero matches. No EEXIST, no forcibly-recreate, nothing. CI's own retry-with-scratch-PEX_ROOT logic fired more than once per run and the corruption reproduced identically every time.
Exact error, representative of this group:
#9 [deps 3/3] RUN PEX_TOOLS=1 /usr/local/bin/python3 /binary-deps.pex venv --scope=deps --compile /bin/app
#9 0.246 Traceback (most recent call last):
#9 0.246 File "<frozen runpy>", line 198, in _run_module_as_main
#9 0.246 File "<frozen runpy>", line 88, in _run_code
#9 0.246 File "/binary-deps.pex/__main__.py", line 280, in <module>
#9 0.246 result, should_exit, is_globals = boot(
#9 0.246 ^^^^^
#9 0.246 File "/binary-deps.pex/__main__.py", line 247, in boot
#9 0.246 from pex.variables import ENV, Variables
#9 0.246 File "/binary-deps.pex/.bootstrap/pex/variables.py", line 18, in <module>
#9 0.246 File "/binary-deps.pex/.bootstrap/pex/common.py", line 26, in <module>
#9 0.246 File "/binary-deps.pex/.bootstrap/pex/enum.py", line 13, in <module>
#9 0.246 'global_flag_repr', 'global_enum_repr', 'global_str', 'global_enum',
#9 0.246 ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
#9 0.246 File "/binary-deps.pex/.bootstrap/pex/exceptions.py", line 12, in <module>
#9 0.246 ModuleNotFoundError: No module named 'pex.version'
#9 ERROR: process "/bin/sh -c PEX_TOOLS=1 /usr/local/bin/python3 /binary-deps.pex venv --scope=deps --compile /bin/app"
did not complete successfully: exit code: 1
A .bootstrap chroot missing pex/version.py in every run in this group, same failure signature as the earlier builds you and I discussed, reproduced across multiple retries against a fresh scratch PEX_ROOT before the run gave up.
That's exactly the shape your "second mechanism" concern predicts: a corrupted artifact with the identical downstream symptom, but no trace in the loud atomic_directory EEXIST path #3263 instruments which is consistent with the pep_427.py:1119-1131 skip-on-exists copies (safe_copy(overwrite=False), elif not os.path.exists(dst_file), the symlink-EEXIST swallow) quietly preserving stale content into a populate that never trips `os.mkdir`'s EEXIST check in the first place, so there's nothing for #3263 to warn about.refined-lifeguard-66639
09/01/2026, 7:32 PMechoing-psychiatrist-92346
09/01/2026, 7:44 PM20:07:14.84 [INFO] Canceled: Building 79 requirements for .../migration/alembic_pex.pex from the
3rdparty/python/deps_lock.txt resolve: PyMySQL, aiohttp, alembic, anthropic, ... (18.3s)
20:07:14.84 [WARN] .../pex/atomic_directory.py:290: PEXWarning: [pid:<pid>, ...]: After obtaining
an exclusive lock on ..., failed to establish a work directory at ... due to: [Errno 17] File exists
Pants is killing one pex build which lines up exactly with another pex process tripping over a leftover .lck.work dir. Also Pants canceled a pex build very shortly into building it, and that's the exact pex that shows up corrupted later in the same run.refined-lifeguard-66639
09/01/2026, 7:47 PMbrief-scientist-13682
09/01/2026, 7:49 PMechoing-psychiatrist-92346
09/01/2026, 7:49 PMbrief-scientist-13682
09/01/2026, 7:49 PMbrief-scientist-13682
09/01/2026, 7:50 PMechoing-psychiatrist-92346
09/01/2026, 7:52 PMbrief-scientist-13682
09/01/2026, 7:53 PMrefined-lifeguard-66639
09/01/2026, 7:54 PMrefined-lifeguard-66639
09/01/2026, 7:54 PMbrief-scientist-13682
09/01/2026, 7:54 PMrefined-lifeguard-66639
09/01/2026, 7:54 PMbrief-scientist-13682
09/01/2026, 7:54 PMrefined-lifeguard-66639
09/01/2026, 7:55 PMbrief-scientist-13682
09/01/2026, 7:55 PMbrief-scientist-13682
09/01/2026, 7:56 PMbrief-scientist-13682
09/01/2026, 7:56 PMbrief-scientist-13682
09/01/2026, 8:00 PMbrief-scientist-13682
09/01/2026, 8:01 PMechoing-psychiatrist-92346
09/02/2026, 9:04 PMos.walk failing silently.
.deps/mypy-1.8.0-…whl blew up but is a perfectly valid zip with 2 members: .layout.json plus one .so where the healthy packed wheel has 1300, including 81 directory entries. The poisoned one has zero, i.e. the walk never descended past the chroot root. No dist-info/METADATA`, which is what PEX_TOOLS=1 … venv chokes on.
deterministic_walk (pex/common.py) wraps os.walk and no call site passes onerror, so a failed directory listing gets discarded and that subtree is silently omitted then the walk then finishes "successfully" with a strict subset of the tree. PEXBuilder._add_dist registers only the chroot's root-level files, _build_packedapp zips that subset, and atomic_directory takes its success path and finalizes it. _add_dist returns the fingerprint of the complete chroot, not a hash of what it actually collected so the truncated wheel gets cached under a key asserting it's whole.
The fix is small github.com/pex-tool/pex/pull/3269
default onerror to re-raise so we fail at the listing instead of shipping a broken PEX. Reproduced it against pex's own builder and added regression tests.
tbh it is still slightly unknown what makes these fail on our build agents, my best guess is fd exhaustion from the parallel docker builds or ENOENT from concurrent cache mutation.brief-scientist-13682
09/02/2026, 10:25 PMuv run dev-cmd package and report back!echoing-psychiatrist-92346
09/03/2026, 5:04 PMbrief-scientist-13682
09/03/2026, 5:07 PMechoing-psychiatrist-92346
09/03/2026, 5:39 PMpants package on the pex build now instead of 20 min later at the docker build stage. I'm only sampling our main pipeline so far and haven't caught one there yet. Pulling data across all branches/pipelines later today for a real sample.
I put in a guard yesterday as well to detect and evict incomplete cache entries preflight which has been helping as well to clean out the poison, but this should make it so no incomplete packages get in in the first placebrief-scientist-13682
09/03/2026, 10:12 PMos.walk fail fast fix is now available here: https://github.com/pex-tool/pex/releases/tag/v2.101.4brief-scientist-13682
09/04/2026, 6:44 PMrefined-lifeguard-66639
09/08/2026, 4:32 PMbrief-scientist-13682
09/08/2026, 4:36 PMrefined-lifeguard-66639
09/08/2026, 4:39 PMechoing-psychiatrist-92346
09/08/2026, 6:12 PMbrief-scientist-13682
09/08/2026, 6:16 PMechoing-psychiatrist-92346
09/08/2026, 6:21 PMbrief-scientist-13682
09/08/2026, 6:24 PMbrief-scientist-13682
09/08/2026, 6:25 PM