<@U07QKSQKB7S> :wave: hello I don't think aiden is...
# general
r
@brief-scientist-13682 👋 hello I don't think aiden is in here; but if you want to talk about github.com/pex-tool/pex/pull/3262 and what we're seeing I'm around
b
If you are using the uv resolver, then this is likely github.com/pantsbuild/pants/issues/23657 I'd like to verify or deny that 1st.
Ok - you say no uv in ticket. So are you using network mounts in CI (e.g. NFS or similar)?
r
Nope EBS and instance-store only
I had Fable try to build a repro last night; I've not been able to look through what it did before Aiden filed the patch
b
Ok, well I'll need more details likely, but the fix in the PR is patchy - just addressed
--layout packed
and not loose or zipapp, and getting to the bottom of the corruption is the preferred route.
r
Yeah for sure
b
So Pex will fail in environments where flock doesn't work (NFS angle), and it will fail if fed bad input wheels, or if it has external party delete from its cache. But those are the only known cases over many years.
r
Yeah that makes absolute sense. AFAICT neither of those is the case that we're seeing here.
b
IIUC your company is ~long time Pants user; so what changed?
That is clearly useful missing context.
Maybe?: github.com/pex-tool/pex/…/v2.97.3 Pants 2.33.0 is on 2.97.1
Worth a try anyway, you upgrade with, for example:
Copy code
[pex-cli]
version = "v2.101.1"
known_versions = [
    "v2.101.1|macos_x86_64|1fc9c277152e450ab5de2728eac6130b2433d8e94c491d9eaeeae63433a25c15|5314994",
    "v2.101.1|macos_arm64|1fc9c277152e450ab5de2728eac6130b2433d8e94c491d9eaeeae63433a25c15|5314994",
    "v2.101.1|linux_x86_64|1fc9c277152e450ab5de2728eac6130b2433d8e94c491d9eaeeae63433a25c15|5314994",
    "v2.101.1|linux_arm64|1fc9c277152e450ab5de2728eac6130b2433d8e94c491d9eaeeae63433a25c15|5314994"
]
(That's in pants.toml)
r
when we upgraded we moved to pex 2.97.3
b
That's fishy. Why and why that version in particular.
r
Latest patch of the version that pants wanted to be on that was all.
b
Ok, so just lucky timing.
r
Yeah we upgraded on the 20th IIRC
b
Ok, well, then there is no known bug that would cause Pex to write bad atomic directories; so the existing PR just patches over the unknown cause and only does it for packed. Those are both problems. I'm trying to think what I might ask you to provide that could help me debug.
Do I have to be thinking about Pants remote cache usage or not? Do you use that?
r
Yes we do. I have reason to believe what brought this to our attention was that I fixed our cache's deployment to be much more stable (ie not crashlooping). This then meant that builds were actually able to read/write consistently so when the bad pex got built and pushed to the cache it was then distributed to other builds.
b
Ok. So that I know the full code path here, these are useful to answer: + What
pip_version
do you set in pants.toml if any? + What
interpreter_constraints
are in-play? + Do you enable resolves (locks)? Basically, the more of pants.toml I can see the better.
r
Roger that. 👀
pant.toml [pex-cli]
Copy code
[pex-cli]
known_versions = [
  "v2.97.3|macos_x86_64|61bc0afdb879b791c30211a809fddc542e36f6b4533240c85fb56753f6d65c05|5038874",
  "v2.97.3|macos_arm64|61bc0afdb879b791c30211a809fddc542e36f6b4533240c85fb56753f6d65c05|5038874",
  "v2.97.3|linux_x86_64|61bc0afdb879b791c30211a809fddc542e36f6b4533240c85fb56753f6d65c05|5038874",
  "v2.97.3|linux_arm64|61bc0afdb879b791c30211a809fddc542e36f6b4533240c85fb56753f6d65c05|5038874",
]
version = "v2.97.3"
global_args = ["--tmpdir=/tmp"]
pants.ci.toml [pex-cli]
The --tmpdir stuff is because of the same issue George-Cristian reported with sockets on macos.
Copy code
[pex-cli]
global_args = []
b
Ok, and I typo-ed and corrected above - sorry - its pip_version I need - you already let me know Pex was 2.97.3
r
Ah 👀
b
And no ICs?
r
pants.toml [python]
Copy code
[python]
enable_resolves = true
resolves_generate_lockfiles = true

interpreter_constraints = [">=3.10,<3.13"]
tailor_requirements_targets = false
looking @ pip
b
In CI, do you know which is the real Python version used? If you do not set
pip_version
the version used depends on the Python version selected by Pex from the range.
r
Ok yeah we don't set it. It should be py 3.10 😭
b
Ok, I'll assume 3.10 which means Pex picks Pip 26.1.
Last question before I do re-analysis of Pex bug paths writing cache: Do you enable the sandboxer?
r
yes
wait*
b
r
Pants sandbox right?
b
The linked option above.
r
Ok no we don't. Thought i have seen
pants-sandbox
dirs around. Which is why i initially said yes
b
Yeah - there is a name overload there. Ok, great - thanks for the info. I'll spend a few hours staring at the code to see how there could be a write bug and report back.
r
One thing I was looking at is Pant's request hedging and cancellation of pex procs
And if somehow that could leave state behind that would get picked up by a future pex invocation and treated as good to proceed with. But I've not had much time to validate anything there.
b
That has existed for a very long time and should not be new. What was your Pants version upgrade from?
You never gave me full what changed context.
r
Pants went from 2.31.0 -> 2.33.0. We wanted some of the perf changes as well as the opentelemetry trace support.
• On aug 17th a teammate bumped pex-cli for pants to
2.97.1
from
2.81.0
to prepare for the upgrade of pants • On aug 18th I moved pants' cache dirs on our build nodes to be outside the dockerized build environment so that future builds could potentially cache-hit on local artifacts. • On the 19th I fixed our bazel-remote-cache servers to be properly provisioned / have a functioning loadbalancer • On aug 20th the same teammate then bumped pants to
2.33.0
from
2.31.0
and raised pex to
2.97.3
• On the 26th that
--tmpdir
change for pex was merged but without the override in the
pants.ci.toml
. We thought that caused some of these issues that cropped up so we reverted & re-implemented with the
pants.ci.toml
global args override of
[]
• In various attempts to try to limit the scope of the build issue we see with these malformed artifacts we've moved most build state out of shared directories with the host and back into the container. None of that worked. • Last night we finally were able to find a build that produced such a malformed artifact and persisted it into the remote-cache rather than consuming one. This helped us understand that we're still seeing the issue and prompted in part Aiden's proposed patch.
More digging through build logs today has shown that we've seen this kind of malformed artifact as far back as April (when we were on pants v2.28.0). But i'm less sure of the setup from back then as I started in Aug.
It is my current belief that the attempts to achieve more cache hits made this problem more prominent for us because the pex step succeeds and ends up being cached. By sharing the cache on build nodes, and repairing the remote-cache more than just the build that produced the artifact are being affected.
b
Last night we finally were able to find a build that produced such a malformed artifact and persisted it into the remote-cache rather than consuming one.
Good datapoint. How do you know the build produced the artifact? This was a build that had 0 local pants caches?
r
Correct it cache missed on the build step for this - I will try to get you as many specifics as I'm able
b
A remote cache miss can pull from local cache; i.e. ~/.cache/pants/named-caches which Pex uses
So I just want to confirm 0 caches.
r
Yeah we've disabled local caches entirely at this point and have just remote going.
Though it is my understanding that
named-caches
is independent of local cache
b
Yeah we've disabled local caches entirely
What is the config that does that?
Yes - exactly
r
Ok cool - yeah we've disabled pants' local cache opt. I do not think we've disabled pex's
As there is no option for the named caches other than to route it to a specific dir.
b
Ok. So Pex may or may not be pulling from local cache. That would have to be a volume mount into the container though.
So - are named caches currently shared in from host and does the host have them persisted between CI runs.
r
I have just discovered as we're speaking here that it is in fact shared from the host 😭
Apparently buildkite mounts the checked out repo from the host so the default dir of
buildroot/.pants.d/named-caches
is not in fact ephemeral like we thought it was.
b
Ok, and what type of sharing is this linux host, linux container, windows host linux container - and what is the the volume mount technology being used?
r
Linux bind mount -> docker
b
Sorry - fs driver for docker
r
no worries we're both learning some things about the here 🙃
b
I think the docker fs driver story has settled, but in the early days there were ones you could pick IIRC even on linux that would not respect locks
🤝 1
At any rate, now that the local cache hole is discovered, a run with that wiped would be good
It may have pre-Pex 2.97.3 artifacts IIUC.
r
It def won't have pre-pex2.97 artifacts because we've fully rolled the fleet of nodes
b
Ok.
r
But there might be other shenanigans like a build getting oomkilled or something weird
b
Well, Pex should be resilient to that.
The atomic_directory works on a tmp dir only moved into final position after the directory populating task completes without raising.
That mv is an atomic fs rename
r
I was able to repro this failure by having ~30 pex procs in parallel building the same pex and killing ~1/3 of them. But I don't trust that this is representative of what actually happens to us
b
Please provide the script!
I'd love a repro to work from - that's great.
r
well when I say I - i mean a robot that I haven't validated the work of; didn't want to pass you something I didn't personally review.
b
I hate AI. Ok, but what you're relaying is what it claimed?
🤝 1
I can work with that if its all I've got.
r
Ok, but what you're relaying is what it claimed?
Yes, it had something that it managed to repro with. Once I get you the docker fs question answer I'll actually review what it has.
b
Ok, great. I'll work on reproing the repro.
🙌 2
🙏 1
r
We're using overlay2. The builds are running on
m8gd.4xlarge
and
m8id.4xlarge
hosts in AWS. All the disk at play here should be disk from the dedicated NVME side of those nodes (ofc they still have a root EBS as well)
b
If the open issue is in play for a lock, and the lock is on the read-only layer and gets copied up, the fd changes which breaks flock IIUC, but not sure yet.
r
More changes to the story here, sorry. It appears all of our apps that have run into this are not using our repo's configured python 3.10 but are instead on 3.12 Which changes the pip pex end up using right?
b
I think that will also select 26.1. Thanks for the info though. I will note you should be selecting the latest Pip supported though instead of relying on defaulting. Pip has been getting love in the last 1.5 years and is getting faster.
r
Roger that - given my core mission of build speed that's super good intel 🙂
b
So that I'm clear - still reading up on overlayfs2 - a single host named caches is shared to 1 container or multiple concurrent containers?
r
1 container at a time.
But how many pex cmds pants runs in the container is up to pants right?
b
Yes. Yeah Pants / Pex parallelism on a normal host has been hammered for a long time. I'm just probing the point of the flock'ing. If the flock is questionable all bets are off.
🤝 1
r
OK I gotta go do a few weekend chores and such. But i'll be back in a few hours / be on mobile.
b
I would love that AI claimed repro. I can't repro it yet. Using:
Copy code
#!/usr/bin/env bash

set -euo pipefail

COUNT=30
MOD=3

export PEX_ROOT="${PEX_ROOT:-/tmp/pex-root}"

pex3 lock create --pip-version latest-compatible ansible --indent 2 -o /tmp/lock.json
rm -rf "${PEX_ROOT}"

declare -a kill_pex_pids
declare -a wait_pex_pids
for ((x=1; x<COUNT; x++)); do
    pex ansible -c ansible --lock /tmp/lock.json --venv prepend --layout packed -- --version &
    pid=$!
    if (( x % MOD == 0 )); then
        kill_pex_pids+=($pid)
    else
        wait_pex_pids+=($pid)
    fi
done

for pid in "${kill_pex_pids[@]}"; do
    sleep 2
    kill ${pid}
done

for pid in "${wait_pex_pids[@]}"; do
    wait ${pid}
done
I get things like this with exit code 0:
Copy code
/home/jsirois/bin/tools.venv/lib/python3.14/site-packages/pex/atomic_directory.py:268: PEXWarning: [pid:60339, tid:131625590040256, cwd:/home/jsirois/support/pex/slack/8-29-2026]: After obtaining an exclusive lock on /tmp/pex-root/downloads/2/.fb06b66c8da04172d9e72a21d7d06186d8919e32ae5ab5cdf5b9d920be805ac2.atomic_directory.lck, failed to establish a work directory at /tmp/pex-root/downloads/2/fb06b66c8da04172d9e72a21d7d06186d8919e32ae5ab5cdf5b9d920be805ac2.lck.work due to: [Errno 17] File exists: '/tmp/pex-root/downloads/2/fb06b66c8da04172d9e72a21d7d06186d8919e32ae5ab5cdf5b9d920be805ac2.lck.work'
  pex_warnings.warn(
/home/jsirois/bin/tools.venv/lib/python3.14/site-packages/pex/atomic_directory.py:268: PEXWarning: [pid:60347, tid:126699643143872, cwd:/home/jsirois/support/pex/slack/8-29-2026]: After obtaining an exclusive lock on /tmp/pex-root/downloads/2/.9e7dd367f7dc5d5e9fc5ae1baf8af9c4edc09e916a73a40108a3f32e3ad93f10.atomic_directory.lck, failed to establish a work directory at /tmp/pex-root/downloads/2/9e7dd367f7dc5d5e9fc5ae1baf8af9c4edc09e916a73a40108a3f32e3ad93f10.lck.work due to: [Errno 17] File exists: '/tmp/pex-root/downloads/2/9e7dd367f7dc5d5e9fc5ae1baf8af9c4edc09e916a73a40108a3f32e3ad93f10.lck.work'
  pex_warnings.warn(
/home/jsirois/bin/tools.venv/lib/python3.14/site-packages/pex/atomic_directory.py:279: PEXWarning: [pid:60339, tid:131625590040256, cwd:/home/jsirois/support/pex/slack/8-29-2026]: Continuing to forcibly re-create the work directory at /tmp/pex-root/downloads/2/fb06b66c8da04172d9e72a21d7d06186d8919e32ae5ab5cdf5b9d920be805ac2.lck.work.
  pex_warnings.warn(
/home/jsirois/bin/tools.venv/lib/python3.14/site-packages/pex/atomic_directory.py:279: PEXWarning: [pid:60347, tid:126699643143872, cwd:/home/jsirois/support/pex/slack/8-29-2026]: Continuing to forcibly re-create the work directory at /tmp/pex-root/downloads/2/9e7dd367f7dc5d5e9fc5ae1baf8af9c4edc09e916a73a40108a3f32e3ad93f10.lck.work.
  pex_warnings.warn(
/home/jsirois/.pyenv/versions/3.14.7/lib/python3.14/multiprocessing/resource_tracker.py:475: UserWarning: resource_tracker: There appear to be 2 leaked semaphore objects to clean up at shutdown: {'/mp-eleeewhf', '/mp-inn89lbx'}
  warnings.warn(
/home/jsirois/.pyenv/versions/3.14.7/lib/python3.14/multiprocessing/resource_tracker.py:475: UserWarning: resource_tracker: There appear to be 2 leaked semaphore objects to clean up at shutdown: {'/mp-r2k9mixn', '/mp-fv5ir0r1'}
  warnings.warn(
/home/jsirois/.pyenv/versions/3.14.7/lib/python3.14/multiprocessing/resource_tracker.py:475: UserWarning: resource_tracker: There appear to be 2 leaked semaphore objects to clean up at shutdown: {'/mp-qg5m83q7', '/mp-wtniolvk'}
  warnings.warn(
/home/jsirois/.pyenv/versions/3.14.7/lib/python3.14/multiprocessing/resource_tracker.py:475: UserWarning: resource_tracker: There appear to be 2 leaked semaphore objects to clean up at shutdown: {'/mp-7oaiujg4', '/mp-yawzw6h2'}
  warnings.warn(
/home/jsirois/.pyenv/versions/3.14.7/lib/python3.14/multiprocessing/resource_tracker.py:475: UserWarning: resource_tracker: There appear to be 6 leaked semaphore objects to clean up at shutdown: {'/mp-3elo32t7', '/mp-jb8j0soh', '/mp-__qceg7m', '/mp-19byohln', '/mp-7xca8xo8', '/mp-u8yvl5h8'}
  warnings.warn(
Process ForkServerPoolWorker-1:
Traceback (most recent call last):
  File "/home/jsirois/.pyenv/versions/3.14.7/lib/python3.14/multiprocessing/pool.py", line 131, in worker
    put((job, i, result))
    ~~~^^^^^^^^^^^^^^^^^^
  File "/home/jsirois/.pyenv/versions/3.14.7/lib/python3.14/multiprocessing/queues.py", line 397, in put
    self._writer.send_bytes(obj)
    ~~~~~~~~~~~~~~~~~~~~~~~^^^^^
  File "/home/jsirois/.pyenv/versions/3.14.7/lib/python3.14/multiprocessing/connection.py", line 210, in send_bytes
    self._send_bytes(m[offset:offset + size])
    ~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/home/jsirois/.pyenv/versions/3.14.7/lib/python3.14/multiprocessing/connection.py", line 441, in _send_bytes
    self._send(header)
    ~~~~~~~~~~^^^^^^^^
  File "/home/jsirois/.pyenv/versions/3.14.7/lib/python3.14/multiprocessing/connection.py", line 404, in _send
    n = write(self._handle, buf)
BrokenPipeError: [Errno 32] Broken pipe
...
ansible [core 2.21.3]
  config file = None
  configured module search path = ['/home/jsirois/.ansible/plugins/modules', '/usr/share/ansible/plugins/modules']
  ansible python module location = /tmp/pex-root/venvs/3/0006b71e2cfba8fe25d1d711b7e930f1aec28f6f/c40de391655dede9f50dced114f8a1dc8f21dd91/lib/python3.14/site-packages/ansible
  ansible collection location = /home/jsirois/.ansible/collections:/usr/share/ansible/collections
  executable location = /tmp/pex-root/venvs/3/0006b71e2cfba8fe25d1d711b7e930f1aec28f6f/c40de391655dede9f50dced114f8a1dc8f21dd91/pex
  python version = 3.14.7 (main, Aug  5 2026, 17:40:48) [GCC 15.2.0] (/tmp/pex-root/venvs/3/0006b71e2cfba8fe25d1d711b7e930f1aec28f6f/c40de391655dede9f50dced114f8a1dc8f21dd91/bin/python)
  jinja version = 3.1.6
  pyyaml version = 6.0.3 (with libyaml v0.2.5)ansible [core 2.21.3]
  config file = None
  configured module search path = ['/home/jsirois/.ansible/plugins/modules', '/usr/share/ansible/plugins/modules']
  ansible python module location = /tmp/pex-root/venvs/3/0006b71e2cfba8fe25d1d711b7e930f1aec28f6f/c40de391655dede9f50dced114f8a1dc8f21dd91/lib/python3.14/site-packages/ansible
  ansible collection location = /home/jsirois/.ansible/collections:/usr/share/ansible/collections
  executable location = /tmp/pex-root/venvs/3/0006b71e2cfba8fe25d1d711b7e930f1aec28f6f/c40de391655dede9f50dced114f8a1dc8f21dd91/pex
  python version = 3.14.7 (main, Aug  5 2026, 17:40:48) [GCC 15.2.0] (/tmp/pex-root/venvs/3/0006b71e2cfba8fe25d1d711b7e930f1aec28f6f/c40de391655dede9f50dced114f8a1dc8f21dd91/bin/python)
  jinja version = 3.1.6
  pyyaml version = 6.0.3 (with libyaml v0.2.5)
ansible [core 2.21.3]
  config file = None
  configured module search path = ['/home/jsirois/.ansible/plugins/modules', '/usr/share/ansible/plugins/modules']
  ansible python module location = /tmp/pex-root/venvs/3/0006b71e2cfba8fe25d1d711b7e930f1aec28f6f/c40de391655dede9f50dced114f8a1dc8f21dd91/lib/python3.14/site-packages/ansible
  ansible collection location = /home/jsirois/.ansible/collections:/usr/share/ansible/collections
  executable location = /tmp/pex-root/venvs/3/0006b71e2cfba8fe25d1d711b7e930f1aec28f6f/c40de391655dede9f50dced114f8a1dc8f21dd91/pex
  python version = 3.14.7 (main, Aug  5 2026, 17:40:48) [GCC 15.2.0] (/tmp/pex-root/venvs/3/0006b71e2cfba8fe25d1d711b7e930f1aec28f6f/c40de391655dede9f50dced114f8a1dc8f21dd91/bin/python)
  jinja version = 3.1.6
  pyyaml version = 6.0.3 (with libyaml v0.2.5)
...
N.B.: This has the
failed to establish a work directory at ... Continuing to forcibly re-create the work directory at
logging @echoing-psychiatrist-92346 reported in the PR due to the kills as well as multiprocessing complaining about kills, but all the non-killed Pex processes run to successful completion.
This was on my Linux machine - no Docker. Next up is to try that repro but with PEX_ROOT mounted into docker container using overlayfs2...
r
😞 Mac Docker Desktop uses FUSE to mount local APFS directories into containers - the repro was on that FUSE fs. Not in an ext4 bindmount, or the overlayfs2. Since this is not a repro of what we're seeing in production I will keep hammering on this and come back when I've got a useable repro for you.
b
Ok. Thank you. My docker wrapper repro did no repro, but I'm Linux / Linux:
Copy code
#!/usr/bin/env bash

set -euo pipefail

CONTAINER_CONCURRENCY=5

docker build -t pex-cache-write-fail-repro .

mkdir -p $PWD/pex-root
sudo find $PWD/pex-root/ -type f -name "*.lck" | while read lock_file; do
    # pex-root/downloads/2/.b727414169a36b7d524c1c3e31839a521725078d7b2ff038656844266160a992.atomic_directory.lck
    base_dir="$(dirname "${lock_file}")"
    lock_file_name="$(basename "${lock_file}")"
    target_dir_name="$(echo "${lock_file_name}" | cut -d. -f2)"
    sudo rm -rf "${base_dir}/${target_dir_name}"
done

declare -a pids
for ((x=0; x<CONTAINER_CONCURRENCY; x++)); do
    docker run --rm -v $PWD/pex-root:/tmp/pex-root pex-cache-write-fail-repro:latest repro.sh &
    pids+=($!)
done

for pid in "${pids[@]}"; do
    wait $pid
done
r
Yeah I mean that's our build machine setup - which is more annoying to repro when the dev machine is a mac.
b
I will beat my dead horse - its a brain sickness that devs who work on prod Linux apps are given macs.
🤝 1
r
I may show up to work on Monday and just walk over to IT and requisition something I can run ubuntu on; at least to repro this.
Having to do all of it via SSH on remote nodes (that get cleaned up) is such a pain
b
Yup! I used to wipe my IT supplied mac and install Linux, but now that intel is gone and M2 is old - no longer an option.
r
Yeah I loved my 2014 macbook air (with linux on it)
old man yells at cloud
b
Make some noise! Get those macs recycled! Yeah - unfortunately.
r
OK it has reproduced itself in prod and since we've enabled the otel tracing of pants i now have the raw pex command that pants ran
b
Ok. If you can provide that the info can't hurt. Presumably the command runs fine in isolation though.
r
Yeah just removing some internal stuff before I post it here publicly
Copy code
[
  "/pants-named-caches/python_build_standalone/aacb154e49a3600820e795704db17f7177cf27c4e08fc7b159882144f27f022d/bin/python3",
  "./pex",
  "--tmpdir",
  ".tmp",
  "--jobs",
  "{pants_concurrency}",
  "--no-emit-warnings",
  "--pip-version",
  "24.2",
  "--python-path",
  "/usr/local/bin/python3",
  "--output-file",
  "python.klaviyo.redacted.server/app-deps.pex",
  "--emit-warnings",
  "--check=warn",
  "--venv",
  "prepend",
  "--include-tools",
  "--runtime-pex-root=/tmp/pex-redacted-server-app-deps",
  "--requirements-pex",
  "local_dists.pex",
  "--interpreter-constraint",
  "CPython<3.13,==3.12.*,>=3.10",
  "--python-path",
  "/usr/local/bin/python3.12",
  "--entry-point",
  "redacted.server.main:main",
  "--sources-directory=source_files",

  "REDACTED many deps, left grpcio cause that's the one that failed",
  "grpcio",

  "--lock",
  "3rdparty/python/deps_lock.txt",
  "--no-pypi",
  "--index=<https://pypi.org/simple/>",
  "--index=<https://REDACTED>",
  "--manylinux",
  "manylinux2014",
  "--layout",
  "packed"
]
I think relevantly it looks like pants is explicitly setting pip 24.2
which is not a version of pip we talked about before
b
Ok. This is good to see but doesn't shed light.
r
Since we're using the lockfile we should never be able to download and then use a bad wheel right? Cause pex/pex via pip will checksum that against the sha in the lockfile?
b
Correct
So, one old bug in Pants: it used to kill processes, not wait for them to be dead; then tear down sandbox in parallel to the still not dead process.
That led to corrupt partial sandbox getting persisted in caches by still running process.
I raise this to get you thinking about named caches managed outside container and mounted in. You aren't doing something similar? Kill container, don't wait for result, clean host named caches mount?
👀 1
For example.
r
The named caches are currently being written to a mount that's shared with the host. however there is only 1 build container running on the host at a time right now. However that build container's pants execution may run many pex-clis at once.
b
As does a normal Pants. That bit is not novel.
A mount that is shared with host or on host and shared with container?
r
sorry the latter, we mount in the buildkite 'work' directory; presently the named caches are there. in .pants.d/named_caches
Normal ext4/bindmount for that AFAIK
b
I'm not sure what to say. I sort of need a repro in hand. Code inspection + thinking through your setup seems to reveal no crease for write corruption.
r
Yeah it's late here. I think i'm going to put this down for now and sleep on it. We may try to run with some variant of Aidan's patch to get a clearer-error message when the corruption occurs to make it easier to catch 'in action'. If we did that is there anything in particular that would be useful to you to have? (since it seems to be many hours between occurrences in prod)
b
The read repair side can't provide useful info. The damage is done.
His patch effectively just hides the problem.
That's why I'm leery of patches like that.
r
I was thinking to run that verification on the produced pex at build time
b
A patch like that would have hid that old Pants bug for example.
r
ie build it; then try to re-read it
making it possible for the 'build' step to know if the output was 👍 or 👎
b
But that still doesn't tell you what caused the bad write when you get 👎
r
no - but part of my problem there is that I have yet to be able to get on an actual host where the bad build happened
So having some stronger signal would be helpful in that regard. but if there isn't anything off the top of your head you'd want to see that's ok
It would also prevent the 👎 write (that exited 0) from going into our cache and then failing future builds - while allowing us to continue to hunt down the problem. But that's more of an us concern in terms of maintaining a build pipeline that flows smoothly.
b
Gotcha
Nope EBS and instance-store only
@refined-lifeguard-66639 just checking: not EBS Multi-Attach I presume https://docs.aws.amazon.com/ebs/latest/userguide/ebs-volumes-multi.html#considerations
e
(Confirming for Jonathan since I'm working on this with him) - the instances do not use multi-attach, I checked the ASG instances with `aws ec2 describe-volumes`:
Copy code
"Iops": 6000,
            "VolumeType": "gp3",
            "MultiAttachEnabled": false,
🤝 1
b
Ok. Thanks for checking @echoing-xylophone-39473.
e
Thanks for creating the
v2.101.2
Pex update, we're testing this out and will update by tomorrow
b
Any insights / new data from either the v2.101.2 Pex update or @echoing-psychiatrist-92346's change?
r
We initially got bit by the
LAST_ACCESS_FILE
write since we run the app in a container with a RO filesystem, so had to roll back / re-roll out. I believe the re-rollout is happening ~now unsure if we've had enough builds go through to hit a repro yet.
b
I'll ignore that because its unsettling.
LAST_ACCESS_FILE
is old news for Pex; so how was this something new you hit?
r
Not sure @echoing-xylophone-39473 / @echoing-psychiatrist-92346 would have more info there
b
I.E. ~2 years old:
Copy code
:; git grep LAST_ACCESS_FILE
pex/cache/access.py:LAST_ACCESS_FILE = ".last-access"
pex/cache/access.py:    return os.path.join(pex_dir.path, LAST_ACCESS_FILE)
pex/venv/installer.py:                cache_access.LAST_ACCESS_FILE,

:; git blame pex/cache/access.py | grep LAST_ACCESS_FILE
991883c71 (John Sirois 2024-11-04 16:03:58 -0800 101) LAST_ACCESS_FILE = ".last-access"
991883c71 (John Sirois 2024-11-04 16:03:58 -0800 106)     return os.path.join(pex_dir.path, LAST_ACCESS_FILE)

:; git blame -- pex/venv/installer.py | grep LAST_ACCESS_FILE
991883c71 pex/venv/installer.py (John Sirois  2024-11-04 16:03:58 -0800 582)
e
The jump to v2.101.2 caused CrashLoopBackOff on read-only rootfs
Copy code
File "/bin/app/pex", line 336, in boot
    with open(os.path.join(os.path.dirname(__file__), ".last-access"), "a") as fp:
OSError: [Errno 30] Read-only file system: '/bin/app/.last-access'
which I linked to pex 2.98.3 which added an unconditional last-access write to the generated
--venv
boot script with no try/except: github.com/pex-tool/pex/pull/3226 That collides with two things on our side: • we materialize the venv at build time into
/bin/app
, i.e. on the container root filesystem, not under
PEX_ROOT
• appfile-cli (custom tool which generates our kube manifests) sets
ReadOnlyRootFilesystem: true
unconditionally for every rendered container, with no appfile knob to opt out So the boot script tries to create
/bin/app/.last-access
on a read-only filesystem and dies before reaching the entrypoint. Worth noting the write buys us nothing anyway: it exists to feed
pex3 cache prune
, which only ever manages venvs inside
PEX_ROOT
. A build-time venv at
/bin/app
is never pruned. So I set
_PEX_CACHE_SET_LAST_ACCESS=0
in all container env which skips the write and app boots clean. I was also able to reproduced the crash and the fix locally.
Waiting for the actual cache change from 2.101.2 to propagate to check if it fixed the issues we've been seeing the last couple days or not
b
we materialize the venv at build time into
/bin/app
Via what?
PEX_TOOLS=1 /bin/app venv ...
If so then the
.last-access
touch is a bug.
Yeah - that's a bug. I'll self-report and get out a fix and link that back here.
gratitude thank you 2
❤️ 1
github.com/pex-tool/pex/issues/3265 Fix: github.com/pex-tool/pex/pull/3266 Release should be out in a few hours in Pex 2.101.3.
🙇 1
Waiting for the actual cache change from 2.101.2 to propagate to check if it fixed the issues we've been seeing the last couple days or not
I think it's clear, but in case not: I don't think this will fix anything unless there are wonky perms / ownership which you all have ruled out. I only expect more diagnostics in the logs.
e
Yep, sorry should have been clearer, I rigged up the pipeline to emit some metrics for how often we are seeing this based on the diagnostics, that's what I'm waiting to permeate
b
Ok, the
.last-access
fix is available here: https://github.com/pex-tool/pex/releases/tag/v2.101.3
🤝 1
Alright, I'll be away for the rest of the day climbing. I'll check in for any new information this evening. As always, since this is async, the more data the better.
🧗‍♂️ 1
r
Enjoy! we'll def report back once we have some more info.
e
New data point from CI running pex 2.101.2, all of the failed CI runs we're looking at are docker image builds, and they split into two patterns. Some of the failed runs show the diagnostics firing exactly as designed: the `atomic_directory.py:290`/`:301` pair,
EEXIST
on
.lck.work
, then "Continuing to forcibly re-create the work directory" all across several different
downloads/
cache keys. In every case that's the only pair that appears; neither the "Failed to forcibly re-create" nor "Using new random workdir instead" follow-ups fired, meaning
rmtree
+
mkdir
succeeded cleanly on the first retry. Self-heal worked as intended in the runs that ultimately passed. Where the run still failed overall despite this pattern, the log tail cuts off before showing the actual terminal error, so I can't yet say whether that failure is downstream of this same event or unrelated Exact warnings, representative of this group:
Copy code
[WARN] .../pex_root/installed_wheels/.../pex-2.101.2-py2.py3-none-any.whl/pex/atomic_directory.py:290:
  PEXWarning: [pid:<pid>, tid:<tid>, cwd:/tmp/pants-sandbox-<x>]: After obtaining an exclusive lock on
  .../pex_root/downloads/2/.<hash-a>.atomic_directory.lck,
  failed to establish a work directory at
  .../pex_root/downloads/2/<hash-a>.lck.work due to:
  [Errno 17] File exists: '.../downloads/2/<hash-a>.lck.work'

.../pex/atomic_directory.py:301: PEXWarning: [pid:<pid>, tid:<tid>, cwd:/tmp/pants-sandbox-<x>]:
  Continuing to forcibly re-create the work directory at
  .../pex_root/downloads/2/<hash-a>.lck.work.

.../pex/atomic_directory.py:290: PEXWarning: [pid:<pid>, tid:<tid2>, cwd:/tmp/pants-sandbox-<x>]: After
  obtaining an exclusive lock on
  .../pex_root/downloads/2/.<hash-b>.atomic_directory.lck,
  failed to establish a work directory at
  .../pex_root/downloads/2/<hash-b>.lck.work due to:
  [Errno 17] File exists: '.../downloads/2/<hash-b>.lck.work'

.../pex/atomic_directory.py:301: PEXWarning: [pid:<pid>, tid:<tid2>, cwd:/tmp/pants-sandbox-<x>]:
  Continuing to forcibly re-create the work directory at
  .../pex_root/downloads/2/<hash-b>.lck.work.
Same pid across both keys within a run, different tids which is consistent with your abnormal-termination read rather than a live lock race, same pattern you'd already established. Every run in this group shows the identical two-line shape, just different pids and cache keys, nothing new, just repeated independent occurrences. The other failed runs in this batch show a different pattern entirely: the exact symptom class from your earlier thread which was
ModuleNotFoundError: No module named 'pex.version'
on a
.bootstrap
zip during
venv --scope=deps
inside the Docker build but grepping the entire job log (and its sibling arch build in the same run) for `PEXWarning`/`atomic_directory.py` turns up zero matches. No EEXIST, no forcibly-recreate, nothing. CI's own retry-with-scratch-PEX_ROOT logic fired more than once per run and the corruption reproduced identically every time. Exact error, representative of this group:
Copy code
#9 [deps 3/3] RUN PEX_TOOLS=1 /usr/local/bin/python3 /binary-deps.pex venv --scope=deps --compile /bin/app
#9 0.246 Traceback (most recent call last):
#9 0.246   File "<frozen runpy>", line 198, in _run_module_as_main
#9 0.246   File "<frozen runpy>", line 88, in _run_code
#9 0.246   File "/binary-deps.pex/__main__.py", line 280, in <module>
#9 0.246     result, should_exit, is_globals = boot(
#9 0.246                                       ^^^^^
#9 0.246   File "/binary-deps.pex/__main__.py", line 247, in boot
#9 0.246     from pex.variables import ENV, Variables
#9 0.246   File "/binary-deps.pex/.bootstrap/pex/variables.py", line 18, in <module>
#9 0.246   File "/binary-deps.pex/.bootstrap/pex/common.py", line 26, in <module>
#9 0.246   File "/binary-deps.pex/.bootstrap/pex/enum.py", line 13, in <module>
#9 0.246     'global_flag_repr', 'global_enum_repr', 'global_str', 'global_enum',
#9 0.246 ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
#9 0.246   File "/binary-deps.pex/.bootstrap/pex/exceptions.py", line 12, in <module>
#9 0.246 ModuleNotFoundError: No module named 'pex.version'
#9 ERROR: process "/bin/sh -c PEX_TOOLS=1 /usr/local/bin/python3 /binary-deps.pex venv --scope=deps --compile /bin/app"
    did not complete successfully: exit code: 1
A
.bootstrap
chroot missing
pex/version.py
in every run in this group, same failure signature as the earlier builds you and I discussed, reproduced across multiple retries against a fresh scratch
PEX_ROOT
before the run gave up. That's exactly the shape your "second mechanism" concern predicts: a corrupted artifact with the identical downstream symptom, but no trace in the loud
atomic_directory
EEXIST path #3263 instruments which is consistent with the
pep_427.py:1119-1131
skip-on-exists copies (
safe_copy(overwrite=False)
,
elif not os.path.exists(dst_file)
, the symlink-EEXIST swallow) quietly preserving stale content into a populate that never trips `os.mkdir`'s EEXIST check in the first place, so there's nothing for #3263 to warn about.
r
Aiden, do we see any logs from pants indicating that per got terminated outside of its control?
e
Yea I think so, and we see these two logs arrive at the same time which is kinda sus
Copy code
20:07:14.84 [INFO] Canceled: Building 79 requirements for .../migration/alembic_pex.pex from the
  3rdparty/python/deps_lock.txt resolve: PyMySQL, aiohttp, alembic, anthropic, ... (18.3s)
20:07:14.84 [WARN] .../pex/atomic_directory.py:290: PEXWarning: [pid:<pid>, ...]: After obtaining
  an exclusive lock on ..., failed to establish a work directory at ... due to: [Errno 17] File exists
Pants is killing one pex build which lines up exactly with another pex process tripping over a leftover .lck.work dir. Also Pants canceled a pex build very shortly into building it, and that's the exact pex that shows up corrupted later in the same run.
r
Yeah those pants cancellations are absolutely expected for how pants does things (its management of compute + hedging remote cache requests). But that's not the same as a pex getting OOM killed out from under pants for example.
b
It is the same from Pex's perspective and neither method of kill breaks the assumptions of atomic_directory which is flock + atomic rename or nothing.
e
also that 2.101.3 readonly patch is working great
b
Unfortunately, climbing cancelled.
😭 1
📉 1
So it sounds like we know nothing new.
e
yea, not really, weve put a mitigation in place to disable the remote cache on retries though (not that thats a solution)
b
So, focusing on docker image builds. That's via docker in docker. Has that fact changed recently?
r
Its DinD but by mounting the socket
so the containers end up being siblings on the same host
b
Same question remains. Is that new?
r
Nope.
b
K
r
Yeah there's not a ton to go on 😞
b
So ... if I were an employee with access I'd be eliminating the dnd and just running, say, tests that had no dep on a docker image being built.
And see if issue still occurs
Pants builds lots of packed PEXes and cancels things running tests.
So, do you guys have access to these PEXes?
I'd love to see one to look at the exact missing things for example. May or may not shed light.
e
Ok, I was able to pull a couple poisoned pex's from the cache and they seem to point to
os.walk
failing silently.
.deps/mypy-1.8.0-…whl
blew up but is a perfectly valid zip with 2 members:
.layout.json
plus one
.so
where the healthy packed wheel has 1300, including 81 directory entries. The poisoned one has zero, i.e. the walk never descended past the chroot root. No dist-info/METADATA`, which is what
PEX_TOOLS=1 … venv
chokes on.
deterministic_walk
(
pex/common.py
) wraps
os.walk
and no call site passes
onerror
, so a failed directory listing gets discarded and that subtree is silently omitted then the walk then finishes "successfully" with a strict subset of the tree.
PEXBuilder._add_dist
registers only the chroot's root-level files,
_build_packedapp
zips that subset, and
atomic_directory
takes its success path and finalizes it.
_add_dist
returns the fingerprint of the complete chroot, not a hash of what it actually collected so the truncated wheel gets cached under a key asserting it's whole. The fix is small github.com/pex-tool/pex/pull/3269 default
onerror
to re-raise so we fail at the listing instead of shipping a broken PEX. Reproduced it against pex's own builder and added regression tests. tbh it is still slightly unknown what makes these fail on our build agents, my best guess is fd exhaustion from the parallel docker builds or ENOENT from concurrent cache mutation.
b
Excellent. That explanation holds water. I will be interested to see what raising on walk error reveals about your CI setup! Please don't wait on me - out climbing. Rip your own release with
uv run dev-cmd package
and report back!
🫡 2
🧗‍♀️ 1
🧗 1
e
For a quick update, I built the package and deployed it out, we have not seen any instances of the poisoning since, im going to continue letting it percolate for the rest of the day though then collect some hard numbers
b
Ok. But you should see hard errors right? That's the trade, no slow poison error, a quick hard error.
e
yep and it should fail at
pants package
on the pex build now instead of 20 min later at the docker build stage. I'm only sampling our main pipeline so far and haven't caught one there yet. Pulling data across all branches/pipelines later today for a real sample. I put in a guard yesterday as well to detect and evict incomplete cache entries preflight which has been helping as well to clean out the poison, but this should make it so no incomplete packages get in in the first place
b
Alright @echoing-psychiatrist-92346 your
os.walk
fail fast fix is now available here: https://github.com/pex-tool/pex/releases/tag/v2.101.4
🙌 2
gratitude thank you 3
👀 1
Surely something blew up by now.
r
I took a look over the long weekend + fri and I think the proactive check + cleanup mechanisms we put in place to stabilize builds for folks are preventing this error from being hit right now. @echoing-psychiatrist-92346 would be in a better spot to confirm though since he did the majority of that work.
b
Thanks for the report. That's definitely very disappointing from a debug perspective. Stability is stability though, so it's good you have that back.
r
Yeah I think with this hard error now in place unwinding some of those mitigations is now safe (ie we avoid most cache poisoning scenarios). so we can 👀 that.
e
sorry, was out for the extended weekend so just catching back up, but yea like jonathan said we have been catching and remediating the bad caches now automatically through our ci pipelines, but you can see the decline in stage failures starting with the mitigation, then with the implementation of the updated pex version (to further limit the bad pex from making it to the remote cache)
b
I will be very interested to learn what the now hidden errors are. It seems likely some other user is doing whatever you're doing in CI to trigger directory walk errors, which seem overwhelmingly likely to be the culprit here (the only other known is a file system setup that invalidates flock / atomic directory rename assumptions - and you all have ruled that out).
e
let me rig somthing up to try and capture / alert on these in our pipelines
b
Thanks. I assume this would be interesting to you all as well. Disk-fulls, ulimits too low, weird file perms ... something like that must be happening in CI.
Given a Pants upgrade, my money is on ulimits. Pants can change its behavior in ways that ramp up resource usage.