Hi John! Appreciate the message, sorry for the la...
# general
s
Hi John! Appreciate the message, sorry for the late response
b
Did you get answers on what was going on over at namespace? The whole posix lock thing has been a problem in the past and I've sunk alot of time into it over the years and so I'm generally interested.
s
Yep! So we configure our cache volume size as 50GB. On their end, this means "at least 50GB of free space + the cached contents". So if our job runs using a cache volume with 25GB of contents, we get at least a 75GB volume for that job. After the job finishes, they only put the cache volume back in the pool if it remains under the chosen size (50GB for us)
For us, that volume is also used for caching docker images. They have something that cleans up stale docker images if needed to attempt to bring it under the 50GB But the other files shouldn't be touched
We can see the lineage of a cache volumes (and jobs they were used in I think), so I'm adding some logging to see if the files are mysteriously disappearing or not
b
And you confirmed no funny business with storage drivers for docker that do not support posix locks on the volume mounts?
The short of it being the only 2 things I know of that should cause this are nuke behind Pex's back and PEX_ROOT on filesystem that does not support posix locks.
s
I did not - I will ask them, thanks!
If I shell in to the CI machine, is there an easy CLI command I can do to sanity check the posix locks thing? Google or AI can answer this for me
b
I'll spend a little time looking at your case again today, but will warn you I spent ~1.5 years on a prior case when I actually worked on Pants - assumed it was a bug in my Pex code and did all sorts of diagnostics over that 1.5 year period, and, in the end, it was a Pants engine bug in that ase that got fixed ... so I am a bit wary of sinking too much time!
Yeah, just use df:
Copy code
$ df   
Filesystem                         1K-blocks      Used  Available Use% Mounted on
overlay                           3933680272 724180796 3009604644  20% /
tmpfs                                  65536         0      65536   0% /dev
shm                                    65536         0      65536   0% /dev/shm
...
Here root
/
is
overlay
, for example and
/dev/shm
is
shm
.
Just report back what the filesystem is for the mount that houses
/home/runner/.cache/pants
Google or AI can answer this for me
Huge BS to that. No it cannot! Unless it can hack into the machine in question or it has backdored namespace's business or your business?
You need access to the machine to get this answer.
s
sorry I meant the shell command to run to test it out, it suggested running
flock
b
Oh, right.
s
Copy code
sh-5.1$ df
Filesystem     1K-blocks     Used Available Use% Mounted on
...
/dev/vdi        50471392 23559424  26895584  47% /cache
...
Copy code
sh-5.1$ mount | grep vdi
/dev/vdi on /cache type ext4 (rw,relatime)
/dev/vdi on /opt/hostedtoolcache type ext4 (rw,relatime)
/dev/vdi on /home/runner/.gradle/caches type ext4 (rw,relatime)
/dev/vdi on /home/runner/.gradle/wrapper type ext4 (rw,relatime)
/dev/vdi on /home/runner/.local/share/pnpm/store/v10 type ext4 (rw,relatime)
/dev/vdi on /home/runner/.cache/nce type ext4 (rw,relatime)
/dev/vdi on /home/runner/.cache/pants/named_caches type ext4 (rw,relatime)
b
Ok, now
grep /dev/vdi /etc/mtab
s
Copy code
sh-5.1$ flock /home/runner/.cache/pants/named_caches/.darren echo "testing"
testing
This seemed to work
Copy code
sh-5.1$ grep /dev/vdi /etc/mtab
/dev/vdi /cache ext4 rw,relatime 0 0
/dev/vdi /opt/hostedtoolcache ext4 rw,relatime 0 0
/dev/vdi /home/runner/.gradle/caches ext4 rw,relatime 0 0
/dev/vdi /home/runner/.gradle/wrapper ext4 rw,relatime 0 0
/dev/vdi /home/runner/.local/share/pnpm/store/v10 ext4 rw,relatime 0 0
/dev/vdi /home/runner/.cache/nce ext4 rw,relatime 0 0
/dev/vdi /home/runner/.cache/pants/named_caches ext4 rw,relatime 0 0
b
ext4
Ok, that's normal filesystem that does support posix locks. Thanks
AI will screw you in its current state. That was a poor answer re flock. The whole problem with NFS mounts is silent lock failure (i.e. lock "succeeds" but doesn't actually lock things, leading to races and corruption). Unless flock somehow detects that, no dice.
s
Yea my next test was to run 2 shells with sleep + echo, see if it actually was locking or not
b
Actual good info: https://www.samba.org/samba/news/articles/low_point/tale_two_stds_os2.html There are a few other articles out there about this.
👀 1
s
is it weird that I don't even have a
/home/runner/.cache/pants/named_caches/pex_root/bootstrap_zips
after
pants test ::
was run?
nvm, everything was from remote bazel cache 🤦
b
Well, pants caches things - did the test runs get cached?
Right.
The thing I'll investigate in your case is the RHS. The LHS is from the PEX_ROOT cache, but the RHS is the tmp .zip Pex is creating to move to the final .zip atomically when outputting a PEX file.
If the RHS is what generate the error message, that's not similar to the 1.5 year investigation case and is novel enough for me to spend time on today to suss out.
Yeah, quick experiment suggests that error could be from either LHS or RHS:
Copy code
# LHS exists (src) RHS (dest dir) does not exist
:; touch x
:; ls y
ls: cannot access 'y': No such file or directory
:; mv x y~/x
mv: cannot move 'x' to 'y~/x': No such file or directory

:/ mkdir y
:; rm x
:; mv x y~/x
mv: cannot stat 'x': No such file or directory
You can see
ls
gives extra prefix info identifying which, but the final phrase after last
:
is the OSError.
s
Eh appreciate it, but don't kill yourself digging for something that might not even exist I wouldn't be surprised if these CI runs are the only place this is happening lol, probably something stupid we're doing
b
See, that's where you can't unsee the thing. I take issues like this, if they are real Pex issues - super seriously. Combine that with a mantra of never trust yourself and I sorta am compelled to fully understand I'm not screwing something up here.
❤️ 1
s
Will keep this thread updated if I spot anything interesting from the logging
Will also mention, we're currently on pants 2.29.1 And when running with PEX warnings enabled, we got these: https://gist.github.com/darrenclark/830483a550337e5a9b67677c27f44f97 Didn't get the
No such file or directory
this time around - will let you know if I get any better logs
b
I'm adding a better diagnostic on the error codepath and I'll get a Pex release out with it today. I'll let you know when that is available, but you won't be able to use it until you get to Pants 2.30.1 or newer. Normally you can freely upgrade Pex without concern over Pants version, but there was a bug in Pex and another in Pants that relied on that bug and ... the 1 instance I can recall in the last 5+ years where you cannot freely upgrade.
s
Okay, cool, thanks, appreciate it! Will try it out next week
b
s
Upgrading to it this morning, thanks John! Perhaps something interesting from my logging:
Copy code
[Errno 2] No such file or directory: '/home/runner/.cache/pants/named_caches/pex_root/bootstrap_zips/0/23f6f08a1ef609bba85fd8d67dd6af6ab14d1a98/.bootstrap' -> 'pytest_runner.pex~/.bootstrap'
and running
tree tree -a -L 4
on that directory afterwards:
Copy code
/home/runner/.cache/pants/named_caches/pex_root/bootstrap_zips
  └── 0
      ├── .665f5cc6c9dbc44f942d0ea8bb0f192ba5292a3d.atomic_directory.lck
      ├── 23f6f08a1ef609bba85fd8d67dd6af6ab14d1a98
      │   ├── .un-compressed.atomic_directory.lck
      │   └── un-compressed
      │       └── .bootstrap
      └── 665f5cc6c9dbc44f942d0ea8bb0f192ba5292a3d
          └── .bootstrap
I've seen an issue like this before. At one point, some of our
pex_binary
targets used
extra_build_args=["--no-compress"]
to speed up reloads during local development (and other
pex_binary
targets didn't). This caused similar issues, so we made all
pex_binary
use the
--no-compress
flag Could this be a conflict between the pytest pex packaging & packaging for our main apps?
We have a macro that does:
Copy code
def our_pex_binary(**kwargs):
   # etc...
   # etc...
   # etc...

   pex_binary(
       # etc...

        # TODO: Enabled for dev for faster reload.  We should remove the
        # '--no-compress' flag when building a docker image for deployment
        # maybe?  Need to be careful, if we mix --no-compress & not, we
        # get random build failures.
        extra_build_args=["--no-compress"],
    )
I wonder if we should move this to
pants.toml
via:
Copy code
[pex-cli]
global-args = ["--no-compress"]
🤔
Alternatively, I wonder if the pex caching stuff needs to incorporate
extra_build_args
(or at least
--no-compress
) as part of that hash in the folder name 🤔 If this sounds like a decent idea, I'm happy to open a PR on the appropriate project!
b
Aha. The
--no-compress
option & the tree output make it pretty clear this is a Pex bug. Thanks for those details. I'll need to analyze more closely when I get to a keyboard, but I should have a fix out in the by tomorrow latest.
Yeah. easy repro with that info:
Copy code
# Create a packed PEX with --no-compress:
:; pex --pex-root /tmp/repro --no-compress --layout packed -o repro.pex
repro.pex

# Now attempt to create a standard packed PEX with the same Pex version (the .bootstrap tends to vary by Pex version):
:; pex --pex-root /tmp/repro --layout packed -o repro.pex
[Errno 2] Failed to link /tmp/repro/bootstrap_zips/0/b02d0030281010c1aa3be923a7002e5e62718dcf/.bootstrap -> repro.pex~/.bootstrap: No such file or directory
login: jsirois uid: 1000 gid: 1000
effective: uid: 1000 gid: 1000
groups: [4, 24, 27, 30, 46, 100, 114, 984, 993, 1000]
<unknown path type> err '/tmp/repro/bootstrap_zips/0/b02d0030281010c1aa3be923a7002e5e62718dcf/.bootstrap':
    [Errno 2] No such file or directory: '/tmp/repro/bootstrap_zips/0/b02d0030281010c1aa3be923a7002e5e62718dcf/.bootstrap'
dir ok? 'repro.pex~':
    mode: 0o40775 (drwxrwxr-x) owner: 1000 group: 1000
    os.stat_result(st_mode=16893, st_ino=35952373, st_dev=64513, st_nlink=3, st_uid=1000, st_gid=1000, st_size=4096, st_atime=1771347922, st_mtime=1771347922, st_ctime=1771347922)
Ok, fix is up here: https://github.com/pex-tool/pex/pull/3106 Thanks @some-insurance-58590 - I'll ping when the release is out in a few hours.
s
wow, that was fast 🙌 Thank you so much!
b
These things tend to be easy with good info. Your details were crucial. Generally actually looking at the filesystem like you did is the key. Most folks don't bother to do that.
❤️ 1
Alright @some-insurance-58590 this should allow you to mix and match
--no-compress
as you see fit with no cache issues: https://github.com/pex-tool/pex/releases/tag/v2.90.1
❤️ 1
s
Thank you so much John!
These things tend to be easy with good info. Your details were crucial
Happy to help where I can
Fix has been working great, thanks John!
b
Excellent. You're welcome.