Hi y'all, I wanted to flag our use-case where we f...
# general
r
Hi y'all, I wanted to flag our use-case where we frequently see cache-corruption and need to manually clear pants cache and rebuild, and/or delete
.pants.d/pids
This happens to us most commonly when we're running pants on SLURM machines with the pants directory on a shared, mounted, home folder. Users may have multiple terminals on the same node, and even multiple nodes, all running with the same home folder mounted. I suspect something about our set up is causing pants corruption to be very common (eg once a day) and wondering if this is a known issue, and if there's anything we can do to workaround
Example error which requires purging `.cache/pants`:
Copy code
Engine traceback:
  in `export` goal

IntrinsicError: Couldn't find file contents for "playground/src/components/playground/BUILD": Got hash collision reading from store - digest Digest { hash: Fingerprint<3e40830e1c1aa7488a5e93d134299bd4d7e0e627858fa6145780f10777776966>, size_bytes:
 70 } was requested, but retrieved bytes with that fingerprint had length 8. Congratulations, you may have broken sha256! Underlying bytes: [162, 206, 35, 104, 0, 0, 0, 0]
Users typically run
pants export --py-resolve-format=symlinked_immutable_virtualenv --resolve=python-default
and the
source dist/export/...
to activate a virtual env, and then run
PYTHONPATH=. python path-to-my-script.py
in order to iterate locally
b
the pants directory
Just to be 100% sure, which pants directory is this?
.pants.d/
or something else?
r
The whole home directory, including both
~/.cache/pants
and ~/my_repo/.
pants.d
b
Hmm, pants should be tolerant of concurrent usage of the cache, but maybe having multiple machines operating on a shared drive is particularly "hostile" usage.
r
It may be a red herring, I do notice then when I have eg many terminals open on a single machine , where I have multiple python servers running using the python binary in the generated virtual environment. I’m not certain what causes the corruption error but every now and then, trying to run pants throws that error which is unrecoverable, and we then need to spend 10 minutes deleting the cache and recreating the virtualenv.
g
What is the mounting mechanic/fs type?
r
Not totally sure what would be helpful, but
df -T ~
reports
nfs
.
g
Ok. I wonder if that could be a cause, nfs has limited support for concurrent writes. Often leads to data loss/corruption. In particular it has no support for append, locks, etc.
Concurrent writes of the same files, to be clear.
r
got it
is that a problem in general for pants? Ie pants would never be happy on an NFS filesystem, or is it because of possibly multiple daemons (eg vscode, pantsd) trying to read/write at the same time
or multiple pantsd on different TTY sessions
g
Anyting databasy using file locks is going to struggle. Though now that I think of it, if multiple things are using the same home directory, wouldn't you get issues with the pants daemon itself, the pid files, etc? I'd imagine single-machine multiple-terminals complains about pantsd parallel invocations, but what happens if two nodes run Pants commands in the same directory? Wouldn't they see each others pid files?
r
That definitely coudl be happening, however I'm seeing the corruption issues even when only a single machine is working on the repo
g
Judging by the LMDB docs; I'd not trust that to work either.
Do not use LMDB databases on remote filesystems, even between processes on the same host. This breaks flock() on some OSes, possibly memory map sync, and certainly sync between programs on different hosts.
If you want any chance of having this work I'd start by validating that both your client and server are using NFSv4 which does have some semblance of functional flock.
r
oh wow, nice find thanks Tom!
I can see that
/tmp
is formatted as xfs, Maybe I can try sticking
.cache/pants
there, if I can find the configuration option
g
Yeah. If you need/want shared cache I would investigate a remote cache solution instead, there's a few different ones.
r
we're using remote bazel cache, there's other issues with this environment though since we're accessing cache over VPN and these slurm machines can't connect to vpn... but that's not a pants problem.
oh hmm, I just realized these are multi-tenant machines, probably better to use
/tmp/pants-$USER/lmdb_store
or something