Hi folks — we’re debugging a *CI-only remote cache...
# general
m
Hi folks — we’re debugging a CI-only remote cache miss issue for
pants test
and would love a sanity check on whether this matches any known Pants behavior. 🧵
Setup Pants 2.31.0 GitHub Actions CI Remote cache via Depot REAPI (
<grpcs://cache.depot.dev>
) Linux x86_64 in both runs Python 3.13 • We also reproduced the same behavior on Depot’s GitHub runner, so it doesn’t seem specific to our self-hosted runner image Symptom •
pants lint
and
pants package
get good remote cache hits
pants test
gets 0 remote hits between consecutive CI runs Example: 0/374 remote hits for the
test
step, with all 374 results written fresh • Locally, against the same Depot cache, we do get remote cache hits for tests What we’ve already controlled for • Same code / same effective test scope across the compared CI runs • Stable CodeArtifact token across runs • Same Linux x86_64 platform • Removed rotating AWS creds from
[subprocess-environment].env_vars
• Tests only receive needed secrets via
extra_env_vars
• Tried disabling
xdist
,
xdist_enabled=false
Why we’re suspicious of Pants test process construction The failure mode is very test-specific:
check
lint
`package`etc. hit remote cache,
test
does not • The
test
run seems to consist of many
requirements.pex
/
pytest_runner.pex
/
Run Pytest
processes • We’re wondering whether something CI-specific is still entering the process fingerprint for test/PEX setup only Questions Is there any known issue in Pants 2.31 where pytest / PEX setup / dynamic concurrency can cause cross-machine remote cache misses even when source inputs are stable? Are there known cases where interpreter search paths / PATH / concurrency_available / runner-local env end up affecting the cache key for test-related processes but not
check
? 1. Is there a recommended way to diff the exact process definition (argv/env/input digest/platform) for one cache-missing
Run Pytest
or
requirements.pex
action across two CI runs? If helpful, we can share sanitized
pants.log
excerpts. We’re mostly trying to determine whether we should keep digging in our CI env, or whether this sounds like a known Pants remote-caching edge case for
test
.
h
There is no known issue, but obviously there is an issue
Let me see what debug info we expose on the process definition
Hmm, not much... So, step 1 might be to log the request.description (or whatever info from the Process helps you identify it uniquely) and action_digest here and then again here and compare the digests for the read and write of the same process. We can put this into TRACE level logs in the next dev release (some time this weekend). But you can also run Pants from source in CI (clone your fork of the Pants repo and set PANTS_SOURCE=/path/to/pants/repo), and then
pants
will run it from source there while still acting on the code in your repo.
👀 1
If the action digests are identical then we have one kind of problem. If they are not (as I suspect) then we will start investigating that.
👀 1
Copy code
diff --git a/src/rust/process_execution/remote/src/remote_cache.rs b/src/rust/process_execution/remote/src/remote_cache.rs
index c040790617..2e543eda48 100644
--- a/src/rust/process_execution/remote/src/remote_cache.rs
+++ b/src/rust/process_execution/remote/src/remote_cache.rs
@@ -10,6 +10,7 @@ use fs::{DigestTrie, RelativePath, SymlinkBehavior, directory};
 use futures::FutureExt;
 use futures::future::{BoxFuture, TryFutureExt};
 use hashing::Digest;
+use log::{log_enabled, trace};
 use parking_lot::Mutex;
 use protos::pb::build::bazel::remote::execution::v2 as remexec;
 use protos::require_digest;
@@ -505,7 +506,11 @@ impl process_execution::CommandRunner for CommandRunner {
         let use_remote_cache = request.cache_scope == ProcessCacheScope::Always
             || request.cache_scope == ProcessCacheScope::Successful;
 
+        let proc_descr = if log_enabled!(Level::Trace) { Some(request.description.clone()) } else { None };
         let (result, hit_cache) = if self.cache_read && use_remote_cache {
+            if let Some(proc_descr) = proc_descr.as_deref() {
+                trace!("Checking remote cache for process {} against action digest {:?}", proc_descr, &action_digest);
+            }
             self.speculate_read_action_cache(
                 context.clone(),
                 cache_lookup_start,
@@ -527,6 +532,9 @@ impl process_execution::CommandRunner for CommandRunner {
             && self.cache_write
             && use_remote_cache
         {
+            if let Some(proc_descr) = proc_descr.as_deref() {
+                trace!("Updating remote cache for process {} against action digest {:?}", proc_descr, &action_digest);
+            }
             let command_runner = self.clone();
             let result = result.clone();
             let write_fut =
m
We can put this into TRACE level logs in the next dev release (some time this weekend). But you can also run Pants from source in CI (clone your fork of the Pants repo and set PANTS_SOURCE=/path/to/pants/repo), and then
pants
will run it from source there while still acting on the code in your repo.
If you could do this in the next dev release that would be absolutely fantastic! I'll try building from source but I'm stretched pretty thin (aren't we all 🙃 ) but I'd be more than happy to test out the dev release and report back
h
FWIW I could not reproduce in our own CI, which uses bazel-remote backed by S3