Does sharding in pants offer some guarantees aroun...
# general
a
Does sharding in pants offer some guarantees around target selections across runs in my CI? I'm observing that it doesn't get cached So if my 1st run took X mins and all my tests passed, even on my 2nd, 3rd .. runs it still takes X mins even if no code level changes were made. I think this is because target selection is random so the same shard might be getting different targets each time and not being able to fully utilize the CI cache. Will look more into this but any help would be appreciated
h
Are you referring to test sharding using
--shard=x/n
?
a
Yes
h
The input set of test files is sharded by hash, so if the same set of input targets is applied then the same tests should be sharded together every time.
But if no code changes were made then I would expect all those to be resolved from cache
✅ 1
Sounds like this is in CI. What are you using to preserve the cache across runs?
a
Copy code
- name: Set up Pants Caching
        uses: pantsbuild/actions/init-pants@main
        with:
          gha-cache-key: v0-test-shard-${{ matrix.shard }}
          named-caches-hash: ${{ hashFiles('uv.lock', 'pants.toml') }}
          cache-lmdb-store: "false"  # Managed explicitly below so saves survive cancel-in-progress
      - name: Restore Pants LMDB Store
        id: restore-lmdb
        uses: actions/cache/restore@v4
        with:
          path: ~/.cache/pants/lmdb_store/
          key: pants-lmdb-v1-shard-${{ matrix.shard }}-${{ github.event.pull_request.base.sha || github.sha }}
          restore-keys: |
            pants-lmdb-v1-shard-${{ matrix.shard }}-
      - name: Test (shard ${{ matrix.shard }}/2)
        run: |
          pants test \
            --changed-since=origin/${{ github.base_ref }} \
            --changed-dependents=transitive \
            --shard=$((${{ matrix.shard }}))/N 
      - name: Save Pants LMDB Store
        # always() ensures this runs even when the job is cancelled via concurrency
        # cancel-in-progress — otherwise a new push cancels the running job before
        # the init-pants post-step can persist the LMDB, creating a cold-cache cycle.
        # GHA save is a no-op when the exact key already exists (immutable keys),
        # so we always attempt the save without checking cache-hit.
        if: always()
        uses: actions/cache/save@v4
        with:
          path: ~/.cache/pants/lmdb_store/
          key: pants-lmdb-v1-shard-${{ matrix.shard }}-${{ github.event.pull_request.base.sha || github.sha }}
Arrived at this fix / conclusion after observing Pants' built-in cache save runs as a post-step which gets killed on cancellation. So I manage the LMDB store manually. We use
cancel-in-progress: true
for concurrency groups in our CI e.g ( Most recent changes trigger CI all old runs are cancelled to conserve cost ). Thoughts and feedback are welcome on this approach or if pants can handle this natively
h
Yeah, if you get the caching you expect when running manually on a desktop machine with persistent storage, and then you don't get it in CI, that is the thing to look for.
But when you mean "build-in cache save" are you referring to the
init-pants
action? https://github.com/pantsbuild/actions/tree/main/init-pants
I haven't looked at that one in a while, but it uses the standard
<https://github.com/actions/cache>
under the hood. It sounds like that should use
actions/cache/restore
and
actions/cache/save
explicitly as you've done above?
a
Yes, referring to the init-pants action it uses
actions/cache@v5
which saves via a post-step hook that it registers at restore time. Post-steps run on success/failure but do not run on cancellation So with
cancel-in-progress: true
1. Run gets cancelled mid-way → post-step save never fires 2. Next run starts with a cold cache (nothing was saved) 3. Repeat indefinitely until whole workflow completes ( We run numerous checks in parallel and developers start their fixes immediately with A.I thus cancel in progress is quite likely) Might be worth considering this pattern in the init-pants action itself as using
cancel-in-progress: true
(which is pretty common) would hit the same issue with the LMDB store cache
h
Yeah, so I'm thinking that the change you made locally might make sense to upstream to the action
a
https://github.com/pantsbuild/actions/pull/48 Let me know if you have feedback, I could also do a backward compatible approach