<#23694 Idle pantsd spends significant CPU re-exte...
# github-notifications
q
#23694 Idle pantsd spends significant CPU re-extending store leases every 72s Issue created by omarzouk Is your feature request related to a problem? Please describe. An idle
pantsd
(no client attached, nothing run for days) keeps using CPU in short bursts, indefinitely. On a large polyglot monorepo (macOS, Pants 2.33.0a0) we had 5 idle daemons, one per git worktree, sharing
~/.cache/pants/lmdb_store
. Sampling
top
every 5s for 2 minutes showed them using about one core between them, each spiking to 85–120%. System time was ~30%. A 5s
sample
of the oldest daemon (idle 33 days) is dominated by
StoreGCService
→ `lease_files_in_graph`:
Copy code
mdb_page_search_root                     14721
sharded_lmdb::ShardedLmdb::get_raw         954
stat                                       833
mdb_node_search                            214
engine::externs::interface::__pyfunction_lease_files_in_graph
The cause is in `StoreGCService`:
Copy code
lease_extension_interval_secs: float = (float(LOCAL_STORE_LEASE_TIME_SECS) / 100),
LOCAL_STORE_LEASE_TIME_SECS
is 2h, so every daemon re-leases its whole graph every 72 seconds, though the leases it writes are still valid for nearly 2 hours. Each pass is a full, non-incremental sweep: 1.
Scheduler::all_digests
walks every node in the graph (
visit_live
over all
node_indices()
) and collects every digest. 2.
Store::lease_all_recursively
→
expand_local_digests
runs
entry_type
on each digest (two LMDB existence checks plus an FSDB
stat
), then walks every directory digest again. 3.
ByteStore::lease_all
leases one digest at a time, with a separate LMDB write transaction per digest, or an
mtime
update for FSDB files. The cost grows with everything the daemon has ever loaded and never drops while it lives. It is the same work every 72s whether or not anything has run. With several daemons sharing one store (one per worktree is common), each repeats it independently. There is no option to tune this:
StoreGCService
is built in
PantsDaemon._setup_services
with only
local_store_options
, and both the interval and the lease time are constants. Describe the solution you'd like Two small changes, independent of each other: 1. Lease less often. Change the default extension interval from
lease_time / 100
(72s) to something like
lease_time / 2
(1h), and/or expose it as an option such as
[GLOBAL].pantsd_lease_extension_interval
. Renewing at half the lease time still leaves an hour of margin before a lease could expire. The 72s pass is ~50× more frequent than that margin requires. 2. Skip passes when nothing has run. If no session has run since the last pass, the set of digests referenced by the graph has not changed and their leases are still valid.
StoreGCService
could keep a cheap "graph touched" marker (the scheduler bumps a counter when a session runs) and: • re-lease after any session has run since the last pass, as today, because a run can make the graph reference existing store entries without leasing them (
initial_lease: false
on some write paths), and • otherwise skip, unless the last full pass is older than about
lease_time / 2
. With this, a truly idle daemon does one pass per hour instead of 50. Together these make an idle daemon close to free, without changing GC safety: leases are still renewed well before they expire, and they are renewed promptly after any run. Describe alternatives you've considered • Lowering
pantsd_max_memory_usage
to restart daemons more often: this resets the graph but costs cold starts during active work, and does nothing for a daemon that never runs again. • Killing idle daemons from outside Pants (our current workaround, with a shell function that checks
.pants.d/workdir/run-tracker
age). An idle-shutdown option for pantsd would be a natural companion to this issue, but is separate. • Deeper fixes to the pass itself: leasing incrementally as nodes complete, carrying entry types through
NodeOutput::digests()
so the `entry_type`/`expand_directory` step isn't needed (related TODO: #13112), and batching LMDB lease writes into one transaction per shard. Worth doing, but larger changes. Additional context • Related: #13558 (lease extension freezing the UI in a repo with large requirement sets; closed without changing lease cost). • Code referenced (at
release_2.33.0a0
, unchanged on
main
):
src/python/pants/pantsd/service/store_gc_service.py
,
src/rust/engine/src/externs/interface.rs
(
lease_files_in_graph
),
src/rust/engine/src/scheduler.rs
(
all_digests
),
src/rust/fs/store/src/lib.rs
(
lease_all_recursively
,
expand_local_digests
),
src/rust/fs/store/src/local.rs
(
lease_all
),
src/rust/sharded_lmdb/src/lib.rs
(
lease
). Happy to send a PR for (1) and/or (2) if the approach sounds right. pantsbuild/pants