gentle-flower-25372
09/16/2025, 11:21 PMpants --changed-since=<git-ref> --changed-dependents=transitive --filter-target-type=docker_image list
Historical posts I've made in Slack regarding this issue.
https://pantsbuild.slack.com/archives/C046T6T9U/p1740007055261349
https://pantsbuild.slack.com/archives/C0D7TNJHL/p1748874243046649wide-midnight-78598
09/16/2025, 11:30 PMwide-midnight-78598
09/16/2025, 11:32 PMhappy-kitchen-89482
09/17/2025, 4:01 AMhappy-kitchen-89482
09/17/2025, 4:09 AMhappy-kitchen-89482
09/17/2025, 4:09 AMhappy-kitchen-89482
09/17/2025, 4:09 AMgentle-flower-25372
09/17/2025, 2:16 PMfast-nail-55400
09/17/2025, 2:16 PMhappy-kitchen-89482
09/18/2025, 12:15 AMgentle-flower-25372
04/07/2026, 8:40 PMmap_addresses_to_dependents for all 50K targets in our monorepo. Introducing a cache has helped dramatically.
https://github.com/pantsbuild/pants/pull/23228curved-manchester-66006
04/07/2026, 9:25 PMmap_addresses_to_dependentsInteresting. Is that from work unit logs, the new
perf support, or something else?gentle-flower-25372
04/08/2026, 2:22 PMperf support, or something else?
@curved-manchester-66006 What do you mean by "that"? Are you asking how I arrived at the conclusion that the bottleneck was the map_addresses_to_dependents ? I had claude alter the code to benchmark it with logging, etc.curved-manchester-66006
04/08/2026, 2:31 PMmap_addresses_to_dependents conclusiongentle-flower-25372
04/08/2026, 2:33 PMgentle-flower-25372
04/08/2026, 2:33 PMgentle-flower-25372
04/08/2026, 2:41 PM```# Performance Investigation: --changed-dependents=transitive on 53K targets
## The Problem
pants --changed-since=<ref> --changed-dependents=transitive --filter-target-type=docker_image list takes ~3.5 minutes in our monorepo (53K targets), regardless of how few files actually changed. This is ~30-40% of our build pipeline time.
## Root Cause
The bottleneck is map_addresses_to_dependents() in src/python/pants/backend/project_info/dependents.py. When --changed-dependents is used, this rule must build the full reverse dependency graph by calling resolve_dependencies() for every
target in the repo:
@rule(desc="Map all targets to their dependents")
async def map_addresses_to_dependents(all_targets: AllUnexpandedTargets) -> AddressToDependents:
dependencies_per_target = await concurrently(
resolve_dependencies(DependenciesRequest(tgt.get(Dependencies), ...))
for tgt in all_targets # ALL 53K targets
)
This is expensive because each resolve_dependencies() call runs dependency inference (Python import parsing, Docker COPY analysis, etc.). With 53K targets, this takes ~150 seconds of the ~210 second total.
## Why Caching Doesn't Help
- pantsd in-memory cache: AllUnexpandedTargets is https://github.com/pantsbuild/pants/blob/main/src/python/pants/engine/target.py#L1023-L1025. Any filesystem change invalidates it, which cascades to invalidate AddressToDependents. We benchmarked identical commands on warm pantsd β still ~3 minutes each time.
- Remote cache: Only caches process execution results (test runs, compilation), not rule engine computations like dependency inference.
- LMDB local store: Same β only process results, not rule outputs.
## What We Tried
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ¬ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β Approach β Result β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββΌββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ€
β Warm pantsd (identical command twice) β 3m00s β 3m01s (no improvement) β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββΌββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ€
β Forward BFS from filtered targets only (https://github.com/pantsbuild/pants/pull/23224) β 3m39s β 2m42s (26% faster, resolves 24K instead of 53K targets) β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββΌββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ€
β Persistent dep cache on disk (https://github.com/pantsbuild/pants/pull/23228) β 3m22s β 43s (5x faster) β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ΄ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
## Our Solution: Persistent Dependency Inference Cache
https://github.com/pantsbuild/pants/pull/23228 adds an opt-in --incremental-dependents-enabled flag that persists the forward dependency graph to ~/.cache/pants/incremental_dep_graph_v2.json (29MB, 1.3MB compressed). On subsequent runs, only targets with changed source files (by SHA-256 content hash) have their dependencies re-resolved.
Cold (no cache): 3m22s β 52,927 targets resolved, writes cache
Warm (from cache): 43s β 0 targets resolved, 52,927 from cache
The cache is portable across machines (SHA-256 content hashing, not mtime), so it works on ephemeral CI agents via S3.
## The Architectural Gap
The fundamental issue is that Pants' rule engine caches rule results in-memory only and has all-or-nothing invalidation β there's no concept of "re-resolve deps for 50 changed targets and reuse the other 52,877 from last run." map_addresses_to_dependents either returns a fully cached result or recomputes from scratch for all targets.
A native solution would be persistent rule-level caching in the engine (extending LMDB to cache resolve_dependencies results per-target, keyed by content hash). Our PR approximates this at the Python level.```
happy-kitchen-89482
04/08/2026, 4:49 PMresolve_dependencies() overwhelming the scheduler, even though they end up being resolved from cache. The real solution is to fix that (after first proving that it is indeed the issue).happy-kitchen-89482
04/08/2026, 4:50 PMhappy-kitchen-89482
04/08/2026, 4:50 PMgentle-flower-25372
04/08/2026, 4:54 PMgentle-flower-25372
04/08/2026, 4:59 PMgentle-flower-25372
04/08/2026, 4:59 PMhappy-kitchen-89482
04/08/2026, 5:03 PMgentle-flower-25372
04/08/2026, 5:03 PMhappy-kitchen-89482
04/08/2026, 5:04 PMhappy-kitchen-89482
04/08/2026, 5:04 PMgentle-flower-25372
04/08/2026, 5:04 PMhappy-kitchen-89482
04/08/2026, 5:05 PMgentle-flower-25372
04/08/2026, 5:05 PMhappy-kitchen-89482
04/08/2026, 5:05 PMgentle-flower-25372
04/08/2026, 5:05 PMwide-midnight-78598
04/08/2026, 5:13 PMgentle-flower-25372
04/08/2026, 5:14 PMgentle-flower-25372
04/08/2026, 5:15 PMwide-midnight-78598
04/08/2026, 5:17 PMgentle-flower-25372
04/08/2026, 5:18 PMgentle-flower-25372
04/08/2026, 5:19 PMgentle-flower-25372
04/08/2026, 5:19 PMwide-midnight-78598
04/08/2026, 5:21 PMgentle-flower-25372
04/08/2026, 5:22 PMif anything takes longer than a couple seconds, Iβll certainly be looking at youtube by then100%. Honestly most of our users don't ever use pants natively because this has been such a big pain point. We've written a lot of workarounds. Mostly exporting pants virtualenv and magic around that.
gentle-flower-25372
04/08/2026, 5:24 PMgentle-flower-25372
04/08/2026, 5:26 PMwide-midnight-78598
04/08/2026, 5:49 PMgentle-flower-25372
04/08/2026, 6:33 PMgentle-flower-25372
04/08/2026, 6:33 PM```Profiling Data: map_addresses_to_dependents with 53K Targets
I instrumented resolve_dependencies and map_addresses_to_dependents on upstream main (2.32.0.dev7) to understand where the time goes. No changes to caching or batching logic β just timing measurements on the stock code.
Setup
- 52,927 targets (20,850 python_source, 15,422 file, 8,620 resource, 2,945 python_test, plus Docker, Shell, etc.)
- Command: pants --no-pantsd --changed-since=HEAD~3 --changed-dependents=transitive --tag="-integration" --filter-target-type="+python_test" filter
- Timing added to each resolve_dependencies call and to map_addresses_to_dependents phases
Results
map_addresses_to_dependents phase breakdown:
- Resolve all dependencies (the concurrently() call): 122.5s
- Build reverse map (pure Python dict construction): 0.1s
So 99.9% of the time is in the concurrently(resolve_dependencies(...) for tgt in all_targets) call.
Per-call resolve_dependencies stats (52,927 calls):
βββββββββββββββββββββββββ¬βββββββββββββ
β Metric β Value β
βββββββββββββββββββββββββΌβββββββββββββ€
β p50 (median) β 13.4s β
βββββββββββββββββββββββββΌβββββββββββββ€
β Average β 38.4s β
βββββββββββββββββββββββββΌβββββββββββββ€
β p99 β 112.7s β
βββββββββββββββββββββββββΌβββββββββββββ€
β Max β 118.2s β
βββββββββββββββββββββββββΌβββββββββββββ€
β Sum of all calls β 2,030,764s β
βββββββββββββββββββββββββΌβββββββββββββ€
β Wall time β 122.5s β
βββββββββββββββββββββββββΌβββββββββββββ€
β Effective parallelism β ~16,500x β
βββββββββββββββββββββββββ΄βββββββββββββ
The sum-of-all-durations is ~2 million seconds but wall time is only 122s, which is consistent with ~16K+ calls in flight concurrently (as expected from try_join_all on 53K futures).
The per-call latency grows over the course of the run:
βββββββββββββββββββ¬βββββββββββββββ¬βββββββββββββββ
β Progress β Avg per call β p50 per call β
βββββββββββββββββββΌβββββββββββββββΌβββββββββββββββ€
β First 10K calls β 309ms β 119ms β
βββββββββββββββββββΌβββββββββββββββΌβββββββββββββββ€
β At 20K β 219ms β 120ms β
βββββββββββββββββββΌβββββββββββββββΌβββββββββββββββ€
β At 30K β 3,564ms β 139ms β
βββββββββββββββββββΌβββββββββββββββΌβββββββββββββββ€
β At 40K β 21,571ms β 178ms β
βββββββββββββββββββΌβββββββββββββββΌβββββββββββββββ€
β At 50K β 34,824ms β 382ms β
βββββββββββββββββββ΄βββββββββββββββ΄βββββββββββββββ
The median stays relatively low (119ms β 382ms), but the average and tail latency explode as more calls are in flight β calls submitted later wait for the full duration of all earlier calls.
Batching experiment
I also tested whether batching the concurrently() call helps reduce contention:
ββββββββββββββββββββββββββ¬ββββββββββββ
β Batch Size β Wall Time β
ββββββββββββββββββββββββββΌββββββββββββ€
β 53K (unbatched, stock) β 2m40s β
ββββββββββββββββββββββββββΌββββββββββββ€
β 5,000 β 2m44s β
ββββββββββββββββββββββββββΌββββββββββββ€
β 500 β 2m38s β
ββββββββββββββββββββββββββΌββββββββββββ€
β 50 β 2m31s β
ββββββββββββββββββββββββββΌββββββββββββ€
β 1 (fully sequential) β 2m27s β
ββββββββββββββββββββββββββ΄ββββββββββββ
Batching at the Python level doesn't meaningfully help. The overhead is below the concurrently() call β inside the engine's task scheduling, graph node management, and PythonβRust generator protocol for each of the 53K calls.
Interpretation
The per-call work is fast individually (p50 ~120ms), but 53K calls through the engine's concurrent task infrastructure accumulates significant overhead. The reverse map construction itself is trivial (0.1s). The entire cost is the engine processing 53K resolve_dependencies calls, even when the underlying dep inference results are cached.
I don't have a theory for what specifically in the engine is slow at this scale β it could be task scheduling overhead, the generator send/receive protocol, workunit management, or something else in the Rust engine. But the data clearly shows the cost is in the aggregate overhead of 53K concurrent rule calls, not in the dep inference computation itself.```
gentle-flower-25372
04/08/2026, 6:36 PMwide-midnight-78598
04/08/2026, 6:45 PMgentle-flower-25372
04/08/2026, 6:46 PMβ Here's what the data actually says:
The Real Root Cause
It's not 53K operations. It's 1.86 million.
Each of the 53K resolve_dependencies calls is a Python async generator that spawns nested generators β for resolve_target, determine_explicitly_provided_dependencies, infer_dependencies (per inference type), validate_dependencies, etc. The
Rust engine runs each as a separate generator via gen_generator().
βββββββββββββββββββββββββββββββββββββββββββββββββββ¬βββββββββββββββββββ
β Metric β Value β
βββββββββββββββββββββββββββββββββββββββββββββββββββΌβββββββββββββββββββ€
β Top-level resolve_dependencies calls β 52,927 β
βββββββββββββββββββββββββββββββββββββββββββββββββββΌβββββββββββββββββββ€
β Total generators executed β 1,860,000 β
βββββββββββββββββββββββββββββββββββββββββββββββββββΌβββββββββββββββββββ€
β Avg generators per resolve_dependencies β ~35 β
βββββββββββββββββββββββββββββββββββββββββββββββββββΌβββββββββββββββββββ€
β Avg loop iterations per generator β 2.2 β
βββββββββββββββββββββββββββββββββββββββββββββββββββΌβββββββββββββββββββ€
β Total generator_send() calls (GIL acquisitions) β ~4.1 million β
βββββββββββββββββββββββββββββββββββββββββββββββββββΌβββββββββββββββββββ€
β Wall time for all of this β 111 seconds β
βββββββββββββββββββββββββββββββββββββββββββββββββββΌβββββββββββββββββββ€
β Time per GIL round-trip β ~27 microseconds β
βββββββββββββββββββββββββββββββββββββββββββββββββββ΄βββββββββββββββββββ
27 Β΅s per GIL round-trip is not slow β it's actually reasonable for PythonβRust FFI. The problem is purely volume: 4.1 million GIL round-trips, each doing Python::attach() β generator.send() β parse response β dispatch.
Batching at the Python level didn't help because the bottleneck isn't the 53K top-level concurrently() call β it's the 1.86M nested generators those calls spawn. Whether you submit 53K or 500 at a time, each one still creates ~35 child
generators.
The graph mutex isn't the bottleneck β confirmed by batching showing no improvement.
This IS fixable in Pants core, but the fix needs to reduce the number of generator round-trips, not reduce contention. Options:
1. Port resolve_dependencies to Rust β eliminate the Python generators entirely for cached results (Benjy mentioned he enjoys Rust porting)
2. Add a fast-path in generate() β if the graph already has a cached result for this exact node, return it without creating the full generator chain
3. Batch the dep inference β instead of per-target generators, process groups of targets in a single Python callwide-midnight-78598
04/08/2026, 6:48 PMgentle-flower-25372
04/08/2026, 6:48 PMgentle-flower-25372
04/08/2026, 6:48 PM``` I instrumented the Rust engine (src/rust/engine/src/nodes/task.rs generate() function) on upstream main (2.32.0.dev7) to trace what actually happens during the map_addresses_to_dependents call on our 53K-target repo. Here's what I found.
It's not 53K operations β it's 1.86 million.
Each of the 52,927 resolve_dependencies calls is a Python async generator that spawns nested generators for resolve_target, determine_explicitly_provided_dependencies, infer_dependencies (per inference type), resolve_target_parametrizations,
validate_dependencies, etc. The Rust engine executes each as a separate generator via gen_generator().
Top-level resolve_dependencies calls: 52,927
Total generators executed: 1,860,000 (~35 per resolve_dependencies)
Avg loop iterations per generator: 2.2
Total generator_send() calls: ~4,100,000 (each acquires GIL via Python::attach)
Wall time: 111.1s
Cost per GIL round-trip: ~27Β΅s
27Β΅s per round-trip is reasonable for PythonβRust FFI. The problem is pure volume: 4.1 million Python::attach() β generator.send() β parse response β dispatch cycles.
Batching the top-level concurrently() at the Python level doesn't help β I tested batch sizes of 53K, 5K, 500, 50, and 1 (fully sequential). All within noise (~2m30-2m44s). This makes sense: the bottleneck isn't the 53K concurrent futures
contending on something. It's the 1.86M nested generators those futures spawn, regardless of how you batch the outer call.
Batch Size Wall Time
53K (stock) 2m40s
5,000 2m44s
500 2m38s
50 2m31s
1 (seq) 2m27s
The per-call latency distribution confirms this:
Progress Avg/call p50/call
10K calls 309ms 119ms
20K 219ms 120ms
30K 3,564ms 139ms
40K 21,571ms 178ms
50K 34,824ms 382ms
The p50 stays low (individual calls are fast), but the average explodes as later calls include wait time for all concurrent work. The sum of all per-call wall times is ~2M seconds across 111s of wall time β consistent with ~16,500 calls
in-flight simultaneously.
Where the time goes:
- resolve_all (the concurrently() call): 111.1s
- build_reverse (pure Python dict construction): 0.1s
The graph cache IS working β entry.get_node_result() returns Completed results via future::ready() for cached nodes. But even for a cached result, the engine still: creates a Python coroutine β sends to it β gets back a rule call β looks up
the graph node β finds it cached β returns the value β sends back to the coroutine β which yields the final result. That's the 27Β΅s overhead, and it happens 4.1M times.
The fix would need to reduce the number of generator round-trips for cached results β either by short-circuiting in the Rust generate() function when a node is already cached, or by porting the hot-path rules to native Rust to eliminate the
Python generator protocol entirely.```
wide-midnight-78598
04/08/2026, 6:49 PMgentle-flower-25372
04/08/2026, 6:49 PMwide-midnight-78598
04/08/2026, 6:49 PMgentle-flower-25372
04/08/2026, 6:50 PMwide-midnight-78598
04/08/2026, 6:52 PMgentle-flower-25372
04/08/2026, 6:57 PMgentle-flower-25372
04/08/2026, 6:57 PMwide-midnight-78598
04/08/2026, 7:01 PMgentle-flower-25372
04/08/2026, 7:40 PMcurved-manchester-66006
04/08/2026, 10:54 PMhappy-kitchen-89482
04/08/2026, 11:14 PMcurved-manchester-66006
04/09/2026, 1:03 AM$ time pants --pantsd-max-memory-usage=28GiB --changed-since=HEAD~1 --changed-dependents=transitive list
21:00:06.82 [INFO] Initializing scheduler...
21:00:06.95 [INFO] Scheduler initialized.
21:01:23.10 [WARN] No targets were matched in goal `list`.
real 1m19.369s
user 0m0.731s
sys 0m0.211s
$ time pants --pantsd-max-memory-usage=28GiB --changed-since=HEAD~1 --changed-dependents=transitive list
21:01:59.74 [WARN] No targets were matched in goal `list`.
real 0m13.284s
user 0m0.005s
sys 0m0.013shappy-kitchen-89482
04/09/2026, 1:07 AMgentle-flower-25372
04/09/2026, 1:17 AMgentle-flower-25372
04/09/2026, 1:17 AMwide-midnight-78598
04/09/2026, 1:25 AMwide-midnight-78598
04/09/2026, 1:26 AMwide-midnight-78598
04/09/2026, 1:29 AMtest calls run this way?gentle-flower-25372
04/09/2026, 1:51 AMwide-midnight-78598
04/09/2026, 2:00 AM