Does anyone in the pants community accept bounties...
# general
g
Does anyone in the pants community accept bounties for bug fixes? This performance bug is introducing 3-4m delay to our pipelines right now which represents roughly 30-40% of a build pipeline. We use this in our monorepo build to determine which docker containers need to be built.
Copy code
pants --changed-since=<git-ref> --changed-dependents=transitive --filter-target-type=docker_image list
Historical posts I've made in Slack regarding this issue. https://pantsbuild.slack.com/archives/C046T6T9U/p1740007055261349 https://pantsbuild.slack.com/archives/C0D7TNJHL/p1748874243046649
w
There are a few Pants service providers. Myself, @fast-nail-55400 and @happy-kitchen-89482 being a few that I know of
πŸ‘ 1
There are also sponsorship options to consider: https://www.pantsbuild.org/sponsorship
h
I don’t think sponsorship will get a specific bug fixed. But a bounty could work.
I just took a quick look at bug bounty platforms to see if we could set something up easily, but they seem to be more security focused. Also the bounties are typically set by the owners of the code, whereas here they would be set by the users, so not the best fit.
But putting the technicalities aside, I’m sure there is a $ amount that would get someone interested in making this much faster. I could be convinced, for example.
Since porting to Rust is a thing I enjoy
g
Just to circle back here. 1. I was thinking about sponsorship, but like @happy-kitchen-89482 said it didn't seem like that was a means of getting work done, just a way of getting some additional support. 2. @happy-kitchen-89482 have you done work in exchange for money before in the pants world? I'd be interested in having a conversation around what that would look like and figure out a dollar amount that would work.
f
I would suggest just a consulting contract with someone with a Statement of Work with the price set forth. Even with time/materials, there can be hours limits to control costs.
πŸ‘ 1
h
@gentle-flower-25372 I am starting to do this kind of work. Want to chat about it via DM?
πŸ‘ 1
g
For context, I think Claude found the root cause of the performance issue, and it looks like we're bumping up against the limits of pants' architecture. Take that with a grain of salt, though. 50K targets seems to be just genuinely a lot of work. I had claude bang for 3-5 hours trying a bunch of different things and at the end of the day the bottleneck is the call to
map_addresses_to_dependents
for all 50K targets in our monorepo. Introducing a cache has helped dramatically. https://github.com/pantsbuild/pants/pull/23228
c
map_addresses_to_dependents
Interesting. Is that from work unit logs, the new
perf
support, or something else?
g
> Is that from work unit logs, the new
perf
support, or something else? @curved-manchester-66006 What do you mean by "that"? Are you asking how I arrived at the conclusion that the bottleneck was the
map_addresses_to_dependents
? I had claude alter the code to benchmark it with logging, etc.
c
Yes, "that" as the
map_addresses_to_dependents
conclusion
g
I've sicked claude on it like a dog 2-3x over the last 3-6 months and every time it keeps coming back to that being the bottleneck given the sheer volume of targets in our monorepo which is ~50K.
Let me give more concrete summary of the findings so it's clear and concise.
```# Performance Investigation: --changed-dependents=transitive on 53K targets
## The Problem
pants --changed-since=<ref> --changed-dependents=transitive --filter-target-type=docker_image list takes ~3.5 minutes in our monorepo (53K targets), regardless of how few files actually changed. This is ~30-40% of our build pipeline time.
## Root Cause
The bottleneck is map_addresses_to_dependents() in src/python/pants/backend/project_info/dependents.py. When --changed-dependents is used, this rule must build the full reverse dependency graph by calling resolve_dependencies() for every
target in the repo:
@rule(desc="Map all targets to their dependents")
async def map_addresses_to_dependents(all_targets: AllUnexpandedTargets) -> AddressToDependents:
dependencies_per_target = await concurrently(
resolve_dependencies(DependenciesRequest(tgt.get(Dependencies), ...))
for tgt in all_targets # ALL 53K targets
)
This is expensive because each resolve_dependencies() call runs dependency inference (Python import parsing, Docker COPY analysis, etc.). With 53K targets, this takes ~150 seconds of the ~210 second total.
## Why Caching Doesn't Help
- pantsd in-memory cache: AllUnexpandedTargets is https://github.com/pantsbuild/pants/blob/main/src/python/pants/engine/target.py#L1023-L1025. Any filesystem change invalidates it, which cascades to invalidate AddressToDependents. We benchmarked identical commands on warm pantsd β€” still ~3 minutes each time.
- Remote cache: Only caches process execution results (test runs, compilation), not rule engine computations like dependency inference.
- LMDB local store: Same β€” only process results, not rule outputs.
## What We Tried
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ Approach β”‚ Result β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚ Warm pantsd (identical command twice) β”‚ 3m00s β†’ 3m01s (no improvement) β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚ Forward BFS from filtered targets only (https://github.com/pantsbuild/pants/pull/23224) β”‚ 3m39s β†’ 2m42s (26% faster, resolves 24K instead of 53K targets) β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚ Persistent dep cache on disk (https://github.com/pantsbuild/pants/pull/23228) β”‚ 3m22s β†’ 43s (5x faster) β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
## Our Solution: Persistent Dependency Inference Cache
https://github.com/pantsbuild/pants/pull/23228 adds an opt-in --incremental-dependents-enabled flag that persists the forward dependency graph to ~/.cache/pants/incremental_dep_graph_v2.json (29MB, 1.3MB compressed). On subsequent runs, only targets with changed source files (by SHA-256 content hash) have their dependencies re-resolved.
Cold (no cache): 3m22s β€” 52,927 targets resolved, writes cache
Warm (from cache): 43s β€” 0 targets resolved, 52,927 from cache
The cache is portable across machines (SHA-256 content hashing, not mtime), so it works on ephemeral CI agents via S3.
## The Architectural Gap
The fundamental issue is that Pants' rule engine caches rule results in-memory only and has all-or-nothing invalidation β€” there's no concept of "re-resolve deps for 50 changed targets and reuse the other 52,877 from last run." map_addresses_to_dependents either returns a fully cached result or recomputes from scratch for all targets.
A native solution would be persistent rule-level caching in the engine (extending LMDB to cache resolve_dependencies results per-target, keyed by content hash). Our PR approximates this at the Python level.```
h
Claude's top-level conclusion that this function is the issue is surely correct, given the timing data, but its comprehension of the moving parts is very superficial and largely wrong. For starters, dep inference results are already cached (in Rust), per file. So Pants is not reparsing imports or whatever. And even without that, in the case of two consecutive runs with a warm Pantsd, every rule call in the second run should be resolved from the in-memory cache, so "persistent rule-level caching" would not help here (assuming that pantsd has not restarted between the runs - can you confirm?) So, Claude's first paragraph under "## The Architectural Gap" is basically false. Furthermore, the bespoke cache solution is naive - it doesn't integrate with the existing caching mechanism, it assumes that the content of the file is the only thing that should be in the cache key (what about all the options that modify how deps are inferred?), it relies on some out of band env var, not a proper pants option, and it's not properly robust in the face of concurrent runs (one will stomp the other's results, even if the tmpfile means that you won't get an actually corrupt file). So this is not the way to go. The more likely underlying issue is the thundering herd of 50k concurrent calls to
resolve_dependencies()
overwhelming the scheduler, even though they end up being resolved from cache. The real solution is to fix that (after first proving that it is indeed the issue).
@curved-manchester-66006 you had a public "large fake repo" for testing this sort of thing on, no?
Or, @gentle-flower-25372 do you have one we can look at?
g
I'm literally working on it now πŸ‘
but yes I got a reproduction of it.
h
That's helpful, thanks
πŸ‘ 1
g
Please don't overi-ndex on claude's statement "pants architecture is wrong" πŸ™ He's just a silly old bot.
h
Oh I'm not offended by a robot
That's not the issue
g
goooood πŸ˜„
h
the issue is that said robot is very confident even when very wrong
πŸ’― 1
g
I'm also not asserting that either. I don't have enough smarts to have an opinion. This is way out of my league.
h
so we can't take his word for it
g
he's a bad boy for sure.
w
I saw that PR opened and I saw the claudsponse and I thought the same. The result being correct vs the way it got there, I wasn’t sure of. We’ve identified a lot of instances where the nature of the dependency mapping (and everything associated with that) is causing a lot of overhead, and there are a few PRs to address that (I think in one instance, JS dep inference got 20x faster) So yeah, the problem is valid, the solution is… less so.
g
claude == stupid fake "intelligence" that's always 100% they're right humans == smart
curl said their credible reports for security vulns has dropped from 15% to 5% 😒
w
Yeah, they dropped a lot of their bug bounty, but some of that was just people throwing anything over the wall - which is rough. Anyways, there are a few in-flight PRs (or upcoming) moving some more intrinsics to Rust, and I believe the intention is to expand how dep inference can be done in some backends, which would help with anything inference related
g
If true, that would be amazing. This has been a really big performance killer for us over the last year. We've mostly just dealt with it and in a few cases we've actually written small static code analysis tools to bypass calls to pants.
At the very least a solid reproduction came out of this exercise. I was struggling before of how to do that.
I'm going to push up the reproduction to a git repo now.
πŸ‘ 2
w
I have some of the same issues - I’ve written a lot of tooling that kinda bypasses Pants for certain operations, and then calls back in for other ones. Not that my repos are big, I’m just inherently impatient and if anything takes longer than a couple seconds, I’ll certainly be looking at youtube by then
g
if anything takes longer than a couple seconds, I’ll certainly be looking at youtube by then
100%. Honestly most of our users don't ever use pants natively because this has been such a big pain point. We've written a lot of workarounds. Mostly exporting pants virtualenv and magic around that.
We essentially only use pants in CI for builds and testing.
w
So, this is running without the pants daemon. What happens with it enabled?
g
@wide-midnight-78598 The repro documents with and without. The daemon doesn't help.
More output from claude based on pushback/feedback here.
```Profiling Data: map_addresses_to_dependents with 53K Targets
I instrumented resolve_dependencies and map_addresses_to_dependents on upstream main (2.32.0.dev7) to understand where the time goes. No changes to caching or batching logic β€” just timing measurements on the stock code.
Setup
- 52,927 targets (20,850 python_source, 15,422 file, 8,620 resource, 2,945 python_test, plus Docker, Shell, etc.)
- Command: pants --no-pantsd --changed-since=HEAD~3 --changed-dependents=transitive --tag="-integration" --filter-target-type="+python_test" filter
- Timing added to each resolve_dependencies call and to map_addresses_to_dependents phases
Results
map_addresses_to_dependents phase breakdown:
- Resolve all dependencies (the concurrently() call): 122.5s
- Build reverse map (pure Python dict construction): 0.1s
So 99.9% of the time is in the concurrently(resolve_dependencies(...) for tgt in all_targets) call.
Per-call resolve_dependencies stats (52,927 calls):
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ Metric β”‚ Value β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚ p50 (median) β”‚ 13.4s β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚ Average β”‚ 38.4s β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚ p99 β”‚ 112.7s β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚ Max β”‚ 118.2s β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚ Sum of all calls β”‚ 2,030,764s β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚ Wall time β”‚ 122.5s β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚ Effective parallelism β”‚ ~16,500x β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
The sum-of-all-durations is ~2 million seconds but wall time is only 122s, which is consistent with ~16K+ calls in flight concurrently (as expected from try_join_all on 53K futures).
The per-call latency grows over the course of the run:
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ Progress β”‚ Avg per call β”‚ p50 per call β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚ First 10K calls β”‚ 309ms β”‚ 119ms β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚ At 20K β”‚ 219ms β”‚ 120ms β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚ At 30K β”‚ 3,564ms β”‚ 139ms β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚ At 40K β”‚ 21,571ms β”‚ 178ms β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚ At 50K β”‚ 34,824ms β”‚ 382ms β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
The median stays relatively low (119ms β†’ 382ms), but the average and tail latency explode as more calls are in flight β€” calls submitted later wait for the full duration of all earlier calls.
Batching experiment
I also tested whether batching the concurrently() call helps reduce contention:
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ Batch Size β”‚ Wall Time β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚ 53K (unbatched, stock) β”‚ 2m40s β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚ 5,000 β”‚ 2m44s β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚ 500 β”‚ 2m38s β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚ 50 β”‚ 2m31s β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚ 1 (fully sequential) β”‚ 2m27s β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
Batching at the Python level doesn't meaningfully help. The overhead is below the concurrently() call β€” inside the engine's task scheduling, graph node management, and Python↔Rust generator protocol for each of the 53K calls.
Interpretation
The per-call work is fast individually (p50 ~120ms), but 53K calls through the engine's concurrent task infrastructure accumulates significant overhead. The reverse map construction itself is trivial (0.1s). The entire cost is the engine processing 53K resolve_dependencies calls, even when the underlying dep inference results are cached.
I don't have a theory for what specifically in the engine is slow at this scale β€” it could be task scheduling overhead, the generator send/receive protocol, workunit management, or something else in the Rust engine. But the data clearly shows the cost is in the aggregate overhead of 53K concurrent rule calls, not in the dep inference computation itself.```
This feels like more BS
w
I mean, the discussion about thousands of in-flight requests was what we were talking about earlier. Async task to/from Rust. Not saying it’s the case here, but it certainly is problematic.
πŸ‘€ 1
g
I spanked claude and seems like we have some real info for the first time
Copy code
● Here's what the data actually says:

  The Real Root Cause

  It's not 53K operations. It's 1.86 million.

  Each of the 53K resolve_dependencies calls is a Python async generator that spawns nested generators β€” for resolve_target, determine_explicitly_provided_dependencies, infer_dependencies (per inference type), validate_dependencies, etc. The
  Rust engine runs each as a separate generator via gen_generator().

  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
  β”‚                     Metric                      β”‚      Value       β”‚
  β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
  β”‚ Top-level resolve_dependencies calls            β”‚ 52,927           β”‚
  β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
  β”‚ Total generators executed                       β”‚ 1,860,000        β”‚
  β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
  β”‚ Avg generators per resolve_dependencies         β”‚ ~35              β”‚
  β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
  β”‚ Avg loop iterations per generator               β”‚ 2.2              β”‚
  β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
  β”‚ Total generator_send() calls (GIL acquisitions) β”‚ ~4.1 million     β”‚
  β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
  β”‚ Wall time for all of this                       β”‚ 111 seconds      β”‚
  β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
  β”‚ Time per GIL round-trip                         β”‚ ~27 microseconds β”‚
  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

  27 Β΅s per GIL round-trip is not slow β€” it's actually reasonable for Python↔Rust FFI. The problem is purely volume: 4.1 million GIL round-trips, each doing Python::attach() β†’ generator.send() β†’ parse response β†’ dispatch.

  Batching at the Python level didn't help because the bottleneck isn't the 53K top-level concurrently() call β€” it's the 1.86M nested generators those calls spawn. Whether you submit 53K or 500 at a time, each one still creates ~35 child
  generators.

  The graph mutex isn't the bottleneck β€” confirmed by batching showing no improvement.

  This IS fixable in Pants core, but the fix needs to reduce the number of generator round-trips, not reduce contention. Options:
  1. Port resolve_dependencies to Rust β€” eliminate the Python generators entirely for cached results (Benjy mentioned he enjoys Rust porting)
  2. Add a fast-path in generate() β€” if the graph already has a cached result for this exact node, return it without creating the full generator chain
  3. Batch the dep inference β€” instead of per-target generators, process groups of targets in a single Python call
w
🀷 Maybe. All of the numbers sound weird, but I’d buy the overall premise - as it lands on what we’d already been discussing.
g
The bottom 3 might be BS, but the rest looks decent.
A simple summary from claude
``` I instrumented the Rust engine (src/rust/engine/src/nodes/task.rs generate() function) on upstream main (2.32.0.dev7) to trace what actually happens during the map_addresses_to_dependents call on our 53K-target repo. Here's what I found.
It's not 53K operations β€” it's 1.86 million.
Each of the 52,927 resolve_dependencies calls is a Python async generator that spawns nested generators for resolve_target, determine_explicitly_provided_dependencies, infer_dependencies (per inference type), resolve_target_parametrizations,
validate_dependencies, etc. The Rust engine executes each as a separate generator via gen_generator().
Top-level resolve_dependencies calls: 52,927
Total generators executed: 1,860,000 (~35 per resolve_dependencies)
Avg loop iterations per generator: 2.2
Total generator_send() calls: ~4,100,000 (each acquires GIL via Python::attach)
Wall time: 111.1s
Cost per GIL round-trip: ~27Β΅s
27Β΅s per round-trip is reasonable for Python↔Rust FFI. The problem is pure volume: 4.1 million Python::attach() β†’ generator.send() β†’ parse response β†’ dispatch cycles.
Batching the top-level concurrently() at the Python level doesn't help β€” I tested batch sizes of 53K, 5K, 500, 50, and 1 (fully sequential). All within noise (~2m30-2m44s). This makes sense: the bottleneck isn't the 53K concurrent futures
contending on something. It's the 1.86M nested generators those futures spawn, regardless of how you batch the outer call.
Batch Size Wall Time
53K (stock) 2m40s
5,000 2m44s
500 2m38s
50 2m31s
1 (seq) 2m27s
The per-call latency distribution confirms this:
Progress Avg/call p50/call
10K calls 309ms 119ms
20K 219ms 120ms
30K 3,564ms 139ms
40K 21,571ms 178ms
50K 34,824ms 382ms
The p50 stays low (individual calls are fast), but the average explodes as later calls include wait time for all concurrent work. The sum of all per-call wall times is ~2M seconds across 111s of wall time β€” consistent with ~16,500 calls
in-flight simultaneously.
Where the time goes:
- resolve_all (the concurrently() call): 111.1s
- build_reverse (pure Python dict construction): 0.1s
The graph cache IS working β€” entry.get_node_result() returns Completed results via future::ready() for cached nodes. But even for a cached result, the engine still: creates a Python coroutine β†’ sends to it β†’ gets back a rule call β†’ looks up
the graph node β†’ finds it cached β†’ returns the value β†’ sends back to the coroutine β†’ which yields the final result. That's the 27Β΅s overhead, and it happens 4.1M times.
The fix would need to reduce the number of generator round-trips for cached results β€” either by short-circuiting in the Rust generate() function when a node is already cached, or by porting the hot-path rules to native Rust to eliminate the
Python generator protocol entirely.```
w
I just spotted the cold vs warm cache. That’s more than a little surprising though
g
Yeah I agree. But if the rest is true, the coroutine and GIL pieces, then it makes sense.
w
Also, fwiw, I find the Claude outputs in this thread relatively useless. Lots of words, not saying a lot we don’t really know.
πŸ‘ 1
g
heard. I'll stop dumping.
w
The real one I’m interested in right now is why cold vs warm is identical - I haven’t had a chance to try the reproduction yet, though. Generally, I don’t see similar problems where touching 1 file re-runs the world, but there might be a memory limit being hit, where if you have pants daemon running, and we’re using up all available memory, it starts from scratch. I recall that was an issue a while back that I wasn’t able to reproduce
g
so I do see memory hitting 32GB (limit of host I'm testing on). I was literally watching memory using htop and it keeps filling until it hits 32GB of memory.
but it's surprising it would need more than 32GB of memory in the reproduction. That said, I don't know what would be cached so I shouldn't have an opinion.
w
Yeah, hitting the memory ceiling triggering this would make sense. I’d have to dig up the notes around this - it was a while ago. The memoization structures when there are a lot of targets can add up.
g
I think @happy-kitchen-89482 mentioned something similar a couple of months ago, maybe more.
c
Tried to consolidate prior discussions and dump my notes in https://github.com/pantsbuild/pants/issues/23236
h
How certain are you that pantsd doesn't help? When I run locally I see pantsd restart every run, presumably due to memory overconsumption
c
What I get:
Copy code
$ time pants --pantsd-max-memory-usage=28GiB --changed-since=HEAD~1 --changed-dependents=transitive list
21:00:06.82 [INFO] Initializing scheduler...
21:00:06.95 [INFO] Scheduler initialized.
21:01:23.10 [WARN] No targets were matched in goal `list`.

real    1m19.369s
user    0m0.731s
sys     0m0.211s

$ time pants --pantsd-max-memory-usage=28GiB --changed-since=HEAD~1 --changed-dependents=transitive list
21:01:59.74 [WARN] No targets were matched in goal `list`.

real    0m13.284s
user    0m0.005s
sys     0m0.013s
☝️ 1
h
Yeah, that makes more sense to me...
g
Ah so if we set a memory limit the cache will help?
So maybe it's just the cache efficiency bug? πŸ›
w
This all makes much more sense than what I was reading in that repro re: warm cache. This is flagged in our CI docs as something to consider: https://www.pantsbuild.org/stable/docs/using-pants/using-pants-in-ci#tuning-resource-consumption-advanced https://www.pantsbuild.org/stable/reference/global-options#pantsd_max_memory_usage Nonetheless, that much RAM... In this economy?
@gentle-flower-25372 What's the workflow you have where you run this so frequently?
Or is this just an alias you have setup, so all your (e.g.)
test
calls run this way?
g
We call pants dependencies in ci to detect which docker images need to be rebuilt vs retagged. We also use this to generate a list of pants python_test targets we need to run. We implemented custom sharding because the pants algo leaves many shards that are empty.
πŸ‘ 1
w
Right, that makes sense. I can’t recall if dependencies work under visibility rules - I was looking at this last week, and just forgot. As in, whether visibility rules constrain the β€œworld” of dependencies, to limit it