Hello. I inherited a codebase that is using Pants...
# general
b
Hello. I inherited a codebase that is using Pantsbuild and I am trying to figure it out. It appears the previous users didn't understand it very well. Some problems I'm seeing so far: 1. They have protobuf files that come from a submodule (separate repo that is not Pants Build). Instead of using protobuf_sources they just manually build all the proto bindings in Python, move this to a different folder, and then use that python code in a python_source target. So, we have a versioning control issue where we have to manually sync and rebuild constantly. 2. They defined resolve files for every single project, and then defined full requirement files for every single project, and it's led to this thing where almost everything has a resolve that is parameterized on 4 or 5 resolves, and libraries for some reason depend on the code that imports them, just a nightmare. 3. This seems to have led to surprisingly long "generate-lockfiles" runs, like 3+ minutes. I'm unclear if this is normal. So, first, am I right that the above two scenarios are not how Pantsbuild is to be used? I am reading the docs, I have used Bazel before and I just want a sanity check that things like above are not intended. These are fairly high level questions. I've been pounding my head at this for several days now, and so I just want to make sure that I'm not crazy that this seems wrong. The second set of question: We will have this protobuf submodule in a lot of places. I want to build a test project in our repo that somehow sets it up as a protobuf target in some codebase and uses them. Basically, to build a properly isolated Pantsbuild repo and see how it works without the pretty serious dependency issues that the system currently has. So, default resolve. No one calling it. Just a simple project of pulling in some protos, then some simple src/ that uses them and then some /test that tests it. So the two questions here: 1. What is the recommended way to use a remote project like this proto library. The library does not have BUILD files, and will not. I am not allowed to alter that code. 2. What is the generic recommended file tree for a project associated with a single BUILD file?
tl;dr First set of questions. I just want a sanity check on whether the previous person used Pantsbuild incorrectly or I am misunderstanding something. Second set of question is more concrete.
e
I don't have a lot of experience with protobuf, but I would think the best way to use it here is to define a BUILD file that lives within your own repo, but describes the submodule. ie.
Copy code
protobuf_sources(
    sources=["path/to/submodule/**/*.proto"]
)
You can then use
overrides
(or separate
protobuf_sources
target generators with more specific
sources
globs) if you need to customize any metadata. This ought to make it work fairly smoothly within your monorepo, despite not being instrumented with pants tooling natively.
Thoughts on repo structure/file tree: check out a thread I commented on here. And on multiple resolves here (tl;dr is to try really hard to see if you can use a single global lockfile/resolve. Generally the only reason not to is if your projects explicitly require incompatible versions of a library)
Apologies if that's a big thought dump. Hopefully you can get some things sorted and have a good appreciation for Pantsbuild
b
No. This is very useful. Thank you.
And for clarity on resolve files. Let us say we have two projects, that have some fatal conflict. Django 3/4, for instance. We would still have the top-level default resolution that handles all the dependencies that are not in conflict, then you would use resolves to narrowly specify Django versions to the code. I ask this, because the way it's currently set up is that every project in the codebase has its own resolve specified, and every resolve has a fully defined requirements file. So, a pretty huge amount of cruft.
e
yeah, if you needed project A with Django 3 and project B with Django 4 you could do something like:
Copy code
python_requirements(
    source="requirements.base.txt",
    resolve=parametrize("projectA", "projectB"),
)

python_requirement(
    requirements=["Django==3.2.8"],
    resolve="projectA",
)

python_requirement(
    requirements=["Django==4.2.24"],
    resolve="projectB",
)
(Not shown:
pants.toml
is the place to declare the resolve itself) Then for any (1st party) code specific to project A or B:
Copy code
python_sources(
    resolve="projectA"
)
but any code that is shared library kind of stuff that must work with both A and B:
Copy code
python_sources(
    resolve=parametrize("projectA", "projectB")
)
which will behind the scenes create two
python_source
targets for each file, (one for each resolve). Ultimately this means that this file will get linted twice, mypy checked twice, and even have its unit tests run twice, to ensure that it works in both resolves that it claims to be compatible with
h
Re your first question: Correct. This is not really how Pants was designed to be used. It would prefer to run code generation itself. See https://www.pantsbuild.org/2.27/docs/python/integrations/protobuf-and-grpc
And similarly, while you can have lots of resolves, it shouldn't be necessary unless your subprojects truly need to have non-compatible requirements.
Ideally you'd have one (or a small number of) lockfiles shared and reused across all projects
That also lets those projects more easily use each other
For your second, specific question: This blog post has an example of Pants shelling out to Bazel in an embedded repo and grabbing build products to use downstream. That sounds a little similar to what you're trying to do?
b
Yes, something like that. Though, now I'm actually just working with the team who owns this protobuf repo to just package it up in some easily consumable way. Sort of just dodge the harder problems, but just treating it as a python package. Probably the best result. As for Pants, it's good to get this feedback that it sounds like it was being used wrong. It seems so to me as well. Any blog posts on how to clean up a messed up Pants repo?
h
I don't think we have anything that specific. All happy repos are alike; each unhappy repo is unhappy in its own way...
There are Pants maintainers who provide consulting services if that is helpful
b
Thank you, I think it will be fine. The code base is new enough that I think I can just kill the projects and start it again in a cleaner fashion.