Hey everybody, I have a question about including t...
# general
g
Hey everybody, I have a question about including tooling in a monorepo. I'm coming from
please
, where managing arbitrary build tools was a first-class pattern. If I needed terraform at a specific version, I'd write a genrule that downloads the binary, makes it executable, and exposes it as a named target. Any other target could then declare a dependency on //tools:terraform and get exactly that version, hermetically, without anything needing to be installed on the developer's machine. We ended up encapsulating all tooling this way, multiple Terraform versions, the AWS CLI, custom binaries and it worked cleanly across the monorepo. Looking at Pants, I've found file with http_source, which handles the download and verification story well. But I'm not clear on how to get from "I have a downloaded binary as a file target" to "any shell_command or adhoc_tool can use it as an executable on PATH without requiring a local install." In Please the genrule output was directly consumable as a tool dependency. In pants it feels like I have the primitives but not the abstraction, I'd need to chmod +x it somewhere, figure out how to get it onto PATH in the sandbox, and repeat that ceremony for every tool. Is there a clean pattern for this that I'm missing? Or is the intended path to write a plugin for each tool that needs proper management?
w
Hmm, I want to say that I’ve seen an example of this - but it may have just been downloading resources that are USED by shell/adhoc…
Not ones that drive it… Like, you could download these files, and call them in a shell script. That’s possible, though a bit meh. The plugin approach is always there, but it should be the last approach, not the first, or even second.
I think you’re describing this workflow (https://www.pantsbuild.org/stable/docs/ad-hoc-tools/integrating-new-tools-without-plugins#using-externally-managed-tools ) but not a
system_binary
but rather a downloaded one… Just before I let myself get nerd sniped, this is a workflow where you’ll have multiple instances of the same tool, in different parts of your code?
g
Do people still nerd snipe instead of being helpful? How 2010 of them. My usage pattern in my previous repositories was such that the user had to have basically nothing installed (beyond please and maybe the basic Unix core utilities). That means anything we would use like jq, aws cli, anything, was included as a “tool”. In please that could mean a simple download of a binary that ends up with an output of that binary to a generic gen rule that declares an entry point (executable to call). Our most common use case for these tools was certainly run time not build time, but some did end up getting used at build time. For example a build step that outputs a configuration file that was credited with some shell commands and jq. Does this clear it up at all? I use awscli as an example because it feels easy to wrap one’s head around. Here's a random example. Build rule take constructs an aws profile config. Then we have aws CLI as a tool. Then we also have a genrule that builds a specific aws CLI target with the correct config (defined in a build file). This means all run time commands for the domain area are pre wrapped with the proper auth tools. Is this helpful? I may have simply not answered the question or become more confusing
i think maybe ive come to the conclusion that the only real difference is that please has
entrypoint
w
I think
jq
was a helpful example, as I think that’s a pretty clear use case. I’ll try to take a look at this, because it was something I was wondering about too for some other tooling. I have standalone downloaders I use to do these, outside of Pants. In Pants, we have python-build-standalone that can grab python installations. But I don’t know if we have an exposed capability to use our
externaltool
template downloads
g
I did this a LOT in my please monorepo
let me see if i can dig up an example?
w
Yeah, I get the idea overall though. If we don’t have it out of the box, it feels like a good thing to have in more places
g
I think its fair to say we were using a build system to do more than just build in many ways but it worked so well. it was amazing.
this is from the
genrule
docs for please
Copy code
entry_points		dict	A subset of outputs of this rule that can be used as entry points by other rules. Entry points can be referenced though the `//path/to:rule|entry-point` syntax.
w
Like, fundamentally,
http_source
allows the downloading part, but combining that with other tooling we have feels like a smart idea. I thought we had that already (e.g. instead of
system_binary
, some sort of
downloadable_binary
idea). Someone else might be able to chime in if we natively support that already
g
this is an example of something we did, maybe not the best example but its concise
Copy code
# tools/terraform/0.14.5/BUILD
  remote_file(
    name = "terraform",
    url = "<https://releases.hashicorp.com/terraform/0.14.5/terraform_0.14.5_darwin_amd64.zip>",
    hashes = ["2edf2491d3b..."],
    extract = True,
    entry_points = {"bin": "terraform-0.14.5-darwin-amd64/terraform"}
  )

  # any BUILD file anywhere
  genrule(
    name = "plan",
    tools = {"tf": ["//tools/terraform/0.14.5:terraform|bin"]},
    cmd = "$TOOLS_TF plan ...",
  )
w
The capability as an in-repo plugin would be pretty simple I think… I’ll take a closer look later to see if we support this natively
g
Thanks. I really appreciate it. I just want to make sure im not trying to bang a square into a circle hole by taking on pants as the build system for what i want to do
w
It sounds like it does most of what you want, with the uniqueness part being: • Multiple tools where we currently might only expect 1 • Arbitrary tool downloading (which is absolutely something we do already, via plugins)
g
I dont know on the feelings of AI here but i asked it to crawl the old code base for examples and explain what i was doing that pants may not support well.
Copy code
A few patterns stand out as genuinely hard to replicate in Pants:

  Auth-baked wrappers — The SSO toolkit generates per-account AWS CLI wrappers where credentials and profile config are baked in at build time. You'd call //tools/aws/sso:aws-myaccount and auth just works. In
  Pants you could download the CLI but generating environment-specific wrappers that depend on other build outputs isn't something shell_command can really express cleanly.

  Tool composition via bash_script — There's a bash_script() build_def that takes a dict of tools and exposes each as $ASSET_<KEY> inside the script. This let them compose multiple tools (aws + jq +
  session-manager) into a single runnable artifact with proper dependency ordering. The EMR tooling uses this heavily — ssh, launch, port-forward all built from composed tools.

  Build-time code generation — pyspark_apps_packager and file_factory_templates use custom downloaded tools at build time to produce artifacts. The tool runs during the build graph, not at runtime. Pants has
  adhoc_tool for this but it requires the binary to already be on the system — you can't easily say "download this tool, then use it to generate this artifact."

  Multi-version tool selection — Terraform targets could specify terraform_tool_overrides={"init": "//tools/terraform/0.13", "plan": "//tools/terraform/0.14"} so different phases use different versions. Since
  every tool reference was just a target label, this was essentially free.

  The auth-baked wrapper pattern is probably the most practically useful one that Pants genuinely can't do today without significant workarounds.
without using AI i had come to the same conclusion about the aws sso wrapper thing we were doing, so take that for what its worth
w
I think there was something added to the codebase, to make adding linters basically a few lines of code as a plugin. But, not sure if that’s a 1:1 use case. For the http_source, I don’t know if it will do stuff like unpack tarballs - just can’t recall
f
I think there was something added to the codebase, to make adding linters basically a few lines of code as a plugin. But, not sure if that’s a 1:1 use case.
code_quality_tool
?
(And which is much more useful now that
shell_command
is runnable.)
s
I have no idea if what we did is a best practise or not, but this is the pattern we've been using: •
3rdparty/tools/protoc/BUILD
has these for downloading & unzipping protobuf compiler:
Copy code
# protobuf compiler (protoc & related files)

file(
    name="compressed-protoc",
    source=per_platform(
        linux_arm64=http_source(
            url="<https://github.com/protocolbuffers/protobuf/releases/download/v29.1/protoc-29.1-linux-aarch_64.zip>",
            len=3257573,
            sha256="1f74a3f3355de7c0666bc125611c13532c2598f853521d0d3e621a5b09f24799",
            filename="protoc.zip",
        ),
        linux_x86_64=http_source(
            url="<https://github.com/protocolbuffers/protobuf/releases/download/v29.1/protoc-29.1-linux-x86_64.zip>",
            len=3288942,
            sha256="00c83fe9722d85e96c81b941b29f17a744b33b4ce66e0f18009fd8937de22c60",
            filename="protoc.zip",
        ),
        macos_arm64=http_source(
            url="<https://github.com/protocolbuffers/protobuf/releases/download/v29.1/protoc-29.1-osx-aarch_64.zip>",
            len=2290879,
            sha256="b8fd5976926198a7c4ea5c6eb4bf78959d5faed27bfc618254caa1043f770445",
            filename="protoc.zip",
        ),
    ),
)

shell_command(
    name="protoc",
    command="unzip protoc.zip",
    tools=["unzip"],
    execution_dependencies=[":compressed-protoc"],
    output_directories=["./bin", "./include"],
)
• and then using it in our protobuf directory's BUILD file to generate a protobuf descriptor set file:
Copy code
shell_command(
    name="http_descriptor_set",
    command="""
    {chroot}/3rdparty/tools/protoc/bin/protoc \\
      --proto_path . \\
      -I {chroot}/3rdparty/tools/protoc/include \\
      --include_imports \\
      --descriptor_set_out=http_descriptor_set.binpb \\
      $(find kiid/http -name '*.proto') \\
      $(find google/rpc -name '*.proto')
    """,
    tools=["find"],
    execution_dependencies=["3rdparty/tools/protoc:protoc", ":http_protobufs"],
    output_files=["http_descriptor_set.binpb"],
)
Not sure about how to get
protoc
on the path in the above example tho...
f
put the command in a shell script along with a
export PATH=.....
, and just have the
shell_command
invoke the shell script?