<#23648 Design doc: resolving codegen ambiguity wh...
# github-notifications
q
#23648 Design doc: resolving codegen ambiguity when one generator target serves multiple consumer languages Issue created by robertpi Design doc: resolving codegen ambiguity when one generator target serves multiple consumer languages Status: draft, for discussion Related: #23619 Author: robertpi Problem A single, unparametrized
protobuf_sources()
target cannot currently be depended on by both a
java_sources
target and a
scala_sources
target at the same time. Pants raises an ambiguity error, because two independent pieces of dispatch logic can't tell which generated language a given consumer wants: 1. Classpath-entry dispatch.
ClasspathEntryRequestFactory.classify_impl
(
jvm/compile.py
) has to pick one
ClasspathEntryRequest
implementation (
CompileJavaSourceRequest
or
CompileScalaSourceRequest
) to produce a classpath entry for the shared
protobuf_sources
component. With two JVM languages both consuming it, both impls claim it, and
classify_impl
raises
ClasspathSourceAmbiguity
. 2. Source-hydration/codegen dispatch.
hydrate_sources
(
engine/internals/graph.py
) has to pick one
GenerateSourcesRequest
implementation (
GenerateJavaFromProtobufRequest
or
GenerateScalaFromProtobufRequest
) when a compiler hydrates the
protobuf_sources
component's own sources. This specifically bites
compile_scala_source
, which requests
for_sources_types=(ScalaSourceField, JavaSourceField)
(scalac supports directly compiling Java files coarsened into the same dependency cycle — "mixed compilation"). Both generators satisfy that request, so
hydrate_sources
raises
AmbiguousCodegenImplementationsException
. Today, the only workaround is to declare two separate
protobuf_sources
targets and manually route each JVM consumer's
dependencies
to the right one — which is exactly the kind of BUILD-file-visible plumbing we'd like to avoid. Goals • A single
protobuf_sources
(or any other multi-language codegen source) target can be depended on by consumers in more than one generated language, without any new required field or additional target declaration. • The mechanism generalizes past protobuf/java/scala: it should work for any pair of `GenerateSourcesRequest`/`ClasspathEntryRequest` implementations that both claim the same input when consumed by different concrete requesters. • No change to dependency-inference output or
BUILD
file syntax for the common case. Non-goals • Determining codegen output for a root target (i.e., a
protobuf_sources
target passed directly to `test`/`check`/`compile` on the command line, with no requester). This case has no "consumer" to infer a preference from, and is out of scope — protobuf targets are not meaningful compile/check roots today, so this doesn't arise in practice. • Changing how codegen is registered (the `GenerateSourcesRequest`/`ClasspathEntryRequest`
@union
mechanisms) — this proposal is about dispatch, not the plugin registration model. Background: today's dispatch flow For a JVM compile of
scala_sources(dependencies=["//protos"])
where
protos
is `protobuf_sources()`: •
compile_scala_source
requests classpath entries for its dependencies via
classpath_dependency_requests
→
ClasspathEntryRequestFactory.for_targets(component=protos, ...)
. •
for_targets
looks at every registered
ClasspathEntryRequest
subclass and asks which ones claim
protos
(via `field_sets`/`field_sets_consume_only`). Both
CompileJavaSourceRequest
and
CompileScalaSourceRequest
claim it (protobuf can generate into either), so it's ambiguous — unless something tells it which one to prefer. • Separately,
compile_scala_source
itself calls
determine_source_files(SourceFilesRequest(protos_sources_field, for_sources_types=(ScalaSourceField, JavaSourceField), enable_codegen=True))
to hydrate `protos`'s own generated sources (needed because scalac can be handed a Java file directly, for mixed compilation).
hydrate_sources
collects every registered
GenerateSourcesRequest
whose output type is in
for_sources_types
— again, both Java and Scala protobuf codegen qualify. Both call sites have the same shape: dispatch is unambiguous everywhere except at a shared codegen boundary, and the caller actually has the context needed to disambiguate — it just isn't being passed through. Proposed design: implicit "preferred codegen" resolution Rather than asking the user to declare a preference in the
BUILD
file, infer it from context that already exists at each call site, per dependency edge: 1. Classpath-entry level:
preferred_impl
classpath_dependency_requests
(the code that, for a given compiling component, requests classpath entries for its dependencies) already knows its own
ClasspathEntryRequest
type — e.g. it's running inside
compile_scala_source
, handling a
CompileScalaSourceRequest
. Thread that type down as a
preferred_impl
hint: • `ClasspathEntryRequestFactory.classify_impl`/`for_targets` gain an optional
preferred_impl: type[ClasspathEntryRequest] | None
parameter. • When more than one impl claims a component, and
preferred_impl
is among the claimants, use it instead of raising. • When there's no
preferred_impl
(e.g. resolving a root target with no requester) or the preference doesn't narrow it to exactly one impl, behavior is unchanged — raise
ClasspathSourceAmbiguity
as before. Concretely: a Scala compile's own dependency requests pass `preferred_impl=CompileScalaSourceRequest`; a Java compile's pass
preferred_impl=CompileJavaSourceRequest
. Each compiler gets its own flavor of the shared codegen target's classpath entry, resolved independently per edge. 2. Source-hydration level: preference by
for_sources_types
order
for_sources_types
is a tuple, and its order already carries meaning elsewhere in
hydrate_sources
(
compatible_with_sources_field
walks it in order to decide which type to report as
sources_type
). We extend that same convention to codegen selection: • When more than one registered generator would satisfy an
enable_codegen=True
request, narrow by walking
for_sources_types
in order; take the first entry that exactly one candidate generator's output type subclasses. • If narrowing arrives at exactly one candidate, use it. • If it arrives at zero or more than one at every step (i.e. two generators would produce the same output type — a genuine, unresolvable ambiguity), raise
AmbiguousCodegenImplementationsException
as before. Concretely: `compile_scala_source`'s
for_sources_types=(ScalaSourceField, JavaSourceField)
already expresses "prefer Scala, but accept coarsened Java" — so protobuf's ambiguity resolves to
GenerateScalaFromProtobufRequest
without the caller needing to change anything. Properties • Zero BUILD-file surface. No new field on `protobuf_sources`/`protobuf_source` (or any other codegen target).
python_sources
, which has no reason to know about JVM codegen, is untouched. • Generalizes. Both mechanisms operate on existing types (
ClasspathEntryRequest
,
GenerateSourcesRequest
,
for_sources_types
), not protobuf-specific concepts — any future codegen input shared by two JVM-ish languages gets this behavior for free. • Symmetric and per-edge. A
protobuf_sources
target can be depended on by an arbitrary mix of
java_sources
and
scala_sources
targets simultaneously; each consumer resolves its own preferred codegen independently. • No change to unambiguous cases. Every existing call site that only ever supplies a single relevant generator (e.g. `javac.py`'s
for_sources_types=(JavaSourceField,)
) is untouched — the new logic only activates once ambiguity would otherwise be raised. Prototype status A working prototype of both mechanisms exists (unmerged, pending this design discussion): `jvm/protobuf-classpath-per-consumer-language`, including an end-to-end test (
test_protobuf_consumed_by_java_and_scala
in
jvm/compile_test.py
) proving a single
protobuf_sources
targ… pantsbuild/pants