I want to understand how're people currently manag...
# general
a
I want to understand how're people currently managing coverage reports with pants sharding? 1, Is there an accurate way to combine coverage reports from each shard? ( given that the reports may have overlapping files ) 2, Whether it is more feasible to being use global_reports / fail-under coverage when using sharding, since I'm more interested with the final coverage the total? 3, Any pants native way that supports this currently? or any example?
I'm also noticing that the CI caching seems to be suffering when I use sharding
c
The
pants
repo itself uses
<http://coveralls.io|coveralls.io>
https://coveralls.io/github/pantsbuild/pants?branch=main In that case, coveralls handles aggregation. I'd be curios for a design that works well with
--changed-since
, since that can be harder to reason about.
a
This is very informative, thanks. My solution was to upload the "raw" .coverage binary output with github actions from each shard run of
pants test --changed-since=origin/${{ github.base_ref }} --changed-dependents=transitive --shard=$((${{ matrix.shard }}))
Then download and merge them based on this logic • Combine all shards via
coverage combine
using
3rdparty/.coveragerc
• Exports a merged
coverage-report/coverage.xml
• Fallback chain if no binary
.coverage
files exist: Try using a shard's
coverage.xml
directly • If nothing at all, generate a minimal empty
coverage.xml
via Python inline script (so downstream steps don't fail on a missing file e.g coverage thresholds) I plan on adding diff-cover later down the line when I would want coverage on new/edited lines of code only but for now this setup works well for us Have added approach as a PR if you're interested in reviewing or suggesting