Hi pants team, I've been looking at setting up a ...
# general
e
Hi pants team, I've been looking at setting up a caching service for my CI runs. Looking at bazel-remote (as referred in the pants docs). I know this is a bit outside of pants itself, but has anyone worked with this before and have a rough idea on cpu/memory requirements for bazel remote? My read is that it doesn't require a whole lot of bulk, especially if its not handling a whole lot requests at a time, but I'd love to confirm that with someone that has some experience. Thanks
f
What is your CI provider / system?
e
Github actions (enterprise, not public github), but we are looking at having our devops team setting up a dedicated instance for the cache server in our AWS account (alongside where our github runners are)
h
FWIW, Pants’s own CI is set up to use S3 as a basic remote cache. It works reasonably well.
e
Wait... I can just skip the cache API and just use a raw S3 bucket? This is very interesting and will investigate tomorrow
f
Pants CI uses bazel-remote in front of S3.
e
...and my devops team will love this because no extra infrastructure
h
Sorry, yes, should have clarified that this is in the context of bazel-remote
but with trivial setup
e
Any thoughts on how to judge the size needed for an s3 bucket?
I got some guidance from my devops team that our runners have a persistent/shared volume mounted to them. If I configure to just use a filesystem for the cache's data store, are there any concerns re: concurrent read/writes from multiple CI runners running at once?
g
By cache data store, do you mean lmdb?
e
a
My own plan for bazel-remote is 1 writer, multiple reader endpoints to avoid any multiple writes to the same bucket. I'll be using GCS instead but this is a thread about S3 linked: https://github.com/buchgr/bazel-remote/issues/483
(and only certain workflows will be writing, most workflows will just be reading. I don't want random developers or PR triggered github-actions to be writing to the build cache. not sure if that is what you also want or need @elegant-florist-94385 ).
Psych I'm going to do what tdyas posted so I don't need any extra infra
g
Fwiw our setup with bazel-remote is a single VM, I believe backed by our ceph cluster for bulk storage. Worked brilliantly with bazel + pants sharing a remote, hovering somewhere in the region 200-300 builds/day. Our config both in bazel and pants has been users read and CI does both. If my storage option was a shared mounted drive, I'd just run a server on one VM and have all of those writes/reads go through a single machine to avoid any FS shenanigans causing data corruption.
h
Yeah, the main reason we use S3 is that we use transient runners
f
Our setup is on ec2 w/ ebs & s3 for extended storage. Used to be fully s3 but the read latency was hurting performance, so now most reads go through ebs first but its size limited, so s3 as a fallback. I gotta work on getting this merged, but the s3 read path has some problems today https://github.com/buchgr/bazel-remote/pull/808 which cause extra latency/cpu utilization. There are some non-default settings you'll want in bazel-remote too, that increase throughput quite a bit
Copy code
--zstd_implementation=cgo
--access_log_level=none
And use
--enable_endpoint_metrics
to collect and monitor its internal metrics And shameless plug for https://github.com/pantsbuild/pants/pull/22243 to increase per even further on the pants side if anyone has some cycles now that call-by-name is done 🙂
e
My devops team pointed out that our runners are equipped with a shared volume at
/data
. (Gotta love the communication in a big org, as I've been back-burnering the cache work for months now). So I set up the "local filesystem remote cache" to have pants use this volume directly, and it seems to be working fairly smoothly., despite the warning that it is still very experimental. If anyone else is using this, are there any pitfalls to be aware of? In particular I don't see any information about storage size, or needing to periodically clear space or anything like that
h
It’ll grow without bound, so you will need to GC
S3 is nice because it can do that for you automatically based on age