|
ra8-firmware 0.1.0
Bare-metal firmware for the Renesas RA8 family (RA8D2 / RA8P1)
|
Every machine that runs CI for this project is declared in one file, infra/fleet.yml. Adding a machine is adding a block to it. Changing how much of a machine CI is allowed to use is changing a number in that block. Nothing else has to be edited: the inventory, the playbook selection, the transport and every role variable are derived from it.
This document is the runbook for the five things you will actually want to do: add a host, retune one, give one quiet hours, remove one, and lend one back to agents as a verification host (section 9). Section 7 explains how the numbers were arrived at, so a new machine can be sized without re-deriving anything.
This is the constraint everything else is built around, and it is not theoretical.
docker stop on a busy runner cancels the job it is running. Not "risks cancelling" – cancels, deliberately, through three hops:
A longer docker stop --time buys nothing: the cancel is immediate and intended, and the extra seconds are spent waiting for a job that has already been told to stop. This fleet has already seen the consequence once – WSL's idle timeout reaped the whole VM out from under three live jobs, and GitHub went on reporting those runners busy=true with orphaned work until it timed out.
So scripts/ci/fleet_capacity.sh never signals a busy runner. It polls, and stops an instance in the moment it is idle. A stopped container cannot be handed another job, so a host converges downward one instance at a time as its jobs finish naturally.
"Is this instance busy" is answered locally: Runner.Listener spawns exactly one Runner.Worker process per job, so docker top <container> answers it with no GitHub API call and no credential on the host. That matters – the quiet-hours timer runs unattended on a machine that deliberately holds no PAT.
If the deadline passes with instances still busy, the drain reports what it could not park and exits non-zero. It does not force them. A scale-down that kills jobs is worse than no scale-down at all.
The residual race – a job assigned in the moment between the busy check passing and the stop landing – is one round trip wide and cannot be closed without a credential on the host, so it is detected: after every stop, the script reads that container's log for the stop's own time window and fails loudly if a job started or was cancelled in it.
Nothing token-free can stop GitHub assigning work to a runner that is up – only removing its labels does that, and that is an API call the machines with quiet hours must not be able to make. So the drain waits for an idle moment and takes it, and if the runner is handed another job first it simply waits again.
The poll interval is 3 seconds because a busy runner's idle gap is only a few seconds wide: measured on the Windows host at a 15-second poll, an instance finished one job and started the next between two checks, costing a whole extra job cycle. Three seconds usually catches the first gap. On a saturated queue, expect a drain to take minutes rather than seconds – the NAS took 8m18s and three job cycles – and to report NOT converged rather than force anything if it is still busy at its 90-minute deadline.
That is the trade the design makes on purpose: a slow scale-down that never loses work, over a fast one that sometimes does.
ARC (the k3s pool) is safe by construction and needs none of this. Its runners are ephemeral – one job, then the pod exits – and the controller only deletes runners that hold no job, so lowering maxRunners never interrupts work. The capacity script's k8s arm just moves the number.
connect.address is an IP or a name a resolver can answer. It is never an ~/.ssh/config alias, and check_fleet_declaration.py rejects a bare label outright.
That rule was bought with real work. Every host used to be addressed as ssh: truenas / ssh: star / ssh: k3s-pve, which resolve only through one machine's private config. The estate then split so that neither half worked:
| ansible | fleet host aliases | |
|---|---|---|
| the Mac | no | yes |
| the dev box | yes | no |
So fleet.py status truenas from the dev box died on Could not resolve hostname truenas while the machine itself answered fine on 10.10.10.1. It was a naming gap, not a routing one – and it cost a NAS that sat at 1 of its declared 2 runners with nothing able to converge it back, plus a runner-image rollout done by hand on all three hosts (#513, #518, #526).
win-ci hid it: it is docker_wsl, so fleet.py ships the play into the distro and runs it --connection=local, never asking the control node to resolve anything. Every docker_linux and k8s host does ask.
Everything is derived from the address now, so no command in this tooling needs a name your machine happens to know:
The gate checks both ends: the declared address must be a literal, and every destination the model derives – ssh argv and inventory line alike – must be one too. A future -J star would pass every input rule and still only work where that alias existed.
That writes ~/.ssh/ra8-fleet.config from the declaration and puts one Include line at the top of ~/.ssh/config. Top, because ssh takes the first value it obtains for each keyword: the declaration wins over a stale hand-written alias of the same name, while anything the fragment does not set (your IdentityFile, a Port) still comes from your own block below it. just infra::setup runs it as part of onboarding.
Do not hand-write the aliases into a control node's ~/.ssh/config. That is the same per-machine prerequisite one level down, and it rots the same way. Print what would be installed with just infra::ssh_config_preview.
The fragment is a convenience, not a dependency: it exists so ssh truenas works for a person and for the scripts and docs that already spell hosts that way. Nothing in fleet.py reads it, which is why a machine with an empty ~/.ssh/config can still drive the whole fleet.
Two things a fresh control node does still need, and neither is a name:
Host keys are not a third prerequisite: every fleet ssh command carries StrictHostKeyChecking=accept-new, which pins a key on first use and still refuses one that later changes. That is strictly stronger than the fleet's Ansible transport, which sets host_key_checking = False.
| class | what it is | capacity is changed by |
|---|---|---|
| arc_k8s | an ARC scale set on a k8s cluster | patching maxRunners |
| docker_linux | long-lived runner containers on any Docker host | draining containers |
| docker_wsl | the same, inside a Windows machine's WSL2 distro | draining containers |
| dev_box | shared verification box + dedicated HIL listener | n/a (not general capacity) |
| hil_bench | the hardware-in-the-loop bench Pi (not a runner) | n/a |
A dev_box may carry one hil_runner: block. It owns the repo-scoped registration identity, custom labels, Actions workflow, and relationship to a declared hil_bench host:
fleet_model.role_vars() resolves the bench's declared address and derives the dev_box_hil_runner_* inputs consumed by the Ansible role. The role's identity/topology defaults are intentionally empty so a direct playbook run cannot recreate a second declaration by accident.
The fleet checker validates the relationship in both directions: every job in the owned workflow must request exactly self-hosted plus the declared custom labels, and no unowned workflow may use those labels. The listener remains outside runners:/budget: because it has no instance count or scale operation. First registration is the sole converge that needs an ephemeral GitHub registration token; follow the mode-0600 procedure in infra/README.md.
Structural facts about a machine – where its runner tree lives, which of its pools CI must never touch, whether it may hold a credential – live in infra/ansible/inventory/host_vars/<host>.yml. Those are properties of the machine, not knobs.
The split is enforced: check_fleet_declaration.py fails a host_vars file that re-declares any variable the declaration owns. Extra-vars beat host_vars, so a duplicate would not change behaviour – it would leave a number in the tree that looks authoritative, that somebody will edit, and that will do nothing.
Worked example: a fourth machine arrives – call it bench-tower, a Ryzen 7 5800X with 16 threads and 32 GB, running Ubuntu, at 192.168.1.40 as the deploy user, that should give CI half of itself.
Half of 16 threads and half of 32 GB is a budget of 8 threads and 16 GB, so min(8/4, 16/8) = 2 instances, at 16/2 = 8 GB and 8/2 = 4 CPUs each. Section 7 explains where the two divisors come from.
ra8-ci is what puts it in the existing pool with no workflow edit – that a plain runner carrying the ARC scale set's name joins the same pool was verified empirically, not assumed (see infra/README.md). The second label is the escape hatch: if GitHub ever stops resolving a plain label that collides with a scale-set name, heavy jobs can be pinned to this host explicitly without touching the others.
Only if it has any. Put host-specific facts in infra/ansible/inventory/host_vars/<host>.yml; for example, the bench tower's file would contain:
The role deliberately does not install these, and fails loudly rather than half-deploying:
The image archive is not a prerequisite and must not be copied by hand. runner_image in infra/fleet.yml declares the Ansible-managed producer, image ref and canonical archive once. On each apply, the Docker role validates the producer archive's tag and image ID, compares SHA-256 with the consumer, transfers and atomically publishes it when stale, then reloads the tag and recreates only containers whose image ID differs. The image is never rebuilt per host: .devcontainer/Dockerfile remains the single source of truth, and a second independently-built image would be a drift source with no upside.
For WSL the play itself runs with connection=local inside the distro, so it cannot delegate back to the producer. fleet.py bridges that transport boundary automatically: it streams the same declared archive into /opt/ra8-infra-cache, verifies SHA-256 before publishing it, and passes that cache to the same Docker role. No archive or host address is maintained twice.
This path is for a host that must not hold a long-lived PAT. The token expires after one hour and exists only in the mode-0600 temporary vars file.
Registration produces long-lived runner credentials on the host, and only those are used to run jobs, so the token is needed once and never again – not even across a reboot.
Never pass credentials as -e KEY=VALUE: they are visible in ps on the control node. The typed recipe accepts only a mode-0600 vars file and forwards it as -e @file. A persistent external vars file can use the same recipe:
just infra::register_runner bench-tower ~/.config/ra8/ci.ymlThe normal established-host path needs neither: the PAT lives in OpenBao and infra/ansible/group_vars/all.yml looks it up at run time.
tower-ci-1 and tower-ci-2 should be online, and the per-host section should show both containers running.
Before the role creates or recreates a Docker listener, it proves that the canonical producer image can execute /usr/local/bin/just, the entry point used by every workflow. The rendered Compose entry point repeats that probe on every subsequent container start and exits before Runner.Listener connects if it fails. The capacity controller independently probes the exact image ID of a stopped container before docker start; this also covers restoration after a failed first rollout, when the old Compose entry point may still be in place. A stale or damaged image therefore stays offline instead of accepting jobs and failing them all with command not found.
A dry run cannot tell you everything. The role deliberately executes its read-only preflight probes in --check mode, including Docker and Compose versions, cgroup mode, storage backing, archive identity, and existing container state. Mutation-dependent postconditions remain apply-only: a dry run cannot prove that a new image was loaded, containers were recreated with that image and their declared caps, or runners returned online. Treat a clean infra-check as a real preflight, then require those apply-time readbacks.
Edit one number:
then:
The apply drains the host first, because a converge recreates containers and that would cancel their jobs. It then re-provisions and brings the new count back up. Expect it to take as long as the longest job currently running there.
An apply also reconciles the runner image. Converge the declared producer (k3s-pve) after changing the toolchain, then apply each Docker/WSL consumer; the current canonical archive is transferred automatically. A direct role run still checks Runner.Worker and refuses to replace a stale busy container, so bypassing the fleet drain cannot silently cancel a job.
If you only want the change for now, do not re-provision at all:
That parks the surplus instances as they go idle and leaves the rest running untouched. Nothing is deregistered and nothing is rebuilt – just infra::scale win-ci 3 brings them straight back. This is the right tool for "I want to play a game for an hour".
| infra-apply | infra-scale | |
|---|---|---|
| changes the declaration's meaning | yes – it is the declaration | no |
| re-registers runners | yes | no |
| downtime for instances that stay | yes (all are drained) | none |
| survives a reboot | yes | yes (docker stop beats restart: unless-stopped) |
| undone by the next infra-apply | n/a | yes |
Container caps are compose-file keys, so they need the container to be recreated: just infra::apply <host>. Nothing takes effect live.
pin_cpus: true gives each instance its own cpuset rather than only a CFS quota, and that is load-bearing rather than tuning: cpus: is a quota and nproc does not see it, so every instance would report the host's full thread count, every build would fan out -j<all of them>, and N instances would oversubscribe the machine N-fold. A cpuset changes the affinity mask, so nproc returns the instance's real share; Just exports that bound through CMAKE_BUILD_PARALLEL_LEVEL, so nested CMake builds inherit it. Instance i gets threads [(i-1)*cpus, i*cpus-1], so the host needs at least instances * cpus threads.
These are written into the Windows user's .wslconfig, and they are the outer limit: the per-instance container caps sit under them.
This one needs a restart. .wslconfig is read when the WSL2 VM boots, so a change needs wsl --shutdown – which stops every distro on that machine, not just the runner's. The tooling reports that and refuses to do it behind the owner's back. Run it yourself when the machine is free; the runners come back with the distro.
Note also that .wslconfig is per Windows user. Written under the wrong profile it is silently ignored and every cap in it does nothing, which is why connect.windows_user is declared rather than guessed.
| change | effect |
|---|---|
| just infra::scale x n | live, drains first, no restart |
| instances | needs infra-apply (drains, re-registers) |
| cpus, memory_gb, pin_cpus | needs infra-apply (recreates containers) |
| labels | needs infra-apply (re-registration) |
| quiet_hours | fleet.py apply <host> --tags capacity – no drain, no container touched |
| dev_slice | fleet.py apply <host> --tags dev-slice – no drain, no container touched |
| budget.threads, budget.memory_gb on docker_wsl | needs infra-apply and wsl --shutdown |
| instances on arc_k8s | infra-apply, or live with infra-scale |
The motivating case: the fastest host in the fleet is also the owner's gaming PC. It should stand CI down on request and come back afterwards, without anyone remembering to do it.
Installing or changing a window does not need a full re-provision. The capacity role is tagged, and that path touches no container:
just infra::apply win-ci installs it too, along with everything else. The window is evaluated in the host's own local time, not UTC.
That makes the host's clock load-bearing, and a WSL2 VM's clock is not reliable on its own – it drifts and jumps across host sleep, which is visible in the drain logs as non-monotonic timestamps from a single process. The wsl_ci_host role waits for timesyncd to report NTPSynchronized=yes before it finishes for exactly this reason. If a window ever appears to fire at the wrong time, check the clock before the schedule.
The timer runs every 10 minutes and asks "what should this host be right now?". It does not fire at the window's edges.
Edge-triggered timers are wrong for a machine that might be switched off, asleep or mid-upgrade at the moment one would have fired: the transition is simply missed, and the host sits at the wrong capacity until the next edge – which for a Friday-evening window means all weekend. Polling is level-triggered and idempotent: a host that was off at 18:00 goes quiet at 18:10 instead, and a host already at its target does nothing at all. It also removes the only genuinely fiddly case, a window that crosses midnight, from the timer and puts it in one place that can be reasoned about once.
Entering the window runs the same drain as just infra::scale: instances are stopped as they go idle, never while they hold a job. Leaving it starts them again.
A host that also lends a dev slice (section 9) has that frozen by the same run, because parking three idle containers while an agent's gate suite goes on using the machine would defeat the whole window. The rule is one line: the dev slice is frozen exactly while the host's runner target is zero – so both the timer and a manual just infra::scale win-ci 0 stand down all of it, and both thaw it again the moment the target is non-zero. Note the consequence of the timer converging capacity generally: a manual N=0 is undone within ten minutes, and the dev slice thaws with it. A durable stand-down is a quiet_hours block or an instances: change, exactly as it is for the runners.
The role asserts the timer is active after installing it: a schedule that is written down and not running is exactly the kind of silent nothing this tree keeps finding in its own tooling. What a healthy run looks like:
That run is also the level-triggered design paying for itself: the host was one instance short of its declared capacity outside its window – left parked by an earlier manual scale – and the timer simply put it back. An edge-triggered pair of alarms would have had nothing to say until the next Friday.
Remove quiet_hours: and re-apply, and the host converges to its declared instance count at all times – the timer stays, because the window was only ever one input to it.
Every runner host carries it, window or not, and that is a fix rather than a generalisation. just infra::scale truenas 1 drained ra8-ci-runner-2 during a bench session; restart: unless-stopped deliberately does not undo an explicit stop; truenas declares no window, so under the old shape it had no timer; so nothing ever re-asked what the host should be. The NAS served CI at half its declared capacity for hours, and the only thing in the tree that knew was just infra::status, which prints the drift and exits 0. win-ci, which does declare a window, healed the identical fault every ten minutes without anyone noticing there had been one.
So fleet_capacity.sh window answers "what should this host be right now" on a host with no window too: its declared capacity. The consequence is deliberate – a live just infra::scale is temporary. To stand a host down durably, change instances: in infra/fleet.yml (or give it a quiet_hours block) and re-converge. That is a capacity decision the next person can find; a parked container on a machine nobody is looking at is not.
/etc/systemd/system has to be writable, and the role fails on a host where it is not rather than converging a machine whose capacity nothing re-asserts. The case it was written for is an appliance with a read-only root: TrueNAS SCALE mounts / ro. On the SCALE release the NAS runs, /etc/systemd/system is writable and it does hold the timer – which is why the role measures the directory instead of inferring it from the distribution. Manual just infra::scale truenas <n> works either way: it pipes the capacity script over ssh and needs nothing installed.
One command, and it is a real teardown: it drains the instances, stops the containers, deregisters every runner from GitHub, removes the image and (on TrueNAS) destroys the ZFS dataset, leaving nothing behind. Then delete the host's block from infra/fleet.yml and its host_vars file.
Removal needs a credential, because deregistering is an API call. Either the PAT, or a short-lived removal token for a host that must not hold one:
The recipe passes only the vars-file path to Ansible; the token never appears in the Just, fleet, or ansible-playbook process arguments.
just infra::remove refuses on a host whose roles do not all implement a teardown path (the dev box, the k3s node, the bench). Only their roles own both halves of the lifecycle; removing the others means undoing them by hand, which is the drift the roles exist to prevent. Add a removal path to the role instead.
Instance count is derived, not guessed:
Both constants are measured properties of this tree, and both are named in infra/fleet.yml so a new machine is sized by plugging in two numbers.
build_parallelism = 4 – every heavy workflow pins its own fan-out with CMAKE_BUILD_PARALLEL_LEVEL (see scripts/ci/lib/parallelism.sh). A job cannot use more CPUs than that no matter how many it is given, so CPU beyond it per instance is bought and never spent.
memory_per_instance_gb = 8 – clang-tidy is the memory ceiling in this tree and has been OOM-killed on an 8 GB machine. An instance that OOMs mid-job is worse than one that never existed: it presents as a flaky gate, and the diagnosis cost is out of all proportion to the capacity gained.
| host | budget | formula | declared |
|---|---|---|---|
| truenas | 8 threads, 16 GB | min(2, 2) = 2 | 2 |
| win-ci | 22 threads, 26 GB | min(5, 3) = 3 | 3 |
| k3s-pve pod | 4 CPU, 12 GB limit | min(1, 1) = 1 per pod | 1 x 6 pods |
On the gaming PC, memory is the binding constraint and cores are not: 22 threads would divide fine at four instances, and 26/4 = 6.5 GB would land back under the OOM threshold. Do not go to four until clang-tidy is sharded and per-shard peak memory has been re-measured.
truenas ran a single instance for a long time on the same reserved budget, which left half of a deliberately-sized reservation idle whenever one job was all it could hold. The formula raised it, with the host's total CI footprint unchanged.
Container memory caps are ceilings, not reservations, so the 8 GB divisor protects the case where several heavy jobs coincide on one host – it is not a claim about one job's typical draw. Measured on win-ci: a full cross-build of all apps peaked at 4.06 GiB RSS (218 apps at the time of measurement; the tree grows, the shape does not), and an idle runner listener holds about 95 MiB.
If per-job peak RSS is ever measured properly across the whole gate set and stays well under the divisor, drop the divisor and every host's recommended count rises with it. That is a measurement away, not a redesign. The number is meant to be revised, not treated as magic.
check_fleet_declaration.py recomputes the formula for every host and fails when a declared count or per-instance cap departs from it with no written sizing_note. Some hosts have real reasons to depart from it – what the gate enforces is that a departure is deliberate and legible. A number nobody can re-derive is folklore.
k3s-pve and the dev box are guests on the same 10-core / 16-thread i5-12600K, together allocated 28 vCPU on a 16-thread part – about 1.75x oversubscribed before anything runs, and measured at load average 68-110 under real use. Adding ARC pods there adds contention, not throughput. The same build-cross gate, same commit, same toolchain image, same 8-CPU allocation:
| gate | win-ci | truenas | a k3s-pve pod |
|---|---|---|---|
| cross-build, all apps (218 when measured) | 258s | 808s | 1689s |
| clang-tidy, full width | 96s | ~330s | ~981s (contended) |
Real capacity comes from machines that are not pve1. That is what truenas and the gaming PC are for, and why the ARC pod ceiling does not rise as the fleet grows.
win-ci is the fleet's fastest machine and it is idle most of the time – measured at load ~2.8 of 22 threads in the troughs between CI bursts. A dev slice hands those troughs to agents: a workspace and a gate run on the same box the runners are on, without CI ever losing a scheduling contest.
Delete the block and re-apply, and the slice, its entry point, its shell environment and its reaper are removed. That is also how the 5 GiB goes back to the Ethos-U55 / NPU work (#228) when that starts: it is one block, in one file.
The two resources behave completely differently under contention, so they are declared differently.
CPU is a weight. Every runner container is a scope under system.slice, which carries systemd's default CPUWeight=100. The dev slice is that slice's sibling at 10. On a thread a runner wants, dev work gets 1/11th of it; on a thread no runner wants, it gets all of it; and it yields within one scheduling period of a job arriving, with nothing to schedule and nobody to remember anything. A cpuset or a --cpus quota would have been exactly backwards – idle while CI is busy, and capped while CI is idle.
Memory is a hard wall, because memory does not yield: a page a dev build holds is a page a clang-tidy job cannot have. MemoryMax is therefore taken out of the host's budget before the runners are sized, and check_fleet_declaration.py fails a declaration where instances * memory_gb + dev_slice.memory_gb overruns budget.memory_gb. On win-ci the arithmetic is closed: 3 x 7 + 5 = 26, the VM cap exactly.
Swap is allowed here and nowhere else in the fleet. A runner that swaps is a job that has silently become an order of magnitude slower and presents as a timeout; a dev gate run that swaps is just slow, and slow beats an OOM kill.
I/O is deliberately not declared. The WSL2 kernel exposes no io.weight (it builds neither BFQ nor blk-iocost), so an IOWeight= would be configuration that does nothing. Stated here so nobody re-derives it.
The unit is ra8dev.slice – no dash. systemd reads - in a slice name as hierarchy: ra8-dev.slice is created as a child of an auto-generated ra8.slice, and it is ra8.slice, at the default weight of 100, that would then face system.slice. The careful CPUWeight=10 would be arbitrating between the dev slice and its zero siblings while CI and dev split the machine evenly one level up. Verified on the host: systemd-run --slice=ra8-dev.slice lands in /ra8.slice/ra8-dev.slice/.... The role asserts the name is a single token, and the model and the role are cross-checked against each other by the fleet-declaration gate.
ra8-dev is the whole interface. Everything else – the workspace root, the shared ccache, the bounded -j, and the --cgroup-parent that puts the gate container in the slice – is exported from /etc/profile.d/ra8-dev-slice.sh and needs no thought.
ra8-dev, not bare just ci. Two kinds of work have to land in the slice and they get there differently. Host-side work (just, git, a cross build) is moved there by ra8-dev, which starts it in a transient scope in the slice – a login shell is otherwise in user.slice, beside CI rather than under it. The gate container cannot be moved that way at all: the docker daemon creates its cgroup, not your shell, so it inherits nothing. That half is done by RA8_CI_CONTAINER_ARGS=--cgroup-parent=ra8dev.slice, which scripts/ci.sh appends to its run command. A bare just ci from a login shell therefore still caps the container correctly, but its just and git run outside the slice. Use ra8-dev.
The dev box carries the pinned toolchain natively. Its one Actions listener is dedicated to HIL and deliberately uses that native toolchain; it is not part of the scalable general runner pool. win-ci already holds the exact image its own runners boot – .devcontainer/Dockerfile plus the actions-runner layer – so a second, apt-installed toolchain beside it would be a drift source with no upside, and the ci_runner_docker role refuses to build a second image for precisely that reason. The dev slice therefore reuses the runners' image: the role tags it ra8-ci:latest (the name scripts/ci.sh boots) and asserts both names resolve to one image id, so "a gate run in the dev slice proves something about the runners beside it" is a checked fact.
The consequence is that just quality::local::gate x – which runs a gate natively, because that is what a CI runner does – does not work in this distro. Use just quality::devcontainer::gate x, which runs the same gate on the same clean snapshot of committed HEAD inside that image. (It works on macOS too, where just quality::local::gate has never been able to.)
The capacity script freezes the slice whenever the host's runner target is 0 – inside a declared window, or after a manual just infra::scale x 0. systemctl freeze suspends every process in it – including the gate container, whose scope is a child of the slice – and thawing resumes them exactly where they stood. An agent's suite pauses for the evening instead of dying at 18:00, which is the same "never destroy work in flight" rule the runner drain follows.
Because the capacity timer converges every host to its declared capacity every ten minutes, a manual N=0 freeze is temporary: the next poll restores the runners and thaws the slice. Standing the machine down for an evening is a quiet_hours window, not a manual scale.
The cost, stated rather than hidden: a run frozen for hours resumes with its wall-clock budgets already spent, so a time-budgeted gate can fail on the way out. That is why ra8-dev refuses to start new work while the slice is frozen and says how to check. A paused run is the caller's informed choice; a run that silently began five minutes before a window is not.
just infra::status reports the slice beside the runners:
The whole point of a weight is that CI does not notice. That was tested rather than assumed.
The method. A probe container given a runner's exact caps – cpuset 14-20, 7 CPU, 7 GiB, --pids-limit 8192, and the default system.slice parent, i.e. indistinguishable from ra8-ci-runner-3 – runs three real gates on a clean snapshot of committed HEAD. Identical work, identical caps, identical cores; the only difference between the phases is whether a full ra8-dev just ci is running in the slice at the same time.
Nothing was parked to make room. Two earlier attempts drained runner 3 so the probe owned its cpuset, and the fleet's own capacity timer put it straight back within ten minutes – correctly, which is now the documented behaviour (section 4). So the other contention is measured instead: a sampler reads every runner's busy state every 5 seconds for the whole round, and a round in which runner 3 was ever busy is reported CONTAMINATED rather than quietly averaged in. Only clean rounds are compared below.
| gate | dev slice idle | dev slice loaded | median change | observed range |
|---|---|---|---|---|
| unit-tests | 72.6 s | 77.0 s | +6.0% | +3.2% .. +16.8% |
| misra | 107.0 s | 117.8 s | +10.1% | +4.3% .. +16.4% |
| tidy | 202.2 s | 211.3 s | +4.5% | +1.7% .. +8.7% |
4 clean control rounds, 7 clean loaded rounds. The absolute numbers are a cold-build worst case and are NOT comparable with the same gates' durations in the Actions history – a real runner works incrementally in its own _work tree, the probe rebuilds from a fresh snapshot every time. The comparison between the two columns is the measurement.
Flakiness: none. 33 of 33 gate runs across every round exited 0, in both phases. The slice sat at its MemoryMax for much of the loaded phase (98405 memory.max events in one 20-minute run) with oom_kill 0 – the hard cap plus swap turned the peak into a slow dev run rather than a dead one, and CI never saw it.
The residual cost is not the scheduler, and cannot be tuned away. Two alternatives were measured on the same rig:
Both point the same way: the weight has already reduced the runqueue share to near nothing, and what is left is SMT siblings and shared last-level cache – physics of two workloads on one die, which no cgroup knob reaches. Removing it would mean giving the slice cores no runner uses, and this host has exactly one spare thread. So cpu.idle was not adopted (it buys nothing and its setting does not survive a slice restart) and max_jobs stays at 6, where the smaller number would have cost dev half its throughput for no gain to CI.
For scale, the fleet already pays more than this to itself. In the same run, the contaminated control rounds – dev slice idle, a real CI job on runner 3 sharing the probe's cpuset – came in at +0.3%..+10.1%, +4.4%..+18.0% and +2.7%..+13.6% on the same three gates. Co-tenancy with the dev slice costs a CI job about what co-tenancy with another runner already costs it on this machine, and that is a trade the fleet made deliberately when it put three runners on one die.
The harness is in issue #519.
The capacity timers above decide how many already-deployed listeners should be running. They do not rebuild the canonical image or apply a changed Ansible role. The dev control node therefore also owns ra8-fleet-reconcile.timer. Every six hours it runs a read-only Ansible check of the ordinary runner hosts and applies real drift in dependency order:
The controller runs from /var/lib/ra8-fleet-reconcile/source, a root-owned snapshot installed by just infra::apply dev. It does not fetch Git, run a branch tip, or execute inside GitHub Actions. This is deliberate: a generic runner with the SSH authority to provision all other runners would make any workflow edit infrastructure-admin code. Applying the dev-box role is the reviewed promotion step that replaces the controller snapshot; the timer then brings every ordinary runner to those approved bytes.
The next timer run applies a full convergence after a new snapshot. It also applies the producer daily and every consumer at least once every seven days, even when check mode reports no drift. The producer's context-staging tasks are deliberately non-idempotent in check mode, so their changed count is reported as CHECK-NOISE; the daily real apply and its role assertions are the authoritative test. The weekly consumer pass closes Ansible's other documented check-mode blind spots. Between full passes, real consumer drift is repaired at the next six-hour run. A failed mutation drains that host to zero capacity; a failed producer blocks consumer updates so an unverified archive is never distributed. A read-only failure does not take an unchanged, last-known-good host down.
The native dev-hil listener and star are excluded. Their roles can touch the physical bench and remain behind the signed, human-present whole-bench hold. Continuous runner maintenance must not weaken that boundary merely because the listener happens to be registered with GitHub.
Operator commands are:
The automated service never stores a GitHub PAT or registration token. Existing runner homes retain their own registration. If one is genuinely lost, the role fails loud and the normal typed, short-lived infra::register_runner bootstrap remains the only registration path.
win-ci uses its tailnet address for fleet maintenance. That address is live when the workstation is on its temporary even-port update leg, which is also the only topology where its GitHub runners have internet access. When the cable returns to isolated odd port3, the tailnet and GitHub listeners are expected to be offline; reconciliation reports the read-only reachability failure and does not drain or rewrite the last-known-good installation. Human bench access on that segment remains through star after resolving the current transient DHCP lease, as documented in infra/network/README.md.
Before promoting a new host into infra/fleet.yml, an operator must complete just infra::doctor and one reviewed just infra::check <host> from dev. That onboarding establishes network reachability and host-key trust before the locked-down timer is allowed to maintain the host unattended.