[fathom (gale) — measured, and the remaining work is admin-side rather than repo-side]
#418 moves the cheap required gates onto pulseengine-ci-01. It does not touch the part that matters for queue depth, because that part cannot move yet.
The measurement
Last merge to main (aa48860), 76 jobs: 630 job-minutes, 59 minutes wall-clock. Where it goes:
|
job-min |
share |
Zephyr C Coverage |
39.3 |
|
Zephyr C Coverage (gale-enabled) |
23.4 |
|
34 × qemu_cortex_m3 twister jobs |
360 (median 10.4 each) |
|
| Zephyr subtotal |
~423 |
67% |
| everything else |
~207 |
33% |
And the queue behaviour that makes it hurt: while #416 and #417 were open, 78 checks were queued at once, and the thread sanitizer sat QUEUED for over two hours with no runner assigned. The two self-hosted sanitizer jobs in the same workflow were picked up in three minutes (pulseengine-ci-01-6 and -10, concurrently). Instances -5, -6, -7, -9, -10, -11, -12 have all taken jobs recently, so the pool has real width.
So the machine is there and idle enough; the work that would use it is exactly the work that cannot run there.
Why it cannot run there
17 job definitions use container: — every Zephyr job, LLVM-LTO, all four Renode engine benches, the wasm dist builds, renode-test, size-comparison. The self-hosted runners are themselves containers, so those need container-in-container, which is the parked work on pulseengine-ci-01:
- podman graphroot onto the
containers-ralf volume (the root filesystem cannot hold Zephyr's SDK + workspace — this is the same disk pressure that produced the /mnt design in zephyr-tests.yml on hosted runners, where /mnt is GitHub's ~70 GB ephemeral scratch and does not exist on ours).
- subuid / newuidmap for rootless podman.
Both are host configuration on pulseengine-ci-01, not changes to this repo.
The smaller, separate ask
Three jobs are movable the moment three packages exist in the rust-cpu image, because they only fail on apt under a container with no passwordless sudo:
qemu-system-arm — unblocks mpu-enforcement (a REQUIRED context: "a denied write really faults (REQ-OS-MPU-001 kill-criterion)")
binutils-arm-none-eabi — unblocks the cross-arch seam gate (REQUIRED)
curl — unblocks Rust Coverage (REQUIRED)
With those three, 19 of the 22 required contexts run on our own hardware and a merge stops depending on hosted capacity at all. The remaining three are the container ones above.
What this issue is asking for
Nothing in the repo. Two host changes on pulseengine-ci-01 (podman graphroot + subuid/newuidmap) and optionally three packages in the runner image. #418 is the repo-side half and is independent of both — it lands value now and does not depend on this.
Kill-criterion for the container half, so it is checkable rather than declared done: a container: job (start with renode-test, the smallest) completes on a pulseengine-ci-01-* runner, and zephyr-tests.yml's workspace fits without the /mnt assumptions it currently carries.
[fathom (gale) — measured, and the remaining work is admin-side rather than repo-side]
#418 moves the cheap required gates onto
pulseengine-ci-01. It does not touch the part that matters for queue depth, because that part cannot move yet.The measurement
Last merge to
main(aa48860), 76 jobs: 630 job-minutes, 59 minutes wall-clock. Where it goes:Zephyr C CoverageZephyr C Coverage (gale-enabled)qemu_cortex_m3twister jobsAnd the queue behaviour that makes it hurt: while #416 and #417 were open, 78 checks were queued at once, and the
threadsanitizer sat QUEUED for over two hours with no runner assigned. The two self-hosted sanitizer jobs in the same workflow were picked up in three minutes (pulseengine-ci-01-6and-10, concurrently). Instances-5, -6, -7, -9, -10, -11, -12have all taken jobs recently, so the pool has real width.So the machine is there and idle enough; the work that would use it is exactly the work that cannot run there.
Why it cannot run there
17 job definitions use
container:— every Zephyr job, LLVM-LTO, all four Renode engine benches, the wasm dist builds,renode-test,size-comparison. The self-hosted runners are themselves containers, so those need container-in-container, which is the parked work onpulseengine-ci-01:containers-ralfvolume (the root filesystem cannot hold Zephyr's SDK + workspace — this is the same disk pressure that produced the/mntdesign inzephyr-tests.ymlon hosted runners, where/mntis GitHub's ~70 GB ephemeral scratch and does not exist on ours).Both are host configuration on
pulseengine-ci-01, not changes to this repo.The smaller, separate ask
Three jobs are movable the moment three packages exist in the
rust-cpuimage, because they only fail onaptunder a container with no passwordless sudo:qemu-system-arm— unblocksmpu-enforcement(a REQUIRED context: "a denied write really faults (REQ-OS-MPU-001 kill-criterion)")binutils-arm-none-eabi— unblocks the cross-arch seam gate (REQUIRED)curl— unblocksRust Coverage(REQUIRED)With those three, 19 of the 22 required contexts run on our own hardware and a merge stops depending on hosted capacity at all. The remaining three are the container ones above.
What this issue is asking for
Nothing in the repo. Two host changes on
pulseengine-ci-01(podman graphroot + subuid/newuidmap) and optionally three packages in the runner image. #418 is the repo-side half and is independent of both — it lands value now and does not depend on this.Kill-criterion for the container half, so it is checkable rather than declared done: a
container:job (start withrenode-test, the smallest) completes on apulseengine-ci-01-*runner, andzephyr-tests.yml's workspace fits without the/mntassumptions it currently carries.