Skip to content

67% of CI compute cannot move to our own runners: 17 container: jobs blocked on podman-in-podman on pulseengine-ci-01 #419

Description

@avrabe

[fathom (gale) — measured, and the remaining work is admin-side rather than repo-side]

#418 moves the cheap required gates onto pulseengine-ci-01. It does not touch the part that matters for queue depth, because that part cannot move yet.

The measurement

Last merge to main (aa48860), 76 jobs: 630 job-minutes, 59 minutes wall-clock. Where it goes:

job-min share
Zephyr C Coverage 39.3
Zephyr C Coverage (gale-enabled) 23.4
34 × qemu_cortex_m3 twister jobs 360 (median 10.4 each)
Zephyr subtotal ~423 67%
everything else ~207 33%

And the queue behaviour that makes it hurt: while #416 and #417 were open, 78 checks were queued at once, and the thread sanitizer sat QUEUED for over two hours with no runner assigned. The two self-hosted sanitizer jobs in the same workflow were picked up in three minutes (pulseengine-ci-01-6 and -10, concurrently). Instances -5, -6, -7, -9, -10, -11, -12 have all taken jobs recently, so the pool has real width.

So the machine is there and idle enough; the work that would use it is exactly the work that cannot run there.

Why it cannot run there

17 job definitions use container: — every Zephyr job, LLVM-LTO, all four Renode engine benches, the wasm dist builds, renode-test, size-comparison. The self-hosted runners are themselves containers, so those need container-in-container, which is the parked work on pulseengine-ci-01:

  1. podman graphroot onto the containers-ralf volume (the root filesystem cannot hold Zephyr's SDK + workspace — this is the same disk pressure that produced the /mnt design in zephyr-tests.yml on hosted runners, where /mnt is GitHub's ~70 GB ephemeral scratch and does not exist on ours).
  2. subuid / newuidmap for rootless podman.

Both are host configuration on pulseengine-ci-01, not changes to this repo.

The smaller, separate ask

Three jobs are movable the moment three packages exist in the rust-cpu image, because they only fail on apt under a container with no passwordless sudo:

  • qemu-system-arm — unblocks mpu-enforcement (a REQUIRED context: "a denied write really faults (REQ-OS-MPU-001 kill-criterion)")
  • binutils-arm-none-eabi — unblocks the cross-arch seam gate (REQUIRED)
  • curl — unblocks Rust Coverage (REQUIRED)

With those three, 19 of the 22 required contexts run on our own hardware and a merge stops depending on hosted capacity at all. The remaining three are the container ones above.

What this issue is asking for

Nothing in the repo. Two host changes on pulseengine-ci-01 (podman graphroot + subuid/newuidmap) and optionally three packages in the runner image. #418 is the repo-side half and is independent of both — it lands value now and does not depend on this.

Kill-criterion for the container half, so it is checkable rather than declared done: a container: job (start with renode-test, the smallest) completes on a pulseengine-ci-01-* runner, and zephyr-tests.yml's workspace fits without the /mnt assumptions it currently carries.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions