Skip to content

fix: recover lost runner assignments - #11

Merged
biw merged 10 commits into
mainfrom
fix/lost-assignment-reconcile
Oct 5, 2026
Merged

biw merged 10 commits into
mainfrom
fix/lost-assignment-reconcile

Conversation

@biw

@biw biw commented Oct 5, 2026 •

Copy link
Copy Markdown
Owner

Recover runner assignments when GitHub's workflow_job: in_progress delivery is missed. A JIT runner can execute a different compatible job from the one that caused it to be provisioned; recovery now confirms the executing job through GitHub and applies its cache scope, scheduler ownership, and resource-trace attribution.

  • Use the executing workflow run to avoid unrelated repository backlogs, with resumable, deadline-bounded candidate scanning for older runner images.
  • Reuse installation credentials during each sweep and record HTTP, timeout, and malformed-response failures.
  • Preserve legacy provisioning RPC calls during deployment while protecting replacement attempts from stale workflows.
  • Cover backlog recovery, pagination and cursor wraparound, ownership boundaries, resource attribution, HTTP parsing, authentication, and deployment compatibility.

Based on the latest main, including the merged GitHub-hosted CI and publishing hardening from #12. This PR's remaining diff contains the runner recovery changes; CI uses Node 26 on ubuntu-latest and the shared trusted publisher at @v1.

Validation: on merge commit e706450, both GitHub-hosted Node 26 jobs passed all 339 tests across 31 files, formatting, lint, type checks, and the package build. The Docker image check passed at 1,249,804,440 bytes, below the 1.5 GB limit. Local workflow formatting, actionlint, shell syntax, and diff checks passed. Publishing was correctly skipped for the PR. GitHub reports no merge conflicts. CI run.

Supersedes #10 with its existing changes and the follow-up fixes.

GoalSeekLabo and others added 9 commits October 5, 2026 13:07
…ess webhooks

The job-started hook resolves a runner's cache claim from the
`workflow_job: in_progress` webhook observed by the account scheduler.
GitHub does not re-deliver webhooks the Worker fails to acknowledge, so
a dropped or excessively delayed delivery leaves a legitimately running
job polling `/v1/runner-cache-v2/assignment` until its 30 second budget
expires and the job fails before any repository code runs.

This was observed repeatedly under merge queue bursts: GitHub assigned
queued jobs to still-online JIT runners provisioned for an earlier run
(for example job 111562234268 started on `cf-standard-4-job-111560274992`),
and the corresponding `in_progress` delivery never reached the scheduler,
leaving the claim at HTTP 202 until timeout.

Changes:

- Persist `github_runner_id` on `scheduler_jit_runners` when the runner
  is provisioned so the scheduler can ask GitHub about the runner later.
- On an unresolved cache claim, ask GitHub which job the runner is
  actually executing (`GET /repos/{owner}/{repo}/actions/runners/{id}/jobs`),
  persist that as the observed assignment, and resolve the claim to the
  assigned job's cache scope. Reconciles at most once per runner per 10
  seconds so the hook's 1-second polling cannot amplify API traffic, and
  the result still goes through the same repository/profile/scope checks
  as the webhook path.
- Record `jit-runner-assignment-unmatched`,
  `github-job-started-without-owned-job`,
  `jit-runner-assignment-unreconciled`, and `github-assignment-reconciled`
  events so the previously silent drop paths are observable.

Tests cover a reassigned runner resolving through the GitHub fallback and
the no-in-progress-job case still denying access.
The runner-side jobs listing endpoint does not exist in the GitHub REST
API, so the previous reconciliation never actually ran. Ask GitHub for
each unassigned candidate job instead and accept the one whose detail
reports this runner's name. The confirmed assignment now flows through
workflowJobStarted itself, so a runner serving a different job still
requeues the displaced owner and moves its reservation exactly like the
webhook path.

Also: reconcile no longer depends on the new github_runner_id column
(pre-migration runner rows match by name), failed/terminal candidates are
re-checked after the API call, and token issuance, fetch and JSON parsing
failures all degrade to no assignment rather than throwing.
…nd sweep

- Cursor-paginate candidates by numeric job id so assignments beyond the
  newest eight are still reached across attempts
- Start provisioning workflows for admissions returned by the
  ownership-transfer path so admitted jobs are not stranded
- Adopt a JIT runner executing a known job when no scheduler job owns it
  (mutual cross-assignment), instead of silently dropping the event
- Restrict adoption/transfer to jobs still waiting for a runner so a
  terminal job cannot be resurrected by a late lookup
- Bound the whole sweep (token minting included) by an 8s deadline and
  return before the job-started hook's timeout

Generated with Devin
…n mint

- Accept  candidates: runnerStarted marks jobs running before the
  in_progress webhook, so a lost delivery must still reconcile
- Advance the candidate cursor only past probed rows so a deadline break
  revisits the unprobed tail instead of skipping it forever
- Forward the sweep-deadline signal into installation-token minting via the
  injectable fetch dependency so hung token requests cannot exceed the
  job-started hook budget
- Exclude  jobs from runner adoption: their in-flight
  provisioning workflow would see canStart()=false and tear the job down
- Drop reconcile-attempt state on success and prune entries idle >1h

Generated with Devin
provisioningFailed previously released any active job, so a stale or
retried provisioning workflow could tear down a job that had already been
adopted onto a different JIT runner and was running. Restrict the failure
to jobs still in 'provisioning' and record an event when a stale failure
is ignored.

Generated with Devin
provisioningFailed now requires the runner name from the workflow's
claim (which embeds the -rN attempt suffix) and only fails the job while
it is still provisioning that runner. A superseded workflow can no longer
fail a fresh attempt that re-claimed provisioning, and an adopted or
running job stays untouched.

Generated with Devin
@biw biw changed the title fix(scheduler): recover runner assignments after missed webhooks fix: recover runner assignments and run PR CI on GitHub Oct 5, 2026
@biw biw changed the title fix: recover runner assignments and run PR CI on GitHub fix: recover lost runner assignments Oct 5, 2026
@biw
biw merged commit e82751c into main Oct 5, 2026
5 checks passed
@biw
biw deleted the fix/lost-assignment-reconcile branch October 5, 2026 23:52
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants