Skip to content

Latest commit

 

History

11 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Zhanyl — GPU cluster scheduling and fleet health, on Slurm and Kubernetes

I build the systems that allocate scarce, expensive, heterogeneous compute — and the benchmarks that check whether they actually work.

Each repo below has a status: built (code and tests exist), partial (part of it is built and the rest is marked planned), or planned (nothing exists yet). Each number says how it was produced. None of them comes from a GPU or a production cluster. Where a live control plane was involved, the line says which one.

Website · LinkedIn · X


Scheduling evaluation: Slurm and Kubernetes, same trace

repo status what it is
k8s-gpu-scheduler-lab partial Kubernetes GPU schedulers on identical traces. The control plane and kube-scheduler are real (kind), and the GPU nodes are kwok objects with no hardware. Built: K0 (default kube-scheduler) and four degenerate baselines; a Phase 2 harness with repeats and a bootstrap screen, not yet run on a cluster. Planned: Kueue, Volcano, NVIDIA's Volcano bin-packing, KAI with DRA.
slurm-scheduler-lab built · model only A discrete-event model of Slurm's multifactor priority (Fair Tree and classic fairshare), sched/backfill-style conservative backfill and EASY, and QOS/partition preemption. It reads a slurm.conf and replays synthetic, sacct or k8s-lab traces. Its semantics come from the Slurm docs and source; nothing has been replayed against a live slurmctld.

The two labs share a fleet format and a structural fragmentation definition, pinned by a common golden test vector. slurm-scheduler-lab can replay a k8s-lab trace through the Slurm model (the S0 control, model only). No S0-versus-K0 comparison has been published.

GPU fleet health

These tools share one rule: never act on absent evidence.

repo status what it is
gpu-reaper built · not yet run on a cluster Finds Slurm GPU allocations doing no work, and drains or cancels them only if you enable it. It only observes by default. Missing, stale or unreadable telemetry is never read as idleness. A gap between observations longer than max_sample_gap, which with the defaults means any failed collection cycle, restarts the escalation from the first alert. It cancels only after a drain that was executed and confirmed at least one window earlier. It reads nvidia-smi or dcgm-exporter; the DCGM source is untested against a real exporter. Tested against a simulator, fake Slurm and NVIDIA binaries, and synthetic dcgm-exporter pages.
epilog-gpu-validator built · not validated on hardware A Slurm Epilog check of the GPUs a job just used. It drains a node only on findings it classes as fatal or degraded, and only with --enforce. A GPU query that fails or times out keeps the node in service; nvidia-smi's documented hardware-fault exits (8, 10, 14, 15) count as evidence. It keeps no history, so "persistent" is a judgement about the fault class, not an observation over time. Exercised against simulated nvidia-smi output and fake binaries.
ib-slurm-exporter built · synthetic sysfs only Attributes InfiniBand/RoCE port counters to the Slurm job using the HCA, and withholds the job label on any port it can see is shared. In its default mode, kernel RDMA users (NFS/RDMA, Lustre o2ib, IPoIB) are invisible to it, so it cannot see them sharing a port. Handles Slurm 26.05's SLUID-keyed cgroup paths. Tested against synthetic /sys, /proc and cgroup trees; not yet run on IB hardware or a 26.05 node.
slinky-gitops partial Slurm 26.05 on Kubernetes through Slinky v1.2 on kind, with CI. On Slinky v1.2 the auth-key rotation does not currently take. In three kind CI runs (commits 8e742be to c51903a) the rotation exited non-zero instead of reporting success, and the cluster ran a job afterwards; the run logs, which would show the rollback, were not re-read. The September 2026 rewrite hashes the key in every slurmd pod and in slurmctld, exits 3 after a verified rollback, and has a CI step assert that exit. So far the script has run only in offline tests against a fake kubectl, and the new CI step has not run at all. One NodeSet, no GPUs. The "GitOps" is still planned: no Argo CD yet.

Agentic operations, read-only first

repo status what it is
slurm-rca-bench partial An incident-diagnosis benchmark for the Slurm control plane. 10 scenarios are designed and 6 are runnable. 2 have causal chains measured on a live Slurm 25.11.4 control plane in Docker; the 0.2.0 injection and heal fixes have not been re-run on that cluster. Grading uses a closed vocabulary with no LLM judge, against degenerate baselines. The leaderboard is empty.
slurm-mcp built · fixtures only A read-only MCP server for Slurm state. Its allowlist is enforced in code and tested with 70 adversarial argv, and it exposes three tools with progressive disclosure. Not yet run against a real cluster.
cluster-ops-skills built · not scored 11 runbooks (10 for Slurm, 1 for Slinky) packaged as Agent Skills, each with a mandatory "what not to conclude" section, checked by a validator. Whether they help an agent is unmeasured.
cluster-sre-agent partial A five-configuration ablation, specified before any results, of whether an explicit dependency graph helps an LLM diagnose Slurm incidents. Built: the graph (12 edges, 4 measured on the benchmark's cluster) and a read-only command guard. Planned: agent configurations A–E and any MCP server. No accuracy has been measured.

Other: research-platform (partial), a point-in-time (bitemporal) data layer for quantitative research. It serves as-of queries over an append-only DuckDB store and is checked with mypy --strict. It is Phase 1 of 6; feature lineage and leakage detection are planned.


Things I measured that turned out to be wrong

Each of these is a belief I held, tested, and had to discard. Three of them correct what an earlier version of this profile said.

"Slurm priority weights barely matter; users' --time limits are the real lever." My own simulator contradicts both halves. Backfill does matter: on a synthetic 300-job trace (seed 5), enabling it moved CPU utilisation 72.3% → 83.0% and mean wait 1896.7 → 396.1 min, and across seeds 0–9 it cut mean wait 2.2x-5.7x. Not every weight is flat either. Sweeping each one from 0 to 100,000 changed mean wait by at most 1.26x for age, but by 1.13x-1.84x for fairshare, 1.50x-2.81x for QOS and 2.15x-3.87x for job size. And exact limits did not help waiting. Setting every limit equal to its job's runtime, with nothing else changed, raised CPU utilisation on 7/10 seeds but also raised mean wait on 9/10. That matches the studies Tsafrir reviews (JSSPP 2010), Mu'alem & Feitelson (IEEE TPDS 2001) among them. He argues their padding model, proportional to runtime, leaks runtime information to the scheduler. My padding is proportional to runtime too, so this says nothing yet about real users' limits. Reference model, synthetic traces, no live controller: 16 nodes × 8 CPU / 2 GPU, conservative backfill, Fair Tree. schedlab --compare-backfill --jobs 300 --seed 5, and the seed loops in scripts/reproduce_scheduler_claims.py. Run 2026-09-26. → slurm-scheduler-lab

"A GPU fragmentation percentage is one number." On one kind + kwok run (a real Kubernetes control plane, simulated GPU nodes, n=1), the largest-first baseline scored 0.0% fragmentation under a queue-relative definition and 55.6% under a queue-independent one, on the same data. So a fragmentation figure without its definition cannot be reproduced. That includes the "34% at KubeCon EU 2026" figure the lab was started to test, whose source I could not verify: I found it only in a third-party blog post that names no talk, and NVIDIA's own write-up of its Volcano bin-packing reports node counts (18 → 214 nodes with all four GPUs free), not a percentage. The same configuration, run twice, varied more than most of the differences between schedulers, so no absolute number from that run can be quoted yet. results/results.json in k8s-gpu-scheduler-lab: Run 2, 800 jobs / 1201 pods on a 140-GPU heterogeneous fleet, Phase 1 harness. → k8s-gpu-scheduler-lab

"kubectl rollout restart reaches every pod that holds the key." On Slinky, compute pods are owned by a NodeSet CRD. rollout restart's documented resource types are deployments, daemonsets and statefulsets (kubectl docs). So slurmctld adopted a rotated auth key while slurmd kept the old one, and my rotation script reported success on a cluster that could not run a job. Fixing that exposed a harder failure: a newly created slurmd pod still mounted the previous key. In three later CI runs on kind (commits 8e742be to c51903a) the rotation exited non-zero instead of reporting success, and the cluster ran a job afterwards. The September 2026 rewrite hashes the key in every slurmd pod and in slurmctld, exits 3 after a verified rollback, and has CI assert that exit. It has so far run only in offline tests against a fake kubectl; the kind job has not run it yet. The cause is still a hypothesis (the kubelet's cache of an immutable Secret) that has not been reproduced in isolation. An upstream issue is drafted but not filed. → slinky-gitops

"A stalled accounting database halts Slurm scheduling." I built a benchmark scenario around that folk chain, then measured it. I froze the database (docker pause on MariaDB) for about 15 minutes on a Slurm 25.11.4 control plane in Docker Compose with one compute node, and scheduling did not halt. sacct blocked with no error. slurmctld's DBD agent queue reached 6 and flushed on recovery. Both jobs submitted during the stall ran to completion, though job start latency rose to tens of seconds. The slurm.conf docs say slurmctld queues these messages while slurmdbd is unreachable, with a limit of at least 10000, so this run never came near saturation. Making StateSaveLocation unwritable behaved the opposite way: the one submission attempted during the fault was rejected at once with an I/O error, and no backlog built up. The database stall was probed in full in one emulated run, and a second run confirmed the queue growth; the unwritable state was one run. Two jobs were submitted during the database stall and one during the unwritable state, and a StateSaveLocation stall is untested. The scenario also needed a second correction. It still gave full credit to a shared filesystem that the emulated cluster does not have, and it now credits the frozen database. → slurm-rca-bench

"Adding scenarios fixed my benchmark's degenerate floor." It did, and then two correct fixes undid it. At 5 scenarios, answering db.mysql to every task scored 0.290 while reading nothing. Adding scenarios whose causes lie elsewhere cut that to 0.145 over 10. Then S01's full credit moved to the only layer its injection touches, and four scenarios whose injections cannot produce their ground truth were blocked and left the denominator. Over the 6 runnable scenarios the same constant answer now scores 0.333, above the suite's own 0.25 threshold. The test that enforces the threshold is a strict expected failure until new runnable scenarios fix the suite. Publishing a score without that floor tells the reader nothing. slurm-rca baselines, run 2026-09-26: 0.333 over 6 runnable scenarios (0.200 with --include-blocked). 0.290 (its S01–S05) and 0.145 recomputed at commit 469ea76, before the corrections. → slurm-rca-bench


Four bands. Scheduling evaluation: slurm-scheduler-lab (built, reference model) and k8s-gpu-scheduler-lab (partial: K0 and four degenerate baselines built; Kueue, Volcano and KAI planned), sharing a trace, a fleet and fragmentation definition C. GPU fleet health: gpu-reaper, epilog-gpu-validator and ib-slurm-exporter built but not yet run on a cluster or hardware; slinky-gitops partial, with a rotation that fails safe. The Slurm control plane as measured in slurm-rca-bench (Slurm 25.11.4 in Docker): shared filesystem to accounting database is inferred; the database halting slurmdbd is measured; slurmdbd to slurmctld only degrades, measured; slurmctld to scheduling halts, documented; an unwritable StateSaveLocation rejected the one submission tried, measured once. Agentic operations: slurm-rca-bench partial, slurm-mcp and cluster-ops-skills built, cluster-sre-agent partial with its agent loop planned. On control-plane edges, solid means measured and dashed means documented or inferred; the dotted link between the two labs means shared inputs, not a measurement.


Writing

I publish at zhanyl-tech.github.io: deep dives on HPC and inference, plus shorter lab notes on whatever I'm currently measuring.

Background

  • Education: MS CS (Machine Learning) @ Georgia Tech · CQF (Quantitative Finance) · NVIDIA NCP-AIO
  • Core stack: Python, Go, PyTorch, CUDA, Kubernetes, Slurm
  • Platform: NVIDIA BCM · Run:ai · DCGM · MIG · NVLink/NVSwitch · DOCA/BlueField · InfiniBand · Prometheus
  • Focus: scheduling and resource allocation across Slurm and Kubernetes, GPU cluster reliability, inference infrastructure, and agentic operations

Chess and poker outside of work — both cheaper places to practise reasoning under uncertainty than production is.

About

Profile README

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages