What problem are you trying to solve?
Before adding a learned monitor to real coding sessions, HeadlessCode needs a reproducible environment that can measure early stall prediction without exposing the host or confusing final failure with the correctness of every earlier step.
This issue depends on #10. It builds the restricted pilot, development fixtures, authoritative records, and one frozen logistic predictor. It does not run the held-out four-arm evaluation.
Proposed behavior
Create a monitor-pilot package under scripts/monitor-pilot/ with:
-
Restricted simulator
- In-memory virtual JSON/text files only.
- Typed
read_file, list_files, apply_diff, and attempt_completion operations.
- A simulated
execute_command that accepts only the exact validate token and never spawns a process.
- No host paths, shell, Git, network, dynamic imports, child sessions, runtime tool registration, checkpoints, browser, vision, or auxiliary model calls.
- Unknown or forbidden operations fail closed with zero host effects.
-
Authoritative experiment records
- Versioned manifest, boundary, prediction, intervention, and outcome records.
- Experiment/run/attempt IDs, ordered sequence numbers, model and fixture hashes, budgets, feature and predictor hashes, and final oracle result.
- Append-only records with an exclusive runner lock and corruption detection.
- Interrupted attempts become aborted; uncertain model calls are not replayed. A retry receives a new attempt ID and consumes the remaining global budget.
-
Fixtures
- 24 hand-authored configuration-repair fixtures, six in each family:
- wrong key or path assumption;
- revision-mismatched patch requiring a re-read;
- failed validation with independent constraints;
- legitimate multi-file investigation requiring several reads before an edit.
- Split each family into three training, one calibration, and two held-out test fixtures. Freeze sibling variants in the same split.
- Each fixture has a visible validator and a separate deterministic final oracle that also checks preservation of unrelated state.
-
Frozen learner
- One L2 logistic predictor using only: identical-call streak, same-tool-error streak, read-only streak, unique-call fraction over four turns, error fraction over four turns, turns since successful virtual write, repeated-result fraction, and iteration-budget fraction.
- Standardization fitted on training data only.
- Boundary label: the following three uncorrected turns contain a redundant or invalid retry and no oracle milestone progress. Mask shorter terminal windows.
- Calibration-only probability scaling and threshold selection over 0.50–0.95 in 0.05 steps, requiring precision at least 0.80 and false-positive rate at most 0.10.
- No online updates, feature search, alternate model class, temporal network, LLM monitor, or SNN.
- Use the existing TypeScript runtime or a standalone build-time script without adding a runtime dependency. Export a hashed, versioned artifact and test inference determinism.
Development collection budget
Run only the training and calibration fixtures after all fake-client and boundary tests pass:
- 16 development fixtures × two repetitions = 32 scheduled runs.
- At most 16 model calls, 20,000 total input-plus-output tokens, and five minutes per run.
- Seeds 101 and 202 control arm ordering and provider sampling where supported; record when the provider cannot honor a sampling seed.
- Pin the resolved model digest and settings; do not use an unresolved mutable
latest identity.
Stop as data-insufficient if training has fewer than 20 positive windows across eight trajectories or calibration has fewer than 10 positive windows across four trajectories. Do not automatically add fixtures, tune features, or try another learner.
Alternative approaches
- Train directly from current RSI trajectory exports. Rejected because histories can be duplicated and their trusted label is derived from the current gates rather than independent verification.
- Use real repositories immediately. Rejected because host effects and uncontrolled task variation would make a failed first result difficult to interpret.
- Start with an SNN or recurrent model. Rejected until a linear baseline demonstrates transferable signal.
Security / least-privilege impact
The simulator, not prompts or command filters, is the pilot authority boundary. Model output is data. The final oracle, fixture split, journal, budgets, and predictor artifact remain supervisor-owned. The predictor may request one fixed correction but cannot authorize tools or modify the experiment.
Do not route pilot requests through the normal ToolExecutor or candidate test execution path.
Acceptance criteria
- The restricted executor has tests showing attempted shell, absolute path, traversal, network, delegation, evaluator access, and dynamic tool registration have zero host effects.
- Train/calibration/test separation and fixture-family grouping are validated before collection.
- Feature extraction has no future-label or oracle-state leakage.
- Fake-client tests cover successful long investigations so a read-only streak is not automatically treated as failure.
- Journal corruption, duplicate sequence, stale attempt, concurrent runner, and interrupted-call cases fail closed or become explicitly aborted.
- The development collection stays within its admitted budget.
- Predictor coefficients, scaling, feature order, threshold, model digest, splits, templates, and hashes are frozen before the held-out test.
- Insufficient label counts or no qualifying threshold produces a recorded STOP result rather than an architecture search.
npx tsc --noEmit and the full npm test suite pass.
Additional context
The held-out evaluation is a separate issue so test outcomes cannot silently change the learner or experiment design.
What problem are you trying to solve?
Before adding a learned monitor to real coding sessions, HeadlessCode needs a reproducible environment that can measure early stall prediction without exposing the host or confusing final failure with the correctness of every earlier step.
This issue depends on #10. It builds the restricted pilot, development fixtures, authoritative records, and one frozen logistic predictor. It does not run the held-out four-arm evaluation.
Proposed behavior
Create a monitor-pilot package under
scripts/monitor-pilot/with:Restricted simulator
read_file,list_files,apply_diff, andattempt_completionoperations.execute_commandthat accepts only the exactvalidatetoken and never spawns a process.Authoritative experiment records
Fixtures
Frozen learner
Development collection budget
Run only the training and calibration fixtures after all fake-client and boundary tests pass:
latestidentity.Stop as data-insufficient if training has fewer than 20 positive windows across eight trajectories or calibration has fewer than 10 positive windows across four trajectories. Do not automatically add fixtures, tune features, or try another learner.
Alternative approaches
Security / least-privilege impact
The simulator, not prompts or command filters, is the pilot authority boundary. Model output is data. The final oracle, fixture split, journal, budgets, and predictor artifact remain supervisor-owned. The predictor may request one fixed correction but cannot authorize tools or modify the experiment.
Do not route pilot requests through the normal
ToolExecutoror candidate test execution path.Acceptance criteria
npx tsc --noEmitand the fullnpm testsuite pass.Additional context
The held-out evaluation is a separate issue so test outcomes cannot silently change the learner or experiment design.