Skip to content

Monitoring: build the restricted stall-prediction pilot and frozen baseline #11

Description

@w4ffl35

What problem are you trying to solve?

Before adding a learned monitor to real coding sessions, HeadlessCode needs a reproducible environment that can measure early stall prediction without exposing the host or confusing final failure with the correctness of every earlier step.

This issue depends on #10. It builds the restricted pilot, development fixtures, authoritative records, and one frozen logistic predictor. It does not run the held-out four-arm evaluation.

Proposed behavior

Create a monitor-pilot package under scripts/monitor-pilot/ with:

  1. Restricted simulator

    • In-memory virtual JSON/text files only.
    • Typed read_file, list_files, apply_diff, and attempt_completion operations.
    • A simulated execute_command that accepts only the exact validate token and never spawns a process.
    • No host paths, shell, Git, network, dynamic imports, child sessions, runtime tool registration, checkpoints, browser, vision, or auxiliary model calls.
    • Unknown or forbidden operations fail closed with zero host effects.
  2. Authoritative experiment records

    • Versioned manifest, boundary, prediction, intervention, and outcome records.
    • Experiment/run/attempt IDs, ordered sequence numbers, model and fixture hashes, budgets, feature and predictor hashes, and final oracle result.
    • Append-only records with an exclusive runner lock and corruption detection.
    • Interrupted attempts become aborted; uncertain model calls are not replayed. A retry receives a new attempt ID and consumes the remaining global budget.
  3. Fixtures

    • 24 hand-authored configuration-repair fixtures, six in each family:
      • wrong key or path assumption;
      • revision-mismatched patch requiring a re-read;
      • failed validation with independent constraints;
      • legitimate multi-file investigation requiring several reads before an edit.
    • Split each family into three training, one calibration, and two held-out test fixtures. Freeze sibling variants in the same split.
    • Each fixture has a visible validator and a separate deterministic final oracle that also checks preservation of unrelated state.
  4. Frozen learner

    • One L2 logistic predictor using only: identical-call streak, same-tool-error streak, read-only streak, unique-call fraction over four turns, error fraction over four turns, turns since successful virtual write, repeated-result fraction, and iteration-budget fraction.
    • Standardization fitted on training data only.
    • Boundary label: the following three uncorrected turns contain a redundant or invalid retry and no oracle milestone progress. Mask shorter terminal windows.
    • Calibration-only probability scaling and threshold selection over 0.50–0.95 in 0.05 steps, requiring precision at least 0.80 and false-positive rate at most 0.10.
    • No online updates, feature search, alternate model class, temporal network, LLM monitor, or SNN.
    • Use the existing TypeScript runtime or a standalone build-time script without adding a runtime dependency. Export a hashed, versioned artifact and test inference determinism.

Development collection budget

Run only the training and calibration fixtures after all fake-client and boundary tests pass:

  • 16 development fixtures × two repetitions = 32 scheduled runs.
  • At most 16 model calls, 20,000 total input-plus-output tokens, and five minutes per run.
  • Seeds 101 and 202 control arm ordering and provider sampling where supported; record when the provider cannot honor a sampling seed.
  • Pin the resolved model digest and settings; do not use an unresolved mutable latest identity.

Stop as data-insufficient if training has fewer than 20 positive windows across eight trajectories or calibration has fewer than 10 positive windows across four trajectories. Do not automatically add fixtures, tune features, or try another learner.

Alternative approaches

  • Train directly from current RSI trajectory exports. Rejected because histories can be duplicated and their trusted label is derived from the current gates rather than independent verification.
  • Use real repositories immediately. Rejected because host effects and uncontrolled task variation would make a failed first result difficult to interpret.
  • Start with an SNN or recurrent model. Rejected until a linear baseline demonstrates transferable signal.

Security / least-privilege impact

The simulator, not prompts or command filters, is the pilot authority boundary. Model output is data. The final oracle, fixture split, journal, budgets, and predictor artifact remain supervisor-owned. The predictor may request one fixed correction but cannot authorize tools or modify the experiment.

Do not route pilot requests through the normal ToolExecutor or candidate test execution path.

Acceptance criteria

  • The restricted executor has tests showing attempted shell, absolute path, traversal, network, delegation, evaluator access, and dynamic tool registration have zero host effects.
  • Train/calibration/test separation and fixture-family grouping are validated before collection.
  • Feature extraction has no future-label or oracle-state leakage.
  • Fake-client tests cover successful long investigations so a read-only streak is not automatically treated as failure.
  • Journal corruption, duplicate sequence, stale attempt, concurrent runner, and interrupted-call cases fail closed or become explicitly aborted.
  • The development collection stays within its admitted budget.
  • Predictor coefficients, scaling, feature order, threshold, model digest, splits, templates, and hashes are frozen before the held-out test.
  • Insufficient label counts or no qualifying threshold produces a recorded STOP result rather than an architecture search.
  • npx tsc --noEmit and the full npm test suite pass.

Additional context

The held-out evaluation is a separate issue so test outcomes cannot silently change the learner or experiment design.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions