docs(agents): cover multi-step exercises in the audit agents - #405
Conversation
What the agents reported about their own blind spots during #403/#404: - terminal-fidelity-auditor: rebuild a Git setup by hand (recipe) instead of stopping at NOT-RUNNABLE-HERE; replay the commands exercise steps ask for and check the claims of warn/restart/successMessage; git is not on pwsh's PATH; never launch a file by its association. - curriculum-validator: count validators used through stepAccepts (no false orphans), check steps (validate xor steps, per-step ByEnv symmetry, one solution command per step), unlocks as info, write scripts to a .mts file. - test-runner: a Playwright test.skip(true, reason) is a runtime skip, not a leaked .skip. - CLAUDE.md: the fidelity trigger covers step texts; multi-step exercises must have no dead end out of order. All agents keep their model alias (opus / sonnet), so the Sonnet agents already run on the latest Sonnet. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
There was a problem hiding this comment.
Sorry @thierryvm, you've used your own review budget of 250,000 diff characters for the last 7 days.
You can request another review in 54 minutes by commenting @sourcery-ai review. Upgrade to get a review now.
Guide de l’évaluateurLes mises à jour portant uniquement sur la documentation expliquent aux agents d’audit comment valider, rejouer et évaluer des exercices en plusieurs étapes, en particulier les leçons reposant sur Git, tout en corrigeant la détection des ignorations à l’exécution et en élargissant les déclencheurs d’invocation concernés. Diagramme de séquence pour l’audit de fidélité des exercices reposant sur GitsequenceDiagram
participant Auditor as terminal-fidelity-auditor
participant Setup as lessonSetup.ts
participant Engine as Terminal simulator
participant Shell as Isolated Git shell
participant Steps as Exercise steps
Auditor->>Setup: Read lesson setup
Auditor->>Shell: Rebuild repository and commits
Auditor->>Engine: setup.apply(createInitialState(), env)
Auditor->>Steps: Read instruction and environment variants
loop Each requested command sequence
Auditor->>Engine: processCommand(command)
Auditor->>Shell: Run command in rebuilt repository
Engine-->>Auditor: Simulated result
Shell-->>Auditor: Real-shell result
end
Auditor->>Steps: Check warn, restart, and successMessage claims
Auditor-->>Auditor: Classify fidelity differences
Diagramme de flux pour l’audit des exercices en plusieurs étapesflowchart TD
Change[Exercise or lesson change] --> Trigger{Relevant trigger?}
Trigger -->|steps, setup, solutions, UI| Validator[curriculum-validator]
Trigger -->|step shell text or behavior| Fidelity[terminal-fidelity-auditor]
Trigger -->|general code or test change| Runner[test-runner]
Validator --> Structure[Validate step structure and solutions]
Validator --> Orphans[Include stepAccepts validators]
Fidelity --> Replay[Replay commands in simulator and real shell]
Fidelity --> Claims[Check warn, restart, and successMessage claims]
Runner --> Skip[Distinguish runtime test.skip from leaked .skip]
Structure --> Review[Report findings]
Orphans --> Review
Replay --> Review
Claims --> Review
Skip --> Review
Modifications au niveau des fichiers
Conseils et commandesInteragir avec Sourcery
Personnaliser votre expérienceAccédez à votre tableau de bord pour :
Obtenir de l’aide
Original review guide in EnglishReviewer's GuideDocumentation-only updates teach the audit agents how to validate, replay, and review multi-step exercises, especially Git-backed lessons, while also correcting runtime-skip detection and broadening the relevant invocation triggers. Sequence diagram for Git-backed exercise fidelity auditingsequenceDiagram
participant Auditor as terminal-fidelity-auditor
participant Setup as lessonSetup.ts
participant Engine as Terminal simulator
participant Shell as Isolated Git shell
participant Steps as Exercise steps
Auditor->>Setup: Read lesson setup
Auditor->>Shell: Rebuild repository and commits
Auditor->>Engine: setup.apply(createInitialState(), env)
Auditor->>Steps: Read instruction and environment variants
loop Each requested command sequence
Auditor->>Engine: processCommand(command)
Auditor->>Shell: Run command in rebuilt repository
Engine-->>Auditor: Simulated result
Shell-->>Auditor: Real-shell result
end
Auditor->>Steps: Check warn, restart, and successMessage claims
Auditor-->>Auditor: Classify fidelity differences
Flow diagram for multi-step exercise auditingflowchart TD
Change[Exercise or lesson change] --> Trigger{Relevant trigger?}
Trigger -->|steps, setup, solutions, UI| Validator[curriculum-validator]
Trigger -->|step shell text or behavior| Fidelity[terminal-fidelity-auditor]
Trigger -->|general code or test change| Runner[test-runner]
Validator --> Structure[Validate step structure and solutions]
Validator --> Orphans[Include stepAccepts validators]
Fidelity --> Replay[Replay commands in simulator and real shell]
Fidelity --> Claims[Check warn, restart, and successMessage claims]
Runner --> Skip[Distinguish runtime test.skip from leaked .skip]
Structure --> Review[Report findings]
Orphans --> Review
Replay --> Review
Claims --> Review
Skip --> Review
File-Level Changes
Tips and commandsInteracting with Sourcery
Customizing Your ExperienceAccess your dashboard to:
Getting Help
|
Why
During #403 and #404 each audit agent ended its report with its own blind spots. This PR writes those lessons into the agents, so the next multi-step exercise PR is audited correctly from the start.
What changes (docs only)
NOT-RUNNABLE-HERE; replay the commands exercise steps ask for; check claims inwarn/restart/successMessage(a text without…ByEnvalso shows on Windows);gitmay be missing frompwsh -NoProfile's PATH; never open a file through its association. Trigger extended to step texts.stepAccepts(...)are not orphans; new check for multi-step exercises (validate xor steps, per-step ByEnv symmetry, one solution command per step); danglingunlocksas info; write scripts to a temporary.mtsfile. Triggered by changes to the Exercise type, exerciseSteps.ts, lessonSetup.ts, lessonSolutions.ts, LessonPage.tsx.test.skip(true, reason)is a runtime skip, not a leaked.skip(.Models: every agent keeps its alias (
opus×13,sonnet×8), so the Sonnet agents already run on the latest Sonnet; no frontmatter change (agentFrontmatter.test.ts64/64).Verification
Voie C — docs-only, 0 runtime file. Smoke test on the preview (HTTP 200 on key routes) before merge.
🤖 Generated with Claude Code
Résumé par Sourcery
Amélioration des consignes de l’agent d’audit afin de valider de manière fiable les exercices en plusieurs étapes et le comportement du shell dans différents environnements.
Corrections de bugs :
Améliorations :
Documentation :
Original summary in English
Summary by Sourcery
Improve audit-agent guidance for reliably validating multi-step exercises and shell behavior across environments.
Bug Fixes:
Enhancements:
Documentation: