diff --git a/.codex/agents/internal-gateway-critical-master.toml b/.codex/agents/internal-gateway-critical-master.toml index 901ebbc4..15a3bde7 100644 --- a/.codex/agents/internal-gateway-critical-master.toml +++ b/.codex/agents/internal-gateway-critical-master.toml @@ -1,6 +1,6 @@ name = "internal-gateway-critical-master" description = "Generic adaptive critical-analysis specialist for plans, proposals, decisions, designs, workflows, requirements, and assumption sets." -sandbox_mode = "workspace-write" +sandbox_mode = "read-only" developer_instructions = ''' ## Required Skill @@ -25,15 +25,15 @@ When that happens, emit the explicit no-context failure report from the skill. ## Operating Posture -Prefer read-only analysis and recommendations. If the user explicitly asks for -an edit, command, or other action, adapt when the available tools, authority, -and safety conditions permit it. This preference is not an absolute prohibition. +Stay report-only in every invocation: make no edit and run no mutating +command. Put every remedy, including a fix the user asked for, under the +report's next actions for the subject owner. ## Output Emit one readable Markdown report in the user's language, following the -skill's fixed layout for the conclusion line, findings, residuals, open -questions, and next actions. Do not emit internal working notes or +skill's fixed layout for the conclusion line, lens line, findings, residuals, +open questions, and next actions. Do not emit internal working notes or machine-only metadata outside the report. ''' diff --git a/.github/CHANGELOG.md b/.github/CHANGELOG.md index 7a68813e..e70486ac 100644 --- a/.github/CHANGELOG.md +++ b/.github/CHANGELOG.md @@ -8,6 +8,84 @@ Use this format for new updates: - One bullet per meaningful change. - Include file/path scope when useful. +## 2026-10-01 + +- Reworked `.github/copilot-instructions.md` and + `.github/instructions/copilot-code-review.instructions.md` for GitHub.com + Copilot code review: removed comment-format directives that code review + ignores, split the always-on content into a review contract (intent check, + noise suppression for linter-enforced issues, intentional fixtures, and + generated files) and severity anchors plus cross-cutting security checks, + and removed the duplicated rules between both files. +- Replaced bundle-checker directives that code review cannot run in the JSON, + YAML, Markdown, and Makefile instructions with concrete review checks; + widened `internal-json` to `**/*.json` with a JSONC exception. +- Made GitHub Actions checks concrete for script injection and + `pull_request_target`/`workflow_run` trust, Terraform checks for `moved`, + `removed`, and `import` blocks, sensitive outputs, and `default_tags`, and + reframed authoring-style `description:` values as review checks. +- Added `.github/instructions/internal-copilot-skill-authoring.instructions.md` + for skill bundles: protected imported bundles, `SKILL.md` frontmatter and + triggers, reachability, self-containment, and tests. + +## 2026-09-27 + +- Reworked .github/skills/internal-gateway-writing-plans/ and .github/skills/internal-gateway-execute-plans/ around shared plan and chat contracts, plan-adjacent run state, read-only handoff checks, native-only execution, bounded repair and stop rules, and checkpoint diffs; added the plan-contract and chat-templates references plus root protocol and ignored-scratch checkpoint tests; removed both retired executor and writer eval-run records. Migration: legacy .superpowers/sdd runs are not resumed; start a new run beside the retained plan. +- Consolidated the five Google Cloud skills into `.github/skills/internal-gcp/`: the router became the single implicit owner with an always-on GCP baseline, a decision mode, a freshness rule, and explicit handoffs; retired `internal-gcp-governance`, `internal-gcp-operations`, `internal-gcp-organization-structure`, and `internal-gcp-strategic` into `references/governance.md`, `references/operations.md`, and `references/structure.md`, with sourced facts for Org Policy dry-run support, IAM deny and Principal Access Boundary composition, VPC Service Controls dry-run, Privileged Access Manager, Data Access audit defaults, Network Connectivity Center, and Backup and DR. +- Replaced the five GCP eval packs with one pack (14 requirements, 10 cases, 3 fixtures, 20 trigger queries) and added `.github/skills/internal-gcp/tests/test_eval_pack_boundaries.py`; added a minimal eval pack to `.github/skills/internal-cloud-policy/` and GCP boundary lines to `internal-cloud-policy` and `internal-terraform`. +- Consolidated the six Azure skills into two: `.github/skills/internal-azure/` became the single implicit Azure control-plane owner with an always-on baseline, a shared workflow, a freshness rule, a three-field output, and `references/structure.md`, `references/governance.md`, `references/operations.md`, `references/decision-mode.md`, and `references/adjacent-owners.md`; retired `internal-azure-governance`, `internal-azure-operations`, `internal-azure-organization-structure`, and `internal-azure-strategic` and removed the Terraform fixtures from the old routing matrix. Migration: invoke `/internal-azure` for every retired lane. +- `.github/skills/internal-azure-devops/` now triggers directly and implicitly, keeps only Azure DevOps-specific guidance in `references/pipelines.md`, and hands off CLI execution and tool-neutral delivery strategy; added eval packs for both Azure skills (12 and 10 requirements, 10 and 6 cases, 20 and 16 trigger queries) and the defective fixture `internal-azure-devops/fixtures/pipeline-secret-inline.yml`. +- Consolidated the seven AWS skills into two: `.github/skills/internal-aws/` became the single implicit AWS platform owner with nine core rules, two output modes, explicit handoffs, and `references/organization.md`, `references/governance.md`, `references/evidence.md`, `references/current-facts.md`, and `references/decisions.md`; retired `internal-aws-governance`, `internal-aws-mcp-research`, `internal-aws-operations`, `internal-aws-organization-structure`, and `internal-aws-strategic` and removed the router `references/routing-matrix.md`. Added sourced guidance for RCPs (scope, exclusions, `RCPFullAWSAccess`), declarative policies, data perimeters, Control Tower control implementations, and mechanism-specific validation. Migration: invoke `/internal-aws` for every retired lane. +- `.github/skills/internal-aws-lambda/` now triggers directly and implicitly, hands off trust and language-only work, adds VPC and freshness rules, and merges `references/common-mistakes.md` into `references/sharp-edges.md`; added eval packs for both AWS skills (14 and 10 requirements, 14 and 8 cases, 21 and 20 trigger queries), four defective fixtures, and `.github/skills/internal-aws/tests/test_eval_pack_boundaries.py`. A runtime pilot recorded `baseline-previous` and `with-skill` case runs (22 of 22 with the new skills, 20 of 22 with the retired family) and held-out activation trials under `tests/evaluation/runs/`. +- Consolidated the six GitHub skills into three implicit owners without a + router: `.github/skills/internal-github-actions/` (workflows, composite + actions, failed-run root cause; references reduced from 15 to 9, including + new `references/run-debugging.md`), `.github/skills/internal-github-pr/` + (template, readiness from fresh `reviewDecision` and check evidence, squash, + terminal state), and new `.github/skills/internal-github-platform/` (decide, + control, and prove modes with sourced references for rulesets, CODEOWNERS, + environments, tokens and OIDC including immutable subject claims, runners, + audit evidence, and security-feature rollout). Retired `internal-github`, + `internal-github-governance`, `internal-github-operations`, and + `internal-github-strategic`. Migration: invoke the deliverable owner + directly; `internal-review-code` now invokes `/internal-github-actions` as + its observer-only contributor. +- Added eval packs for the three GitHub skills (12, 9, and 15 requirements; + 8, 9, and 12 cases; 22, 21, and 25 trigger queries; 12 defective fixtures) + and `tests/test_github_skill_routing_consistency.py`, which checks + identical-prompt sibling near-misses and retired-name references. +- Fixed the ordering check in + `.github/instructions/internal-codeowners.instructions.md`: GitHub applies + the last matching pattern, so catch-alls come first. + +## 2026-09-26 + +- Extended `.github/skills/internal-bash/references/review-anti-patterns.md` with conditioned rows `SH-C04`, `SH-C05`, `SH-M09` to `SH-M12`, and `SH-m08`; added an optional, declared `Compatibility target: Bash 3.2` to `internal-bash` and `internal-bash-script`; added eval cases with purpose-built defective fixtures and a shared dialect-parity case to both packs; added a download-and-execute check to `.github/instructions/internal-bash.instructions.md`. +- Rebuilt `.github/skills/internal-gateway-writing-plans/` and `.github/skills/internal-gateway-execute-plans/` as thin overlays on `superpowers-writing-plans`, `superpowers-executing-plans`, and `superpowers-subagent-driven-development`; retired the Execution Manifest v3, its parser, structural checker, bundled runtime, YAML status sibling, fixtures, and v3 tests, including `tests/internal_gateway/test_bundle_alignment.py` and `tests/internal_gateway/test_execute_plans_v3_contract.py`. +- Execution now runs in the current checkout without commits, records `START` in the superpowers ledger, and stops when `HEAD` moves, a task file was dirty at start, or a change leaves the plan's `Files:` perimeter. +- `internal-gateway-execute-plans` now records object-only checkpoint trees (`CP0..CPn`) for resume, per-task and fix-round review diffs, perimeter checks, and `CP0`-based final attribution; resume continues on partial state inside the first incomplete task and otherwise stops for reconciliation. +- `internal-gateway-writing-plans` now replaces the imported executor header, offers the imported `Subagent-driven` or `Native` handoff, routes every answer to `/internal-gateway-execute-plans`, and converts legacy manifest plans into new plans on explicit request; both eval packs gained cases and a `with-skill` run record. +- The external-resource sync now prefixes relative sibling-skill paths in every `obra-superpowers` file, including scripts, so `task-start` and `task-done` resolve in this repository. + +## 2026-09-16 + +- Split approved Terraform import execution into the secondary `.github/skills/internal-terraform-import/` bundle while retaining `.github/skills/internal-terraform/` as the stable decision and routing wrapper; the relocated runner now defaults to non-mutating assessment and requires complete execute evidence for live import or apply. + +## 2026-09-11 + +- Hardened `.github/skills/internal-gateway-writing-plans/` and `.github/skills/internal-gateway-execute-plans/` against the post-mortem plan-generation defects: the writer checker and the executor parser now emit separator-aware and prose-aware manifest-fence messages with removal hints and bounded 160-character compact finding messages. +- Added the manifest-section hygiene rule, the canonical complete 16-field manifest skeleton with canonical nested values, and the ordered evidence-gate command checklist carrying the literal tokens `structure=passed(0 blocking)` and `execution=passed(0 blocking)` to both `references/manifest-v3.md` files, enforced by full post-preamble reference equality. +- Pinned the post-mortem mutations (`manifest-separator-after-fence`, `partial-manifest-missing-fields`) in the writer regression corpus and added `tests/internal_gateway/test_bundle_alignment.py` with the skeleton parser pin and the differential parity harness for structural mutations. +- Added the plan-completeness contract to both gateway bundles: mandatory `## Target Census`, `## Execution Authorization`, and `## Completeness Audit` sections plus the optional `## Task Graph`, enforced by new blocking checks in `.github/skills/internal-gateway-writing-plans/scripts/check_plan_structure.py` and covered by its regression corpus. +- Aligned both `references/manifest-v3.md` copies with checker behavior: documented the accepted no-hit tokens, the backticked completeness-audit command, the Task Graph emission convention, and the bundle-local pytest suite as the third evidence gate. +- Moved the writer gateway's delegated-authoring mechanics to `references/delegation.md` to keep `SKILL.md` within the 220-line body budget, and repaired the writer `agents/openai.yaml` projection with a complete legacy-material clause, normalized indentation, and the completeness and delegation sync. +- Hardened the writer gate against contradictory plans: execution-gating prose in `## Global Constraints` now blocks in every mode, scope-limiting prose blocks only when it contradicts a declared `modify` target, and the authorization mode must pair with the writer/executor handoff owner in `.github/skills/internal-gateway-writing-plans/scripts/check_plan_structure.py`. +- Added native executor authorization findings `authorization-missing` and `authorization-invalid` for plans with `modify` targets, each naming the `plan-normalization: authorization-backfill` repair, and widened `handoff.next_owner` acceptance to the writer and executor owners in `.github/skills/internal-gateway-execute-plans/scripts/plan_execution.py`. +- Added the parser-validated `plan-normalization` deviation vocabulary (`authorization-backfill`, `census-backfill`, `audit-backfill`, `orphan-rebind`, `constraint-supersession`) requiring the pre-edit `sha256:` semantic fingerprint and the superseded or backfilled content, documented identically in both `references/manifest-v3.md` copies. +- Encoded the dual-gate OK evidence contract and the execution-ready-only handoff offer in the writer surfaces and the bounded normalization loop with typed deviations in the executor surfaces, keeping both `SKILL.md` bodies within the 220-line budget. +- Swept proven-dead legacy paths after a caller/test/current-plan audit: removed the `explicit-single-plan` bootstrap apparatus, the `Preflight` and `Preflight Gate` heading aliases, and the unused delegation compatibility helper, retaining the live rejection codes. +- Extended `tests/internal_gateway/test_bundle_alignment.py` with the missing-authorization, authoring-only, gating-prose, and scope-conflict corpus entries and raised `MINIMUM_EXECUTOR_BLOCKING_CASES` to the new observed count of 16. + ## 2026-08-12 - Narrowed `.github/skills/internal-tf/` to language-only structure guidance, moved cross-platform provider lockfile evidence and executable import-safe guards to `.github/skills/internal-terraform/references/operational-validation.md`, and preserved the wrapper as the fail-safe primary for mixed and operational work. diff --git a/.github/DEPRECATION.md b/.github/DEPRECATION.md index cd44d54a..b54bae6e 100644 --- a/.github/DEPRECATION.md +++ b/.github/DEPRECATION.md @@ -41,6 +41,35 @@ Immediate removal is allowed only for security or compliance issues. The removal ## Current deprecations +- `.github/skills/internal-aws-governance/`, `.github/skills/internal-aws-mcp-research/`, + `.github/skills/internal-aws-operations/`, + `.github/skills/internal-aws-organization-structure/`, and + `.github/skills/internal-aws-strategic/`: **Removed immediately under the + approved architecture-migration exception** on 2026-09-27 with explicit user + authorization. `.github/skills/internal-aws/` is now the single AWS platform + owner, with `references/organization.md`, `references/governance.md`, + `references/evidence.md`, `references/current-facts.md`, and + `references/decisions.md`. `.github/skills/internal-aws-lambda/` is + unchanged in name and now triggers directly. Roll back by restoring the last + Git revision containing the seven bundles. +- `.github/skills/internal-gcp-governance/`, `.github/skills/internal-gcp-operations/`, + `.github/skills/internal-gcp-organization-structure/`, and + `.github/skills/internal-gcp-strategic/`: **Removed immediately under the + approved architecture-migration exception** on 2026-09-27 with explicit user + authorization. `.github/skills/internal-gcp/` is now the single Google Cloud + owner, with `references/structure.md`, `references/governance.md`, and + `references/operations.md`. Roll back by restoring the last Git revision + containing the five bundles. +- `.github/skills/internal-azure-governance/`, `.github/skills/internal-azure-operations/`, + `.github/skills/internal-azure-organization-structure/`, and + `.github/skills/internal-azure-strategic/`: **Removed immediately under the + approved architecture-migration exception** on 2026-09-27 with explicit user + authorization. `.github/skills/internal-azure/` is now the single Azure + control-plane owner, with `references/structure.md`, + `references/governance.md`, `references/operations.md`, and + `references/decision-mode.md`. `.github/skills/internal-azure-devops/` is + unchanged in name and now triggers directly. Roll back by restoring the last + Git revision containing the six bundles. - `.github/skills/ibm-terraform-test/`: **Removed immediately under the approved architecture-migration exception** on 2026-08-05 with explicit user authorization. Native tests and all non-language Terraform work now use diff --git a/.github/INVENTORY.md b/.github/INVENTORY.md index f783a72c..e054ad9d 100644 --- a/.github/INVENTORY.md +++ b/.github/INVENTORY.md @@ -11,6 +11,7 @@ This file is the exact path inventory for the live GitHub Copilot catalog in thi - `.github/instructions/internal-codeowners.instructions.md` - `.github/instructions/internal-codeql.instructions.md` - `.github/instructions/internal-copilot-agent-authoring.instructions.md` +- `.github/instructions/internal-copilot-skill-authoring.instructions.md` - `.github/instructions/internal-copilot-skill-reference-authoring.instructions.md` - `.github/instructions/internal-dependabot.instructions.md` - `.github/instructions/internal-docker.instructions.md` @@ -34,6 +35,8 @@ This file is the exact path inventory for the live GitHub Copilot catalog in thi - `.github/skills/addyosmani-code-review-and-quality/SKILL.md` - `.github/skills/addyosmani-code-simplification/SKILL.md` +- `.github/skills/addyosmani-idea-refine/SKILL.md` +- `.github/skills/addyosmani-planning-and-task-breakdown/SKILL.md` - `.github/skills/anthropic-docx/SKILL.md` - `.github/skills/anthropic-pdf/SKILL.md` - `.github/skills/anthropic-pptx/SKILL.md` @@ -57,19 +60,11 @@ This file is the exact path inventory for the live GitHub Copilot catalog in thi - `.github/skills/awesome-copilot-secret-scanning/SKILL.md` - `.github/skills/awesome-copilot-security-review/SKILL.md` - `.github/skills/grill-me/SKILL.md` +- `.github/skills/grilling/SKILL.md` - `.github/skills/internal-agent-creator/SKILL.md` -- `.github/skills/internal-aws-governance/SKILL.md` - `.github/skills/internal-aws-lambda/SKILL.md` -- `.github/skills/internal-aws-mcp-research/SKILL.md` -- `.github/skills/internal-aws-operations/SKILL.md` -- `.github/skills/internal-aws-organization-structure/SKILL.md` -- `.github/skills/internal-aws-strategic/SKILL.md` - `.github/skills/internal-aws/SKILL.md` - `.github/skills/internal-azure-devops/SKILL.md` -- `.github/skills/internal-azure-governance/SKILL.md` -- `.github/skills/internal-azure-operations/SKILL.md` -- `.github/skills/internal-azure-organization-structure/SKILL.md` -- `.github/skills/internal-azure-strategic/SKILL.md` - `.github/skills/internal-azure/SKILL.md` - `.github/skills/internal-bash-script/SKILL.md` - `.github/skills/internal-bash/SKILL.md` @@ -87,29 +82,23 @@ This file is the exact path inventory for the live GitHub Copilot catalog in thi - `.github/skills/internal-gateway-idea/SKILL.md` - `.github/skills/internal-gateway-simple-task/SKILL.md` - `.github/skills/internal-gateway-writing-plans/SKILL.md` -- `.github/skills/internal-gcp-governance/SKILL.md` -- `.github/skills/internal-gcp-operations/SKILL.md` -- `.github/skills/internal-gcp-organization-structure/SKILL.md` -- `.github/skills/internal-gcp-strategic/SKILL.md` - `.github/skills/internal-gcp/SKILL.md` - `.github/skills/internal-github-actions/SKILL.md` -- `.github/skills/internal-github-governance/SKILL.md` -- `.github/skills/internal-github-operations/SKILL.md` +- `.github/skills/internal-github-platform/SKILL.md` - `.github/skills/internal-github-pr/SKILL.md` -- `.github/skills/internal-github-strategic/SKILL.md` -- `.github/skills/internal-github/SKILL.md` - `.github/skills/internal-go/SKILL.md` - `.github/skills/internal-java-project/SKILL.md` - `.github/skills/internal-java-spring-boot-development/SKILL.md` - `.github/skills/internal-java/SKILL.md` - `.github/skills/internal-json/SKILL.md` +- `.github/skills/internal-knowledge/SKILL.md` - `.github/skills/internal-kubernetes-deployment/SKILL.md` - `.github/skills/internal-kubernetes/SKILL.md` - `.github/skills/internal-lesson-codification/SKILL.md` - `.github/skills/internal-makefile/SKILL.md` - `.github/skills/internal-markdown/SKILL.md` -- `.github/skills/internal-markitdown/SKILL.md` - `.github/skills/internal-mermaid/SKILL.md` +- `.github/skills/internal-microsoft-markitdown/SKILL.md` - `.github/skills/internal-nodejs-project/SKILL.md` - `.github/skills/internal-nodejs/SKILL.md` - `.github/skills/internal-oop-design-patterns/SKILL.md` @@ -122,6 +111,7 @@ This file is the exact path inventory for the live GitHub Copilot catalog in thi - `.github/skills/internal-skill-creator/SKILL.md` - `.github/skills/internal-subagent-contract/SKILL.md` - `.github/skills/internal-tdd/SKILL.md` +- `.github/skills/internal-terraform-import/SKILL.md` - `.github/skills/internal-terraform/SKILL.md` - `.github/skills/internal-tf/SKILL.md` - `.github/skills/internal-wayfinder-report/SKILL.md` @@ -130,29 +120,37 @@ This file is the exact path inventory for the live GitHub Copilot catalog in thi - `.github/skills/local-agent-sync-install-ai-resources/SKILL.md` - `.github/skills/local-copilot-log-analyzer/SKILL.md` - `.github/skills/local-sync-repos/SKILL.md` +- `.github/skills/mattpocock-ask-matt/SKILL.md` - `.github/skills/mattpocock-code-review/SKILL.md` - `.github/skills/mattpocock-codebase-design/SKILL.md` +- `.github/skills/mattpocock-diagnosing-bugs/SKILL.md` - `.github/skills/mattpocock-domain-modeling/SKILL.md` - `.github/skills/mattpocock-grill-with-docs/SKILL.md` - `.github/skills/mattpocock-handoff/SKILL.md` - `.github/skills/mattpocock-implement/SKILL.md` - `.github/skills/mattpocock-improve-codebase-architecture/SKILL.md` +- `.github/skills/mattpocock-prototype/SKILL.md` - `.github/skills/mattpocock-research/SKILL.md` +- `.github/skills/mattpocock-resolving-merge-conflicts/SKILL.md` +- `.github/skills/mattpocock-retro/SKILL.md` - `.github/skills/mattpocock-setup-matt-pocock-skills/SKILL.md` - `.github/skills/mattpocock-tdd/SKILL.md` - `.github/skills/mattpocock-teach/SKILL.md` +- `.github/skills/mattpocock-to-questionnaire/SKILL.md` - `.github/skills/mattpocock-to-spec/SKILL.md` - `.github/skills/mattpocock-to-tickets/SKILL.md` - `.github/skills/mattpocock-triage/SKILL.md` +- `.github/skills/mattpocock-wait-what/SKILL.md` - `.github/skills/mattpocock-wayfinder/SKILL.md` -- `.github/skills/mattpocock-writing-great-skills/SKILL.md` +- `.github/skills/mattpocock-wizard/SKILL.md` +- `.github/skills/mattpocock-writing-for-agents/SKILL.md` - `.github/skills/openai-docs/SKILL.md` - `.github/skills/openai-gh-address-comments/SKILL.md` - `.github/skills/openai-gh-fix-ci/SKILL.md` - `.github/skills/search-company-knowledge/SKILL.md` - `.github/skills/superpowers-brainstorming/SKILL.md` +- `.github/skills/superpowers-diagnosing-superpowers/SKILL.md` - `.github/skills/superpowers-dispatching-parallel-agents/SKILL.md` -- `.github/skills/superpowers-executing-plans/SKILL.md` - `.github/skills/superpowers-finishing-a-development-branch/SKILL.md` - `.github/skills/superpowers-receiving-code-review/SKILL.md` - `.github/skills/superpowers-requesting-code-review/SKILL.md` @@ -162,37 +160,120 @@ This file is the exact path inventory for the live GitHub Copilot catalog in thi - `.github/skills/superpowers-using-git-worktrees/SKILL.md` - `.github/skills/superpowers-using-superpowers/SKILL.md` - `.github/skills/superpowers-verification-before-completion/SKILL.md` -- `.github/skills/superpowers-writing-plans/SKILL.md` - `.github/skills/vercel-find-skills/SKILL.md` ### Support-only imported document skills These vendor-prefixed imported document skills remain support-only depth for repositories that explicitly need document workflows. +## Imported Skill Provenance + +### addyosmani/agent-skills — ref 84ee50673804 · tag 0.6.9 · 2026-09-04 · 4 skills + +- `.github/skills/addyosmani-code-review-and-quality/SKILL.md` +- `.github/skills/addyosmani-code-simplification/SKILL.md` +- `.github/skills/addyosmani-idea-refine/SKILL.md` +- `.github/skills/addyosmani-planning-and-task-breakdown/SKILL.md` + +### anthropics/skills — ref 41bbe19d1a1a · 2026-09-03 · 4 skills + +- `.github/skills/anthropic-docx/SKILL.md` +- `.github/skills/anthropic-pdf/SKILL.md` +- `.github/skills/anthropic-pptx/SKILL.md` +- `.github/skills/anthropic-xlsx/SKILL.md` + +### antonbabenko/terraform-skill — ref 0a3a4a66e990 · 2026-07-03 · 1 skills + +- `.github/skills/antonbabenko-terraform-skill/SKILL.md` + +### atlassian/atlassian-mcp-server — ref 798f8138c976 · 2026-09-03 · 1 skills + +- `.github/skills/search-company-knowledge/SKILL.md` + +### github/awesome-copilot — ref f38fb6cf039b · 2026-09-07 · 10 skills + +- `.github/skills/awesome-copilot-agentic-eval/SKILL.md` +- `.github/skills/awesome-copilot-azure-devops-cli/SKILL.md` +- `.github/skills/awesome-copilot-azure-pricing/SKILL.md` +- `.github/skills/awesome-copilot-azure-resource-health-diagnose/SKILL.md` +- `.github/skills/awesome-copilot-azure-role-selector/SKILL.md` +- `.github/skills/awesome-copilot-cloud-design-patterns/SKILL.md` +- `.github/skills/awesome-copilot-codeql/SKILL.md` +- `.github/skills/awesome-copilot-dependabot/SKILL.md` +- `.github/skills/awesome-copilot-secret-scanning/SKILL.md` +- `.github/skills/awesome-copilot-security-review/SKILL.md` + +### mattpocock/skills — ref 24fe0ef7737e · tag v1.3.1 · 2026-10-04 · 25 skills + +- `.github/skills/grill-me/SKILL.md` +- `.github/skills/grilling/SKILL.md` +- `.github/skills/mattpocock-ask-matt/SKILL.md` +- `.github/skills/mattpocock-code-review/SKILL.md` +- `.github/skills/mattpocock-codebase-design/SKILL.md` +- `.github/skills/mattpocock-diagnosing-bugs/SKILL.md` +- `.github/skills/mattpocock-domain-modeling/SKILL.md` +- `.github/skills/mattpocock-grill-with-docs/SKILL.md` +- `.github/skills/mattpocock-handoff/SKILL.md` +- `.github/skills/mattpocock-implement/SKILL.md` +- `.github/skills/mattpocock-improve-codebase-architecture/SKILL.md` +- `.github/skills/mattpocock-prototype/SKILL.md` +- `.github/skills/mattpocock-research/SKILL.md` +- `.github/skills/mattpocock-retro/SKILL.md` +- `.github/skills/mattpocock-setup-matt-pocock-skills/SKILL.md` +- `.github/skills/mattpocock-tdd/SKILL.md` +- `.github/skills/mattpocock-teach/SKILL.md` +- `.github/skills/mattpocock-to-questionnaire/SKILL.md` +- `.github/skills/mattpocock-to-spec/SKILL.md` +- `.github/skills/mattpocock-to-tickets/SKILL.md` +- `.github/skills/mattpocock-triage/SKILL.md` +- `.github/skills/mattpocock-wait-what/SKILL.md` +- `.github/skills/mattpocock-wayfinder/SKILL.md` +- `.github/skills/mattpocock-wizard/SKILL.md` +- `.github/skills/mattpocock-writing-for-agents/SKILL.md` + +### obra/superpowers — ref 8ca22dba9a94 · tag v6.4.2 · 2026-09-25 · 12 skills + +- `.github/skills/superpowers-brainstorming/SKILL.md` +- `.github/skills/superpowers-diagnosing-superpowers/SKILL.md` +- `.github/skills/superpowers-dispatching-parallel-agents/SKILL.md` +- `.github/skills/superpowers-finishing-a-development-branch/SKILL.md` +- `.github/skills/superpowers-receiving-code-review/SKILL.md` +- `.github/skills/superpowers-requesting-code-review/SKILL.md` +- `.github/skills/superpowers-subagent-driven-development/SKILL.md` +- `.github/skills/superpowers-systematic-debugging/SKILL.md` +- `.github/skills/superpowers-test-driven-development/SKILL.md` +- `.github/skills/superpowers-using-git-worktrees/SKILL.md` +- `.github/skills/superpowers-using-superpowers/SKILL.md` +- `.github/skills/superpowers-verification-before-completion/SKILL.md` + +### openai/skills — ref 49f948faa925 · 2026-06-23 · 2 skills + +- `.github/skills/openai-gh-address-comments/SKILL.md` +- `.github/skills/openai-gh-fix-ci/SKILL.md` + +### openai/skills — ref 49f948faa925 · 2026-06-23 · 1 skills + +- `.github/skills/openai-docs/SKILL.md` + +### sickn33/antigravity-awesome-skills — ref b1aebac60a88 · tag v16.9.1 · 2026-09-06 · 7 skills + +- `.github/skills/antigravity-api-design-principles/SKILL.md` +- `.github/skills/antigravity-aws-cost-optimizer/SKILL.md` +- `.github/skills/antigravity-cloudformation-best-practices/SKILL.md` +- `.github/skills/antigravity-golang-pro/SKILL.md` +- `.github/skills/antigravity-grafana-dashboards/SKILL.md` +- `.github/skills/antigravity-kubernetes-architect/SKILL.md` +- `.github/skills/antigravity-network-engineer/SKILL.md` + +### vercel-labs/skills — ref 1682051d48c3 · tag v1.5.24 · 2026-09-06 · 1 skills + +- `.github/skills/vercel-find-skills/SKILL.md` + ## Scripts -- `.github/scripts/audit_copilot_catalog.py` -- `.github/scripts/benchmark_skill_tokens.py` -- `.github/scripts/build_inventory.py` -- `.github/scripts/check_catalog_consistency.py` -- `.github/scripts/detect_token_risks.py` -- `.github/scripts/github_catalog_validation.py` +- `.github/scripts/benchmark-skill-tokens.py` - `.github/scripts/graphify-file-change-hook.sh` - `.github/scripts/install-graphify-hooks.sh` -- `.github/scripts/lib/catalog_checks.py` -- `.github/scripts/lib/cli_runner.py` -- `.github/scripts/lib/fingerprinting.py` -- `.github/scripts/lib/internal_skills.py` -- `.github/scripts/lib/inventory.py` -- `.github/scripts/lib/jsonc.py` -- `.github/scripts/lib/repo_paths.py` -- `.github/scripts/lib/shared.py` -- `.github/scripts/lib/skill_change_scope.py` -- `.github/scripts/lib/sync_exclusions.py` -- `.github/scripts/lib/token_risks.py` -- `.github/scripts/run.sh` -- `.github/scripts/validate_internal_skills.py` -- `.github/scripts/validate_skill_change_scope.py` ## Agents diff --git a/.github/README.md b/.github/README.md index 3aeeb11b..1135c99a 100644 --- a/.github/README.md +++ b/.github/README.md @@ -11,3 +11,12 @@ customization assets maintained in `cloud-strategy.github`. - Repository-local sync agents are documented in [`agents/README.md`](agents/README.md). - Reusable gateway and review workflows live under [`skills/`](skills/). + +## Validation + +Run `make github-catalog-validation` for the repository-owned GitHub catalog +validation. Run `make docs-lint` after Markdown changes. + +No diagram is provided because this directory README is a catalog navigation +surface; current component and flow relationships belong in +[docs/architecture.md](../docs/architecture.md). diff --git a/.github/agents/README.md b/.github/agents/README.md index 164ba926..a23d99bc 100644 --- a/.github/agents/README.md +++ b/.github/agents/README.md @@ -11,3 +11,11 @@ This folder contains repository-owned Copilot custom agents. | `local-sync-external-resources` | Declared external-resource refreshes need preparation, audit, planning, or application. | | `local-sync-install-ai-resources` | Repository-owned AI resources or the portable `AGENTS.md` baseline need local-home synchronization. | | `local-sync-repos` | Consumer repositories need managed baseline alignment or drift assessment. | + +## Validation + +Run `make github-catalog-validation` for the repository-owned GitHub catalog +validation. + +No diagram is provided because this README lists agent entrypoints and use +conditions without a material relationship that needs a diagram. diff --git a/.github/agents/internal-gateway-critical-master.agent.md b/.github/agents/internal-gateway-critical-master.agent.md index 355a4d3a..3066b6d2 100644 --- a/.github/agents/internal-gateway-critical-master.agent.md +++ b/.github/agents/internal-gateway-critical-master.agent.md @@ -1,7 +1,7 @@ --- name: internal-gateway-critical-master description: Use this agent when any plan, proposal, decision, design, workflow, requirement, or assumption set needs an adaptive critical challenge. -tools: [read, search, edit, execute] +tools: [read, search] agents: [] --- @@ -31,17 +31,17 @@ evidence exists at all. ## Operating Boundary -Prefer read-only analysis and recommendations. If the user explicitly requests -an edit, command, or other action, adapt when the available tools, authority, -and safety conditions permit it. Do not expose internal working notes or treat -the preferred read-only posture as an absolute prohibition. +Stay report-only in every invocation: make no edit and run no mutating +command. Put every remedy, including a fix the user asked for, under the +report's next actions for the subject owner. Do not expose internal working +notes. ## Output Emit one readable Markdown report in the user's language, following the -skill's fixed layout for the conclusion line, finding blocks, residuals, open -questions, and next actions. Do not emit JSON, machine-only metadata, or -internal notes in chat. +skill's fixed layout for the conclusion line, lens line, finding blocks, +residuals, open questions, and next actions. Do not emit JSON, machine-only +metadata, or internal notes in chat. ## No-context Failure diff --git a/.github/agents/internal-luna-executor.agent.md b/.github/agents/internal-luna-executor.agent.md index 712b67f9..87ddc84b 100644 --- a/.github/agents/internal-luna-executor.agent.md +++ b/.github/agents/internal-luna-executor.agent.md @@ -4,6 +4,8 @@ description: Use this agent when a caller assigns one bounded brief/result task tools: [read, search, web, edit, execute] model: GPT-5.6 Luna agents: [] +user-invocable: false +disable-model-invocation: false --- # Luna Executor diff --git a/.github/copilot-instructions.md b/.github/copilot-instructions.md index c78c0894..30ab095c 100644 --- a/.github/copilot-instructions.md +++ b/.github/copilot-instructions.md @@ -2,61 +2,37 @@ This file is only for GitHub.com Copilot code review. -It is not a general task-execution guide, repository routing guide, planning -workflow, or local agent runtime contract. - -## Review Objective - -Review changed files for defects that matter before merge. - -- Prioritize correctness, security, regressions, missing validation, and - maintainability. -- Prefer actionable findings over broad advice. -- Tie each finding to concrete changed-file evidence. -- Do not restate repository policy unless the diff creates a specific risk. - -## Finding Priority - -Use these buckets when reporting issues: - -- `Critical`: data loss, credential exposure, remote code execution, production - outage, or a merge-blocking contract break. -- `Major`: correctness bugs, security weaknesses, broken validators, missing - required tests, or behavior regressions. -- `Minor`: maintainability, edge-case, resilience, or observability issues that - should be fixed before merge when practical. -- `Nit`: small clarity or style issues that are safe to ignore. -- `Notes`: useful context that is not a defect. - -## Required Checks - -- Check for hardcoded secrets, credentials, keys, tokens, and tenant-sensitive - values. -- Check least privilege, destructive behavior controls, unsafe execution paths, - and missing input validation. -- Check whether changed behavior has appropriate tests, fixtures, docs, or - validators. -- Check contract alignment for changed schemas, generated assets, sync behavior, - prompts, instructions, skills, scripts, and CI workflows. -- Check that fixes are scoped to the requested behavior and do not rewrite large - unaffected areas. - -## Review Discipline - -- Report findings first, ordered by severity. -- For each finding, include the file or changed area, impact, and a concrete fix - direction. -- Avoid speculative findings when the diff does not provide enough evidence. -- Avoid praise, summaries, or style-only comments unless they reveal a real - maintenance risk. -- Escalate repeated instances of the same defect pattern when the repetition - increases risk. +## Comment Only On Evidenced Defects + +- Comment when the diff shows a concrete defect, risk, or missing validation. +- Prioritize correctness, security, regressions, and missing tests over style. +- State the impact and the smallest safe fix in each comment. +- Report a repeated defect once, list the other locations, and raise its + severity when the repetition widens the impact. +- When the diff lacks evidence, ask one short question instead of asserting. + +## Check The Change Against Its Intent + +- Compare the diff with the pull request description and linked issues. Flag + missing requirements, behavior that contradicts the stated intent, and + unrelated changes. +- For changed behavior, flag tests that would still pass if the change were + wrong. +- When a contract changes, such as a schema, CLI, output, frontmatter, or + policy, flag callers, validators, tests, and docs left on the old contract. + +## Skip Noise + +- Skip formatting, import order, and syntax issues that the repository's + configured formatters, linters, and pre-commit hooks already enforce. +- Skip intentional defects in test fixtures, seeded review targets, eval packs, + and gold outputs. Check only that the fixture matches its consuming test. +- Skip generated files unless the generator input and output disagree. ## Non-Scope -- Do not provide implementation plans unless the review finding needs fix - guidance. -- Do not ask the author to follow local runtime workflows that GitHub.com cannot - execute. +- Treat `AGENTS.md` as policy for judging the diff. Its local workflow steps, + such as tool, graph, or planning commands, are not review findings. +- Do not ask authors to run local workflows, agents, or skills. - Do not treat this file as instructions for coding agents, local CLIs, or non-review Copilot chat behavior. diff --git a/.github/dependabot.yml b/.github/dependabot.yml index 7b23352b..6172fa14 100644 --- a/.github/dependabot.yml +++ b/.github/dependabot.yml @@ -7,6 +7,8 @@ updates: day: "monday" time: "08:00" timezone: "Europe/Rome" + cooldown: + default-days: 7 open-pull-requests-limit: 10 labels: - "dependencies" @@ -28,6 +30,8 @@ updates: day: "monday" time: "08:30" timezone: "Europe/Rome" + cooldown: + default-days: 7 open-pull-requests-limit: 5 labels: - "dependencies" diff --git a/.github/instructions/copilot-code-review.instructions.md b/.github/instructions/copilot-code-review.instructions.md index 5b99e5eb..2346643e 100644 --- a/.github/instructions/copilot-code-review.instructions.md +++ b/.github/instructions/copilot-code-review.instructions.md @@ -1,32 +1,30 @@ --- -description: Global defect-first Copilot code review baseline for repository changes. +description: Severity anchors and cross-cutting security and change-hygiene checks for every changed file. applyTo: "**" excludeAgent: "cloud-agent" --- -# Review Objective +# Cross-Cutting Review Checks This file is optimized for Copilot code review and should produce only evidenced findings on matching changed files. -- Find correctness bugs, security issues, regressions, and missing validation/tests before merge. -- Keep findings concise, severity-ordered, and tied to concrete file evidence. -- Prioritize: correctness, security, simplicity, maintainability. +## Severity Anchors -## Required Output Shape +- Critical: secret exposure, remote code execution, data loss, an unguarded destructive action, or a merge-blocking contract break. +- Major: a correctness bug, security weakness, broken validator or CI gate, behavior regression, or changed behavior without tests. +- Minor: an edge-case, resilience, observability, or maintainability risk worth fixing before merge. +- Nit: an optional clarity issue. Raise it only when it can cause misreading. -- Use severity buckets: `Critical`, `Major`, `Minor`, `Nit`, `Notes`. -- For each finding include: file/location, impact, and concrete fix guidance. -- Focus on actionable issues; avoid policy restatement without evidence. +## Security -## Required Checks +- Least privilege and no hardcoded secrets are merge-blocking expectations. +- Flag new permissions, roles, token scopes, or network exposure broader than the change needs. +- Flag credentials, tokens, keys, and tenant or account identifiers in code, configuration, fixtures, logs, or docs. +- Flag untrusted input that reaches shell commands, file paths, templates, queries, or deserialization without validation. +- Flag delete, overwrite, force-push, or state-removal operations without a guard, dry run, or explicit confirmation. -- Least privilege and no hardcoded secrets. -- Input validation, unsafe execution paths, and destructive behavior controls. -- Contract alignment with repository owners, validators, and active tests. -- Missing test coverage for changed behavior and missing docs for behavior changes. +## Change Hygiene -## Review Discipline - -- Do not rewrite large unaffected areas to satisfy style-only preferences. -- Escalate repeated anti-patterns (3+ occurrences in one diff) by one severity level. -- Prefer the smallest safe remediation that preserves requested behavior. +- Flag rewrites of unaffected code, dead code, and leftover debug output. +- Flag dependency additions that are unpinned, unused, or broader than the change needs. +- Flag user-visible behavior changes without matching docs, changelog, or migration notes when the repository keeps them. diff --git a/.github/instructions/internal-azure-devops-pipelines.instructions.md b/.github/instructions/internal-azure-devops-pipelines.instructions.md index abec2ce1..1162ee44 100644 --- a/.github/instructions/internal-azure-devops-pipelines.instructions.md +++ b/.github/instructions/internal-azure-devops-pipelines.instructions.md @@ -1,5 +1,5 @@ --- -description: Best practices for Azure DevOps Pipeline YAML files +description: Azure Pipelines review checks for trigger scope, pinning, secrets, stage gating, and deployment approvals. applyTo: "**/azure-pipelines.yml,**/azure-pipelines*.yml,**/*.pipeline.yml" excludeAgent: "cloud-agent" --- diff --git a/.github/instructions/internal-bash.instructions.md b/.github/instructions/internal-bash.instructions.md index bd5a92ed..31a8c119 100644 --- a/.github/instructions/internal-bash.instructions.md +++ b/.github/instructions/internal-bash.instructions.md @@ -1,5 +1,5 @@ --- -description: Bash scripting standards for safe execution, guard clauses, and consistent runtime logs. +description: Shell script review checks for strict mode, quoting, input handling, destructive commands, and download verification. applyTo: "**/*.sh" excludeAgent: "cloud-agent" --- @@ -16,3 +16,4 @@ This file is optimized for Copilot code review and should produce only evidenced - Flag silent failure paths, swallowed exit codes, or ignored command results. - Check temporary-file and cleanup handling for leak and collision risks. - Report unsafe external input use in command construction. +- Report downloaded content that runs without verification against a trusted digest or signature. diff --git a/.github/instructions/internal-codeowners.instructions.md b/.github/instructions/internal-codeowners.instructions.md index dc3af4e8..75d96b17 100644 --- a/.github/instructions/internal-codeowners.instructions.md +++ b/.github/instructions/internal-codeowners.instructions.md @@ -1,5 +1,5 @@ --- -description: CODEOWNERS standards for template placeholders and review-enforcement readiness. +description: CODEOWNERS review checks for valid owners, rule ordering, placeholders, and coverage of critical paths. applyTo: "**/CODEOWNERS" excludeAgent: "cloud-agent" --- @@ -9,7 +9,9 @@ excludeAgent: "cloud-agent" This file is optimized for Copilot code review and should produce only evidenced findings on matching changed files. - Verify entries use valid owner handles and no unresolved template placeholders remain. -- Check rule ordering so specific paths appear before broad catch-all patterns. +- Check rule ordering: GitHub applies the last matching pattern, so broad + catch-all patterns come first and specific paths after them. Flag a + catch-all that follows and shadows specific rules. - Flag duplicate or conflicting patterns that weaken ownership enforcement. - Confirm critical governance paths have explicit owners. - Verify wildcard usage does not unintentionally widen or drop review coverage. diff --git a/.github/instructions/internal-copilot-agent-authoring.instructions.md b/.github/instructions/internal-copilot-agent-authoring.instructions.md index b788c37b..489f4eac 100644 --- a/.github/instructions/internal-copilot-agent-authoring.instructions.md +++ b/.github/instructions/internal-copilot-agent-authoring.instructions.md @@ -1,5 +1,5 @@ --- -description: Use when authoring or revising repository-owned Copilot agents; owns boundary clarity, paired-asset coherence, and minimal duplication. +description: Review checks for repository-owned Copilot agents covering scope boundaries, routing, referenced assets, and duplication. applyTo: ".github/agents/internal-*.agent.md,.github/agents/local-*.agent.md" excludeAgent: "cloud-agent" --- diff --git a/.github/instructions/internal-copilot-skill-authoring.instructions.md b/.github/instructions/internal-copilot-skill-authoring.instructions.md new file mode 100644 index 00000000..a50f05f3 --- /dev/null +++ b/.github/instructions/internal-copilot-skill-authoring.instructions.md @@ -0,0 +1,32 @@ +--- +description: Review checks for skill bundles covering protected imported bundles, SKILL.md frontmatter and triggers, self-containment, reachability, and tests. +applyTo: ".github/skills/**" +excludeAgent: "cloud-agent" +--- + +# Skill Bundle Review Checks + +This file is optimized for Copilot code review and should produce only evidenced findings on matching changed files. + +## Protected Bundles + +- Treat skill directories whose names do not start with `internal-` or `local-` as imported and read-only. +- Flag any change inside an imported bundle unless the pull request description declares an upstream refresh or names that bundle as an authorized edit. + +## SKILL.md + +- Flag missing or empty `name` or `description` frontmatter, and a `name` that differs from the directory name. +- Flag descriptions that do not state when to use the skill, and triggers that overlap a sibling skill without a routing boundary. +- Flag `/skill-name` invocations of skills that do not exist in the repository. +- Flag files under `references/`, `scripts/`, or `assets/` that no `SKILL.md` link reaches. +- Flag a mandatory rule moved behind a conditional reference. That move changes behavior. + +## Self-Containment + +- Flag references, scripts, fixtures, or assets that resolve outside the skill directory, including host-repository scripts, manifests, or docs. Bundles prefixed `local-` are exempt. + +## Tests And Evals + +- Flag behavior changes in a skill or its scripts without matching bundle tests or eval-pack updates. +- Flag tests that assert raw skill prose instead of parsed structure, script output, or evaluation cases. +- Flag eval criteria or expected outputs weakened to make a failing case pass without a stated rationale. diff --git a/.github/instructions/internal-copilot-skill-reference-authoring.instructions.md b/.github/instructions/internal-copilot-skill-reference-authoring.instructions.md index a2c5a746..a9d5e058 100644 --- a/.github/instructions/internal-copilot-skill-reference-authoring.instructions.md +++ b/.github/instructions/internal-copilot-skill-reference-authoring.instructions.md @@ -1,5 +1,5 @@ --- -description: Use when editing internal/local skill references; owns deep reusable detail without duplicating paired agent or SKILL.md contracts. +description: Review checks for internal and local skill references covering SKILL.md duplication, resolvable links, and current procedures. applyTo: ".github/skills/internal-*/references/**/*.md,.github/skills/local-*/references/**/*.md" excludeAgent: "cloud-agent" --- diff --git a/.github/instructions/internal-docker.instructions.md b/.github/instructions/internal-docker.instructions.md index 0b963d78..09eac1c0 100644 --- a/.github/instructions/internal-docker.instructions.md +++ b/.github/instructions/internal-docker.instructions.md @@ -1,5 +1,5 @@ --- -description: Docker and container build standards for secure, reproducible images and pinned digests. +description: Container review checks for pinned images, runtime user, secret leakage, reproducible builds, and Compose exposure. applyTo: "**/Dockerfile,**/Dockerfile.*,**/*.dockerfile,**/.dockerignore,**/docker-compose*.yml,**/docker-compose*.yaml,**/compose*.yml,**/compose*.yaml" excludeAgent: "cloud-agent" --- diff --git a/.github/instructions/internal-github-actions.instructions.md b/.github/instructions/internal-github-actions.instructions.md index 1f376e07..ff63a588 100644 --- a/.github/instructions/internal-github-actions.instructions.md +++ b/.github/instructions/internal-github-actions.instructions.md @@ -1,5 +1,5 @@ --- -description: Baseline standards for GitHub Actions workflows and composite actions with SHA pinning, least privilege, and deterministic execution. +description: GitHub Actions review checks for pinning, least privilege, script injection, untrusted triggers, and deterministic execution. applyTo: "**/workflows/**,**/actions/**/action.y*ml" excludeAgent: "cloud-agent" --- @@ -10,12 +10,13 @@ This file is optimized for Copilot code review and should produce only evidenced - Flag action and container references that are not pinned to immutable SHAs or digests. - Verify `permissions` are least privilege at workflow and job scope. +- Flag `${{ }}` expressions with attacker-controlled data, such as pull request titles, bodies, branch names, or inputs, placed directly in `run:` scripts. Pass them through `env:`. +- Flag `pull_request_target` or `workflow_run` jobs that check out or run pull request code while holding secrets or write permissions. - Check secret usage to ensure no hardcoded sensitive values in workflows. - Flag long-lived cloud credentials in secrets where OIDC should be used. - Flag production deploy jobs without protected `environment` reviewers. - Flag `workflow_dispatch` inputs consumed by shell or deploy steps without validation. -- Report unsafe event triggers or trust-boundary violations for untrusted code. - Verify concurrency, timeout, and cancellation controls for long-running jobs. - Check cache and artifact keys for deterministic behavior and retention clarity. -- Flag required-check naming drift that can break branch protection enforcement. -- Report context misuse that fails at parse time or queue time. +- Flag renamed jobs or workflows that break required-check names in branch protection or rulesets. +- Trace changed reusable workflows and composite actions to their callers; flag callers left on the old inputs, outputs, or secrets. diff --git a/.github/instructions/internal-go.instructions.md b/.github/instructions/internal-go.instructions.md index 62505914..2b4fa70e 100644 --- a/.github/instructions/internal-go.instructions.md +++ b/.github/instructions/internal-go.instructions.md @@ -1,5 +1,5 @@ --- -description: Instructions for writing Go code following idiomatic Go practices and community standards +description: Go review checks for error handling, context use, concurrency safety, API contracts, and dependency scope. applyTo: "**/*.go,**/go.mod,**/go.sum" excludeAgent: "cloud-agent" --- diff --git a/.github/instructions/internal-java.instructions.md b/.github/instructions/internal-java.instructions.md index f7a15c07..db99b91e 100644 --- a/.github/instructions/internal-java.instructions.md +++ b/.github/instructions/internal-java.instructions.md @@ -1,5 +1,5 @@ --- -description: Java project standards with DDD boundaries, readability-first design, and deterministic unit testing. +description: Java review checks for domain boundaries, null and exception safety, API compatibility, tests, and build reproducibility. applyTo: "**/*.java,**/pom.xml,**/build.gradle,**/build.gradle.kts" excludeAgent: "cloud-agent" --- diff --git a/.github/instructions/internal-json.instructions.md b/.github/instructions/internal-json.instructions.md index d3b45a11..f425992c 100644 --- a/.github/instructions/internal-json.instructions.md +++ b/.github/instructions/internal-json.instructions.md @@ -1,6 +1,6 @@ --- -description: JSON formatting and consistency standards for registry and configuration data files. -applyTo: "**/authorizations/**/*.json,**/organization/**/*.json,**/src/**/*.json,**/data/**/*.json" +description: JSON review checks for strict grammar, duplicate keys, numeric interoperability, and consumer contract changes. +applyTo: "**/*.json" excludeAgent: "cloud-agent" --- @@ -8,14 +8,17 @@ excludeAgent: "cloud-agent" This file is optimized for Copilot code review and should produce only evidenced findings on matching changed files. -- Use the bundle checker for `JSON_BOM`, `JSON_ENCODING`, `JSON_SYNTAX`, - `JSON_DUPLICATE_KEY`, `JSON_NON_FINITE`, `JSON_UNSAFE_INTEGER`, - `JSON_NUMBER_RANGE`, and `JSON_UNPAIRED_SURROGATE` findings. -- Review BOM/UTF-8 handling, duplicate keys, strict grammar, and numeric interoperability - at the format boundary. -- Remember that object order is not semantic in JSON; report ordering only - when a consuming owner explicitly defines presentation requirements. -- Separately review schema-sensitive key or type changes, required properties, - identifiers or enums, and content meaning against an evidenced local contract. -- Report secret exposure and contradictory defaults when the changed file - provides evidence; route broader domain semantics to the owning instruction. +- Flag duplicate object keys. Most parsers keep one value silently. +- Flag a UTF-8 byte order mark, non-UTF-8 content, and unpaired surrogate + escapes. +- Flag comments, trailing commas, `NaN`, and `Infinity` in strict JSON. JSONC + files such as `.vscode/*.json`, `tsconfig*.json`, and `devcontainer.json` + allow comments and trailing commas. +- Flag integers outside the range from -(2^53 - 1) to 2^53 - 1 when a + JavaScript consumer reads them. Suggest a string for large identifiers. +- Do not report key order. JSON object order has no meaning unless a consumer + in the repository defines one. +- Flag changed keys, types, required properties, identifiers, or enum values + that break a schema or consumer visible in the repository. +- Flag secrets and contradictory defaults. +- Leave `.tfvars.json` semantics to the Terraform instruction. diff --git a/.github/instructions/internal-kubernetes-manifests.instructions.md b/.github/instructions/internal-kubernetes-manifests.instructions.md index 1bbb4038..5ababc6d 100644 --- a/.github/instructions/internal-kubernetes-manifests.instructions.md +++ b/.github/instructions/internal-kubernetes-manifests.instructions.md @@ -1,5 +1,5 @@ --- -description: Best practices for Kubernetes YAML manifests including labeling conventions, security contexts, pod security, resource management, probes, and validation commands +description: Kubernetes manifest review checks for labels, security context, resources, probes, exposure, and rollout safety. applyTo: "k8s/**/*.yaml,k8s/**/*.yml,manifests/**/*.yaml,manifests/**/*.yml,deploy/**/*.yaml,deploy/**/*.yml,charts/**/templates/**/*.yaml,charts/**/templates/**/*.yml" excludeAgent: "cloud-agent" --- diff --git a/.github/instructions/internal-lambda.instructions.md b/.github/instructions/internal-lambda.instructions.md index 7afdb8ed..921c2104 100644 --- a/.github/instructions/internal-lambda.instructions.md +++ b/.github/instructions/internal-lambda.instructions.md @@ -1,5 +1,5 @@ --- -description: Lambda implementation rules for explicit handlers, input validation, and reusable business logic. +description: AWS Lambda review checks for handler contracts, input validation, retries and idempotency, limits, IAM scope, and logging. applyTo: "**/*lambda*.tf,**/*lambda*.py,**/*lambda*.js,**/*lambda*.ts,**/lambdas/**/*.tf,**/lambdas/**/*.py,**/lambdas/**/*.js,**/lambdas/**/*.ts,**/functions/**/*.tf,**/functions/**/*.py,**/functions/**/*.js,**/functions/**/*.ts" excludeAgent: "cloud-agent" --- diff --git a/.github/instructions/internal-lessons-learned.instructions.md b/.github/instructions/internal-lessons-learned.instructions.md index 6410a3c1..8fe0fc9f 100644 --- a/.github/instructions/internal-lessons-learned.instructions.md +++ b/.github/instructions/internal-lessons-learned.instructions.md @@ -1,5 +1,5 @@ --- -description: Rules for editing the repository retained-learning ledger without turning it into canonical policy. +description: Review checks that keep the retained-learning ledger concise, owned, and free of canonical policy. applyTo: "LESSONS_LEARNED.md" excludeAgent: "cloud-agent" --- diff --git a/.github/instructions/internal-makefile.instructions.md b/.github/instructions/internal-makefile.instructions.md index bea8d3be..a9492fbf 100644 --- a/.github/instructions/internal-makefile.instructions.md +++ b/.github/instructions/internal-makefile.instructions.md @@ -1,5 +1,5 @@ --- -description: Makefile conventions for deterministic targets, readable recipes, and explicit phony declarations. +description: Makefile review checks for phony targets, recipe syntax, variable expansion, ordering, and failure handling. applyTo: "**/Makefile,**/*.mk" excludeAgent: "cloud-agent" --- @@ -8,16 +8,18 @@ excludeAgent: "cloud-agent" This file is optimized for Copilot code review and should produce only evidenced findings on matching changed files. -- Use the bundle checker and distinguish its `phonydeclared` finding from - review-only Make behavior. -- Check target prerequisites, recipe prefix characters, `.PHONY`, variables, - and `$ / $$` expansion intent. -- Review order-only prerequisites, parallelism, recursive Make, and shared - artifacts when target ordering or concurrency matters. -- Separately review deterministic build order, hidden environment coupling, - failure behavior, and undocumented side effects when the changed file - provides evidence. -- Treat shell semantics and domain behavior as human review concerns; the - checker never invokes recipes. -- Remember that `make -n` is not a generic safety boundary: recipes may still - have observable expansion or tool-specific behavior. +- Flag targets that never create a file of the same name but are missing from + `.PHONY`. +- Flag recipe lines that do not start with a tab or the configured + `.RECIPEPREFIX`. +- Flag `$` and `$$` mistakes: in a recipe, `$VAR` expands the Make variable + `V`, and `$$VAR` passes a shell variable. +- Flag missing prerequisites that break ordering under `make -j`, and output + files written by more than one target. +- Flag recursive calls that use plain `make` instead of `$(MAKE)`. +- Flag ignored failures, such as a `-` prefix or `|| true`, where the failure + matters. +- Flag variables read from the environment without a default or a documented + override. +- Do not assume `make -n` has no side effects. `$(shell ...)` and + `+`-prefixed lines still run. diff --git a/.github/instructions/internal-markdown.instructions.md b/.github/instructions/internal-markdown.instructions.md index 485ad09a..5094b0ea 100644 --- a/.github/instructions/internal-markdown.instructions.md +++ b/.github/instructions/internal-markdown.instructions.md @@ -1,5 +1,5 @@ --- -description: Markdown standards for concise, maintainable documentation and explicit command/path formatting. +description: Markdown review checks for links, fences, references, and technical claims that must match the repository. applyTo: "**/*.md" excludeAgent: "cloud-agent" --- @@ -8,15 +8,18 @@ excludeAgent: "cloud-agent" This file is optimized for Copilot code review and should produce only evidenced findings on matching changed files. -- Use the bundle checker for high-confidence structural findings: MD011 for - reversed links, MD042 for empty links, MD051 for invalid fragments, MD052 - for undefined references, and MD053 for duplicate or unused references. -- Check fences, local links/fragments, paths, reference definitions, and - heading structure without treating a Markdown dialect as universal. -- Record dialect awareness when CommonMark, GitHub Flavored Markdown, or a - tool-specific extension changes the interpretation. -- Separately review technical claims, commands, paths, and examples against - repository evidence; report stale references, contradictory guidance, or - behavior presented as enforced without support from code, tests, or validators. -- Report duplicated policy only when its canonical owner is evident. Leave - external targets, editorial judgment, and broader policy ownership to their owners. +- Flag relative links to files that do not exist in the repository and + fragment links to headings that do not exist in the target file. +- Flag empty links, reversed link syntax such as `(text)[url]`, and undefined + reference-style links. +- Flag unclosed code fences and fence changes that turn prose into code or + code into prose. +- Check commands, paths, file names, and options in changed text against the + repository. Flag stale references and guidance that contradicts code, tests, + or validators. +- Flag text that presents a rule as enforced when no validator, test, or CI + check in the repository enforces it. +- Flag policy copied from its canonical owner file when both copies can drift. + Name the owner. +- Assume GitHub Flavored Markdown unless the file targets another renderer. +- Do not review external link targets or editorial style. diff --git a/.github/instructions/internal-nodejs.instructions.md b/.github/instructions/internal-nodejs.instructions.md index e0b24306..ff5844ad 100644 --- a/.github/instructions/internal-nodejs.instructions.md +++ b/.github/instructions/internal-nodejs.instructions.md @@ -1,5 +1,5 @@ --- -description: Node.js project standards with DDD-oriented layering, early returns, and deterministic test practices. +description: Node.js and TypeScript review checks for layering, async errors, contract changes, tests, and dependency drift. applyTo: "**/*.js,**/*.cjs,**/*.mjs,**/*.ts,**/*.tsx,**/package.json,**/tsconfig.json" excludeAgent: "cloud-agent" --- diff --git a/.github/instructions/internal-python.instructions.md b/.github/instructions/internal-python.instructions.md index 9e4b7fcd..ac3a3249 100644 --- a/.github/instructions/internal-python.instructions.md +++ b/.github/instructions/internal-python.instructions.md @@ -1,5 +1,5 @@ --- -description: Python standards for both scripts and application code with DDD boundaries, guard clauses, and pytest defaults. +description: Python review checks for scripts and importable code covering guard clauses, configuration boundaries, dependencies, output contracts, and pytest defaults. applyTo: "**/*.py" excludeAgent: "cloud-agent" --- @@ -11,7 +11,7 @@ This file is optimized for Copilot code review and should produce only evidenced - Verify guard clauses and error handling make failure modes explicit. - Flag unsafe input handling, shell invocation, or filesystem side effects. - Check function and module boundaries for readability and cohesion. -- Flag behavioral configuration buried in helpers, services, or library modules instead of centralized at the correct boundary: a script entrypoint, `Configuration` section, settings module, adapter, application factory, or composition root. +- Flag behavioral configuration buried in helpers, services, or library modules instead of centralized at the correct boundary: a script entrypoint, settings module, adapter, application factory, or composition root. - Do not flag stable domain invariants merely because they are constants near domain code. - Verify type hints and public interfaces stay consistent with call sites. - Flag manual formatting churn that fights the repository formatter; when Ruff is configured, prefer `ruff format` and Ruff diagnostics over subjective style edits. diff --git a/.github/instructions/internal-terraform.instructions.md b/.github/instructions/internal-terraform.instructions.md index 661faceb..35ab2473 100644 --- a/.github/instructions/internal-terraform.instructions.md +++ b/.github/instructions/internal-terraform.instructions.md @@ -1,5 +1,5 @@ --- -description: Terraform authoring standards for readability, typed interfaces, and validation-first delivery. +description: Terraform review checks for typed interfaces, version constraints, destructive changes, least privilege, and state-address moves. applyTo: "**/*.tf" excludeAgent: "cloud-agent" --- @@ -8,16 +8,14 @@ excludeAgent: "cloud-agent" This file is optimized for Copilot code review and should produce only evidenced findings on matching changed files. -- Flag variables or outputs missing a `description`. -- Flag variables missing an explicit `type`. -- Verify variable and output types are explicit and match actual usage. -- Flag provider, module, or version constraints that are missing or too loose. -- Check resource changes for destructive replacement or drift-risk behavior. +- Flag variables missing `type` or `description`, and outputs missing `description`. +- Flag outputs that expose secrets without `sensitive = true`. +- Flag provider, module, or Terraform version constraints that are missing or too loose. +- Flag resource changes that force replacement of stateful resources, such as a changed name, identifier, or immutable argument. +- Flag renamed resources or modules without a `moved` block, and `removed` or `import` blocks without a migration note in the pull request. - Verify IAM and network changes follow least-privilege intent. - Report hidden dependencies that rely on implicit ordering. -- Check naming, tagging, and state-sensitive references for consistency. - Flag hardcoded IDs, ARNs, subscription IDs, or secrets. -- Flag taggable resources without tags. +- Flag taggable resources without tags unless provider `default_tags` or a tagging module covers them. - Flag non-`snake_case` Terraform identifiers. -- Flag `terraform state mv`, `terraform state rm`, or `terraform import` changes without a documented migration note. -- Flag missing validation or precondition logic on critical inputs. +- Flag missing `validation`, `precondition`, or `postcondition` logic on critical inputs. diff --git a/.github/instructions/internal-yaml.instructions.md b/.github/instructions/internal-yaml.instructions.md index 67ce3f90..d57a5351 100644 --- a/.github/instructions/internal-yaml.instructions.md +++ b/.github/instructions/internal-yaml.instructions.md @@ -1,5 +1,5 @@ --- -description: YAML formatting and clarity conventions for stable, maintainable configuration files. +description: YAML review checks for duplicate keys, indentation, implicit typing, block scalars, anchors, and secret exposure. applyTo: "**/*.yml,**/*.yaml" excludeAgent: "cloud-agent" --- @@ -8,15 +8,18 @@ excludeAgent: "cloud-agent" This file is optimized for Copilot code review and should produce only evidenced findings on matching changed files. -- Run the bundle-owned checker for syntax and the `key-duplicates` rule before - reporting automated findings. -- Check indentation, tabs, scalar styles, block scalar/chomping behavior, and - encoding at the format boundary. -- Review anchors/aliases and merge behavior for explicit, portable intent. -- Treat schema/tag routing as a handoff to the owning platform or domain - instruction; generic YAML validity is not schema validation. -- Separately review secret exposure, runtime-changing values, - environment-scope leaks, and domain-policy changes when the changed file - provides evidence. -- Keep those review-only findings distinct from parser findings and route - schema-specific conclusions to the owning platform or domain instruction. +- Flag duplicate mapping keys. Most parsers keep one value silently. +- Flag tab indentation and indentation changes that move a key to another + parent. +- Flag unquoted scalars that YAML 1.1 parsers retype, such as `yes`, `no`, + `on`, `off`, `1.10`, and leading-zero numbers, where the consumer expects a + string. Keys defined by the consumer, such as GitHub Actions `on:`, are + valid. +- Flag block scalar indicators (`|`, `|-`, `>`) whose trailing-newline change + alters a value the consumer uses. +- Flag anchors, aliases, and merge keys (`<<`) that the consumer parser does + not support or that hide an override. +- Flag secrets, environment-scope leaks, and values that change runtime + behavior without a matching description in the pull request. +- Leave schema checks for workflows, Kubernetes, Compose, and CloudFormation + to their path-specific instructions. diff --git a/.github/repo-profiles.yml b/.github/repo-profiles.yml index 0aac73bd..f00b45f3 100644 --- a/.github/repo-profiles.yml +++ b/.github/repo-profiles.yml @@ -11,7 +11,6 @@ _common_skills: &common_skills - skills/superpowers-brainstorming/SKILL.md - skills/internal-review-code/SKILL.md - skills/superpowers-dispatching-parallel-agents/SKILL.md - - skills/superpowers-executing-plans/SKILL.md - skills/superpowers-finishing-a-development-branch/SKILL.md - skills/superpowers-receiving-code-review/SKILL.md - skills/superpowers-requesting-code-review/SKILL.md @@ -21,7 +20,6 @@ _common_skills: &common_skills - skills/superpowers-using-git-worktrees/SKILL.md - skills/superpowers-using-superpowers/SKILL.md - skills/superpowers-verification-before-completion/SKILL.md - - skills/superpowers-writing-plans/SKILL.md profiles: minimal: @@ -30,16 +28,18 @@ profiles: - skills/internal-markdown/SKILL.md - skills/internal-yaml/SKILL.md - skills/internal-json/SKILL.md - - skills/internal-github/SKILL.md - skills/internal-github-actions/SKILL.md + - skills/internal-github-pr/SKILL.md + - skills/internal-github-platform/SKILL.md backend-java: description: Java service repositories. recommended_skills: - skills/internal-java/SKILL.md - skills/internal-java-project/SKILL.md - - skills/internal-github/SKILL.md - skills/internal-github-actions/SKILL.md + - skills/internal-github-pr/SKILL.md + - skills/internal-github-platform/SKILL.md - skills/internal-markdown/SKILL.md backend-java-spring: @@ -48,8 +48,9 @@ profiles: - skills/internal-java/SKILL.md - skills/internal-java-project/SKILL.md - skills/internal-java-spring-boot-development/SKILL.md - - skills/internal-github/SKILL.md - skills/internal-github-actions/SKILL.md + - skills/internal-github-pr/SKILL.md + - skills/internal-github-platform/SKILL.md - skills/internal-markdown/SKILL.md backend-nodejs: @@ -57,8 +58,9 @@ profiles: recommended_skills: - skills/internal-nodejs/SKILL.md - skills/internal-nodejs-project/SKILL.md - - skills/internal-github/SKILL.md - skills/internal-github-actions/SKILL.md + - skills/internal-github-pr/SKILL.md + - skills/internal-github-platform/SKILL.md - skills/internal-markdown/SKILL.md backend-python: @@ -66,8 +68,9 @@ profiles: recommended_skills: - skills/internal-python/SKILL.md - skills/internal-python-project/SKILL.md - - skills/internal-github/SKILL.md - skills/internal-github-actions/SKILL.md + - skills/internal-github-pr/SKILL.md + - skills/internal-github-platform/SKILL.md - skills/internal-markdown/SKILL.md document-heavy: @@ -84,9 +87,11 @@ profiles: description: Infrastructure repositories with Terraform as primary language. recommended_skills: - skills/internal-terraform/SKILL.md + - skills/internal-terraform-import/SKILL.md - skills/internal-tf/SKILL.md - - skills/internal-github/SKILL.md - skills/internal-github-actions/SKILL.md + - skills/internal-github-pr/SKILL.md + - skills/internal-github-platform/SKILL.md - skills/internal-yaml/SKILL.md - skills/internal-cloud-policy/SKILL.md @@ -94,15 +99,12 @@ profiles: description: Azure platform, landing-zone, or governance repositories that need cost-aware guidance. recommended_skills: - skills/internal-terraform/SKILL.md + - skills/internal-terraform-import/SKILL.md - skills/internal-tf/SKILL.md - skills/internal-yaml/SKILL.md - skills/internal-markdown/SKILL.md - skills/internal-azure/SKILL.md - - skills/internal-azure-organization-structure/SKILL.md - - skills/internal-azure-governance/SKILL.md - - skills/internal-azure-operations/SKILL.md - skills/internal-azure-devops/SKILL.md - - skills/internal-azure-strategic/SKILL.md - skills/awesome-copilot-azure-pricing/SKILL.md mixed-platform: @@ -112,7 +114,9 @@ profiles: - skills/internal-java-project/SKILL.md - skills/internal-nodejs/SKILL.md - skills/internal-nodejs-project/SKILL.md - - skills/internal-github/SKILL.md - skills/internal-github-actions/SKILL.md + - skills/internal-github-pr/SKILL.md + - skills/internal-github-platform/SKILL.md - skills/internal-terraform/SKILL.md + - skills/internal-terraform-import/SKILL.md - skills/internal-tf/SKILL.md diff --git a/.github/scripts/audit_copilot_catalog.py b/.github/scripts/audit_copilot_catalog.py deleted file mode 100644 index b74c477e..00000000 --- a/.github/scripts/audit_copilot_catalog.py +++ /dev/null @@ -1,80 +0,0 @@ -#!/usr/bin/env python3 -"""Purpose: run a deeper audit of the Copilot catalog and governance bridge. - -Usage examples: - python3 ./.github/scripts/audit_copilot_catalog.py --root . - python3 ./.github/scripts/audit_copilot_catalog.py --root . --format json -""" - -from __future__ import annotations - -import argparse -from collections import Counter -from pathlib import Path - -from lib.catalog_checks import run_consistency_checks -from lib.cli_runner import has_severity, run_finding_cli -from lib.shared import Finding, find_repo_root, log_info - - -def parse_args() -> argparse.Namespace: - parser = argparse.ArgumentParser(description="Run a governance-focused audit of the Copilot catalog.") - parser.add_argument("--root", default=".", help="Repository root or any path inside it.") - parser.add_argument("--format", choices=["text", "json", "compact"], default="text", help="Output format.") - return parser.parse_args() - - -def main() -> int: - args = parse_args() - root = find_repo_root(Path(args.root)) - findings = run_finding_cli( - detect_fn=lambda: run_consistency_checks(root, include_token_risks=True), - format_name=args.format, - render_text=render_text, - compact_builder=build_compact_payload, - ) - return 1 if has_severity(findings, "blocking") else 0 - - -def render_text(findings: list[Finding]) -> None: - if not findings: - log_info("No catalog findings detected.") - return - current_severity = None - for finding in findings: - if finding.severity != current_severity: - current_severity = finding.severity - print(f"\n{current_severity.upper()}") - print(f"- {finding.path} :: {finding.code}") - print(f" {finding.message}") - print(f" Suggestion: {finding.suggestion}") - - -def build_compact_payload(findings: list[Finding]) -> dict[str, object]: - severity_counts = Counter(finding.severity for finding in findings) - return { - "status": "failed" if severity_counts.get("blocking", 0) else "ok", - "finding_counts": { - "total": len(findings), - "blocking": severity_counts.get("blocking", 0), - "notice": severity_counts.get("notice", 0), - }, - "finding_sample": [ - { - "severity": finding.severity, - "path": finding.path, - "code": finding.code, - "message": finding.message, - } - for finding in findings[:10] - ], - "next_action": ( - "Address blocking findings before apply workflows." - if severity_counts.get("blocking", 0) - else "No blocking findings; optional notices can be triaged." - ), - } - - -if __name__ == "__main__": - raise SystemExit(main()) diff --git a/.github/scripts/benchmark_skill_tokens.py b/.github/scripts/benchmark-skill-tokens.py similarity index 77% rename from .github/scripts/benchmark_skill_tokens.py rename to .github/scripts/benchmark-skill-tokens.py index 7b309b81..d3b867c0 100644 --- a/.github/scripts/benchmark_skill_tokens.py +++ b/.github/scripts/benchmark-skill-tokens.py @@ -35,6 +35,7 @@ { "scenario": "hcl-only", "primary_owner": "internal-tf", + "execution_owner": None, "delegated_owner": None, "delegated_core_owner": None, "loaded_local_references": ["references/common-mistakes.md"], @@ -46,6 +47,7 @@ { "scenario": "tfvars-json-only", "primary_owner": "internal-tf", + "execution_owner": None, "delegated_owner": None, "delegated_core_owner": None, "loaded_local_references": ["references/structure-standard.md"], @@ -57,6 +59,7 @@ { "scenario": "mixed-adoption", "primary_owner": "internal-terraform", + "execution_owner": None, "delegated_owner": "internal-tf", "delegated_core_owner": ANTON_CORE_SKILL, "loaded_local_references": [ @@ -68,22 +71,29 @@ { "scenario": "native-test", "primary_owner": "internal-terraform", + "execution_owner": None, "delegated_owner": None, "delegated_core_owner": ANTON_CORE_SKILL, "loaded_local_references": ["references/operational-validation.md"], - "forbidden_local_references": ["references/existing-infrastructure-adoption.md"], + "forbidden_local_references": [ + "references/existing-infrastructure-adoption.md" + ], }, { "scenario": "state-or-drift", "primary_owner": "internal-terraform", + "execution_owner": None, "delegated_owner": None, "delegated_core_owner": ANTON_CORE_SKILL, "loaded_local_references": ["references/operational-validation.md"], - "forbidden_local_references": ["references/existing-infrastructure-adoption.md"], + "forbidden_local_references": [ + "references/existing-infrastructure-adoption.md" + ], }, { "scenario": "module-architecture", "primary_owner": "internal-terraform", + "execution_owner": None, "delegated_owner": None, "delegated_core_owner": ANTON_CORE_SKILL, "loaded_local_references": [], @@ -95,19 +105,50 @@ { "scenario": "ci-or-provider-operation", "primary_owner": "internal-terraform", + "execution_owner": None, "delegated_owner": None, "delegated_core_owner": ANTON_CORE_SKILL, "loaded_local_references": ["references/operational-validation.md"], - "forbidden_local_references": ["references/existing-infrastructure-adoption.md"], + "forbidden_local_references": [ + "references/existing-infrastructure-adoption.md" + ], }, { "scenario": "ambiguous-adoption-identity", "primary_owner": "internal-terraform", + "execution_owner": None, "delegated_owner": None, "delegated_core_owner": ANTON_CORE_SKILL, "loaded_local_references": ["references/existing-infrastructure-adoption.md"], "forbidden_local_references": [], }, + { + "scenario": "bulk-multi-state-import", + "primary_owner": "internal-terraform", + "execution_owner": "internal-terraform-import", + "delegated_owner": None, + "delegated_core_owner": ANTON_CORE_SKILL, + "loaded_local_references": [ + "references/existing-infrastructure-adoption.md", + "references/operational-validation.md", + "references/import-orchestration.md", + ], + "forbidden_local_references": [], + }, + { + "scenario": "aws-identity-center-import", + "primary_owner": "internal-terraform", + "execution_owner": "internal-terraform-import", + "delegated_owner": None, + "delegated_core_owner": ANTON_CORE_SKILL, + "loaded_local_references": [ + "references/existing-infrastructure-adoption.md", + "references/operational-validation.md", + "references/import-orchestration.md", + "references/aws-identity-center-import.md", + ], + "forbidden_local_references": [], + }, ) CHAIN_RISK_PATTERNS = [ @@ -175,15 +216,17 @@ def build_scenario_report(root: Path) -> list[dict[str, Any]]: scenario_proxy = skill_tokens + chain_tokens - reports.append({ - "scenario": scenario_name, - "expected_owner": expected_owner, - "skill_tokens": skill_tokens, - "bundle_tokens": bundle_tokens, - "chain_tokens": chain_tokens, - "scenario_proxy": scenario_proxy, - "chain_risks": chain_risks, - }) + reports.append( + { + "scenario": scenario_name, + "expected_owner": expected_owner, + "skill_tokens": skill_tokens, + "bundle_tokens": bundle_tokens, + "chain_tokens": chain_tokens, + "scenario_proxy": scenario_proxy, + "chain_risks": chain_risks, + } + ) return reports @@ -203,8 +246,11 @@ def build_terraform_scenario_report(root: Path) -> list[dict[str, Any]]: reports: list[dict[str, Any]] = [] for scenario in TERRAFORM_SCENARIOS: primary_owner = scenario["primary_owner"] + execution_owner = scenario["execution_owner"] delegated_owner = scenario["delegated_owner"] owners = [primary_owner] + if execution_owner: + owners.append(execution_owner) if delegated_owner: owners.append(delegated_owner) @@ -224,11 +270,10 @@ def build_terraform_scenario_report(root: Path) -> list[dict[str, Any]]: { "scenario": scenario["scenario"], "primary_owner": primary_owner, + "execution_owner": execution_owner, "delegated_owner": delegated_owner, "delegated_core_owner": delegated_core_owner, - "loaded_local_references": list( - scenario["loaded_local_references"] - ), + "loaded_local_references": list(scenario["loaded_local_references"]), "forbidden_local_references": list( scenario["forbidden_local_references"] ), @@ -252,7 +297,9 @@ def build_terraform_scenario_report(root: Path) -> list[dict[str, Any]]: "Direct execute": [GATEWAY_SKILL, "internal-gateway-simple-task"], "Define Gate 0": [GATEWAY_SKILL, "grill-me"], "Define idea and critical": [ - GATEWAY_SKILL, "grill-me", "internal-gateway-critical-master", + GATEWAY_SKILL, + "grill-me", + "internal-gateway-critical-master", ], "Plan handoff": [GATEWAY_SKILL, "internal-gateway-writing-plans"], "Approved apply-plan": [GATEWAY_SKILL, "internal-gateway-execute-plans"], @@ -267,7 +314,8 @@ def build_terraform_scenario_report(root: Path) -> list[dict[str, Any]]: "Idea core entry": [IDEA_GATEWAY_SKILL], "Interview support": [IDEA_GATEWAY_SKILL, "grill-me"], "Mandatory critical pass": [ - IDEA_GATEWAY_SKILL, "internal-gateway-critical-master", + IDEA_GATEWAY_SKILL, + "internal-gateway-critical-master", ], "Visible handoff": [IDEA_GATEWAY_SKILL], } @@ -276,13 +324,21 @@ def build_terraform_scenario_report(root: Path) -> list[dict[str, Any]]: "Terminal direct execute": ["result", "evidence", "risk"], "Define checkpoint": ["gate", "brief", "validation", "risk", "checkpoint"], "Plan checkpoint": ["decision", "validation", "risk", "checkpoint"], - "Non-terminal apply-plan stop": ["state", "continuation", "user_action", "evidence", "next_step"], + "Non-terminal apply-plan stop": [ + "state", + "continuation", + "user_action", + "evidence", + "next_step", + ], "Review verdict": ["finding", "confidence", "evidence_gap", "risk", "route"], } def build_gateway_report(root: Path) -> dict[str, Any]: - core_bytes = len((root / ".github" / "skills" / GATEWAY_SKILL / "SKILL.md").read_bytes()) + core_bytes = len( + (root / ".github" / "skills" / GATEWAY_SKILL / "SKILL.md").read_bytes() + ) bundle_dir = root / ".github" / "skills" / GATEWAY_SKILL bundle_bytes = sum( len(p.read_bytes()) for p in bundle_dir.rglob("*") if p.is_file() @@ -296,28 +352,35 @@ def build_gateway_report(root: Path) -> dict[str, Any]: skill_path = root / ".github" / "skills" / skill_name / "SKILL.md" if skill_path.exists(): total_bytes += len(skill_path.read_bytes()) - context_scenarios.append({ - "scenario": name, - "required_skills": deduped, - "bytes": total_bytes, - "estimated_tokens": (total_bytes + ESTIMATED_TOKEN_BYTES - 1) // ESTIMATED_TOKEN_BYTES, - }) + context_scenarios.append( + { + "scenario": name, + "required_skills": deduped, + "bytes": total_bytes, + "estimated_tokens": (total_bytes + ESTIMATED_TOKEN_BYTES - 1) + // ESTIMATED_TOKEN_BYTES, + } + ) output_scenarios: list[dict[str, Any]] = [] for name, fields in GATEWAY_OUTPUT_FIELD_SCENARIOS.items(): field_bytes = sum(len(f.encode("utf-8")) for f in fields) - output_scenarios.append({ - "scenario": name, - "fields": fields, - "field_count": len(fields), - "field_bytes": field_bytes, - }) + output_scenarios.append( + { + "scenario": name, + "fields": fields, + "field_count": len(fields), + "field_bytes": field_bytes, + } + ) return { "core_bytes": core_bytes, - "core_estimated_tokens": (core_bytes + ESTIMATED_TOKEN_BYTES - 1) // ESTIMATED_TOKEN_BYTES, + "core_estimated_tokens": (core_bytes + ESTIMATED_TOKEN_BYTES - 1) + // ESTIMATED_TOKEN_BYTES, "bundle_bytes": bundle_bytes, - "bundle_estimated_tokens": (bundle_bytes + ESTIMATED_TOKEN_BYTES - 1) // ESTIMATED_TOKEN_BYTES, + "bundle_estimated_tokens": (bundle_bytes + ESTIMATED_TOKEN_BYTES - 1) + // ESTIMATED_TOKEN_BYTES, "required_context_scenarios": context_scenarios, "output_field_scenarios": output_scenarios, } @@ -327,9 +390,11 @@ def build_idea_gateway_report(root: Path) -> dict[str, Any]: core_path = root / ".github" / "skills" / IDEA_GATEWAY_SKILL / "SKILL.md" core_bytes = len(core_path.read_bytes()) if core_path.exists() else 0 bundle_dir = root / ".github" / "skills" / IDEA_GATEWAY_SKILL - bundle_bytes = sum( - len(p.read_bytes()) for p in bundle_dir.rglob("*") if p.is_file() - ) if bundle_dir.exists() else 0 + bundle_bytes = ( + sum(len(p.read_bytes()) for p in bundle_dir.rglob("*") if p.is_file()) + if bundle_dir.exists() + else 0 + ) context_scenarios: list[dict[str, Any]] = [] for name, skills in IDEA_GATEWAY_SCENARIOS.items(): @@ -339,18 +404,23 @@ def build_idea_gateway_report(root: Path) -> dict[str, Any]: skill_path = root / ".github" / "skills" / skill_name / "SKILL.md" if skill_path.exists(): total_bytes += len(skill_path.read_bytes()) - context_scenarios.append({ - "scenario": name, - "required_skills": deduped, - "bytes": total_bytes, - "estimated_tokens": (total_bytes + ESTIMATED_TOKEN_BYTES - 1) // ESTIMATED_TOKEN_BYTES, - }) + context_scenarios.append( + { + "scenario": name, + "required_skills": deduped, + "bytes": total_bytes, + "estimated_tokens": (total_bytes + ESTIMATED_TOKEN_BYTES - 1) + // ESTIMATED_TOKEN_BYTES, + } + ) return { "core_bytes": core_bytes, - "core_estimated_tokens": (core_bytes + ESTIMATED_TOKEN_BYTES - 1) // ESTIMATED_TOKEN_BYTES, + "core_estimated_tokens": (core_bytes + ESTIMATED_TOKEN_BYTES - 1) + // ESTIMATED_TOKEN_BYTES, "bundle_bytes": bundle_bytes, - "bundle_estimated_tokens": (bundle_bytes + ESTIMATED_TOKEN_BYTES - 1) // ESTIMATED_TOKEN_BYTES, + "bundle_estimated_tokens": (bundle_bytes + ESTIMATED_TOKEN_BYTES - 1) + // ESTIMATED_TOKEN_BYTES, "context_scenarios": context_scenarios, } @@ -376,12 +446,16 @@ def build_description_report(root: Path) -> list[dict[str, Any]]: desc_chars = len(description) desc_tokens = (desc_chars + ESTIMATED_TOKEN_BYTES - 1) // ESTIMATED_TOKEN_BYTES - reports.append({ - "skill": skill_name, - "description_chars": desc_chars, - "description_tokens": desc_tokens, - "description": description[:100] + "..." if len(description) > 100 else description, - }) + reports.append( + { + "skill": skill_name, + "description_chars": desc_chars, + "description_tokens": desc_tokens, + "description": description[:100] + "..." + if len(description) > 100 + else description, + } + ) return reports @@ -423,7 +497,9 @@ def main(argv: list[str] | None = None) -> int: "summary": { "total_scenarios": len(scenario_reports) + len(terraform_scenario_reports), "total_skills_measured": len(description_reports), - "highest_description_tokens": description_reports[0]["description_tokens"] if description_reports else 0, + "highest_description_tokens": description_reports[0]["description_tokens"] + if description_reports + else 0, "highest_scenario_proxy": max( [ *(r["scenario_proxy"] for r in scenario_reports), diff --git a/.github/scripts/detect_token_risks.py b/.github/scripts/detect_token_risks.py deleted file mode 100644 index eb1fd871..00000000 --- a/.github/scripts/detect_token_risks.py +++ /dev/null @@ -1,47 +0,0 @@ -#!/usr/bin/env python3 -"""Purpose: detect token-heavy overlap and duplication in Copilot governance assets. - -Usage examples: - python3 ./.github/scripts/detect_token_risks.py --root . - python3 ./.github/scripts/detect_token_risks.py --root . --strict --format json -""" - -from __future__ import annotations - -import argparse -from pathlib import Path - -from lib.cli_runner import run_finding_cli, should_fail -from lib.shared import Finding, find_repo_root, log_warn -from lib.token_risks import detect_token_risks - - -def parse_args() -> argparse.Namespace: - parser = argparse.ArgumentParser(description="Detect token efficiency risks in Copilot governance assets.") - parser.add_argument("--root", default=".", help="Repository root or any path inside it.") - parser.add_argument("--strict", action="store_true", help="Return a non-zero exit code when any finding is reported.") - parser.add_argument("--format", choices=["text", "json"], default="text", help="Output format.") - return parser.parse_args() - - -def main() -> int: - args = parse_args() - root = find_repo_root(Path(args.root)) - findings = run_finding_cli( - detect_fn=lambda: detect_token_risks(root), - format_name=args.format, - render_text=render_text, - ) - return 1 if should_fail(findings, strict=args.strict, blocking_severity=None) else 0 - - -def render_text(findings: list[Finding]) -> None: - if not findings: - return - for finding in findings: - log_warn(f"{finding.path} :: {finding.code} :: {finding.message}") - print(f" Suggestion: {finding.suggestion}") - - -if __name__ == "__main__": - raise SystemExit(main()) diff --git a/.github/scripts/graphify-file-change-hook.sh b/.github/scripts/graphify-file-change-hook.sh index e3df307a..88d08e3f 100755 --- a/.github/scripts/graphify-file-change-hook.sh +++ b/.github/scripts/graphify-file-change-hook.sh @@ -15,11 +15,11 @@ shift || true cd "$REPO_ROOT" if ! command -v graphify >/dev/null 2>&1; then - exit 0 + exit 0 fi if [ ! -d "$REPO_ROOT/graphify-out" ]; then - exit 0 + exit 0 fi TMP_ROOT="${TMPDIR:-/tmp}" @@ -32,84 +32,84 @@ LOG_FILE="$STATE_DIR/$REPO_ID.log" NEEDS_UPDATE_FILE="$REPO_ROOT/graphify-out/needs_update" if ! mkdir "$LOCK_DIR" 2>/dev/null; then - exit 0 + exit 0 fi trap 'rmdir "$LOCK_DIR" 2>/dev/null || true' EXIT is_doc_like_path() { - case "$1" in - ""|.git/*|graphify-out/*|tmp/*) - return 1 - ;; - *.md|*.mdx|*.rst|*.adoc|*.txt|*.pdf|*.png|*.jpg|*.jpeg|*.webp|*.svg|*.gif) - return 0 - ;; - docs/*|raw/*) - return 0 - ;; - esac + case "$1" in + "" | .git/* | graphify-out/* | tmp/*) return 1 + ;; + *.md | *.mdx | *.rst | *.adoc | *.txt | *.pdf | *.png | *.jpg | *.jpeg | *.webp | *.svg | *.gif) + return 0 + ;; + docs/* | raw/*) + return 0 + ;; + esac + return 1 } is_relevant_path() { - case "$1" in - ""|.git/*|graphify-out/*|tmp/*) - return 1 - ;; - esac - return 0 + case "$1" in + "" | .git/* | graphify-out/* | tmp/*) + return 1 + ;; + esac + return 0 } list_changed_paths() { - case "$EVENT" in - post-commit) - if git rev-parse --verify HEAD^ >/dev/null 2>&1; then - git diff-tree --no-commit-id --name-only -r HEAD - else - git ls-tree -r --name-only HEAD - fi - ;; - post-checkout) - local old_ref="${1:-}" - local new_ref="${2:-}" - if [ -n "$old_ref" ] && [ -n "$new_ref" ] \ - && git rev-parse --verify "$old_ref" >/dev/null 2>&1 \ - && git rev-parse --verify "$new_ref" >/dev/null 2>&1; then - git diff --name-only "$old_ref" "$new_ref" - fi - ;; - post-merge) - git diff-tree --no-commit-id --name-only -r HEAD - ;; - *) - return 0 - ;; - esac + case "$EVENT" in + post-commit) + if git rev-parse --verify HEAD^ >/dev/null 2>&1; then + git diff-tree --no-commit-id --name-only -r HEAD + else + git ls-tree -r --name-only HEAD + fi + ;; + post-checkout) + local old_ref="${1:-}" + local new_ref="${2:-}" + if [ -n "$old_ref" ] && [ -n "$new_ref" ] && + git rev-parse --verify "$old_ref" >/dev/null 2>&1 && + git rev-parse --verify "$new_ref" >/dev/null 2>&1; then + git diff --name-only "$old_ref" "$new_ref" + fi + ;; + post-merge) + git diff-tree --no-commit-id --name-only -r HEAD + ;; + *) + return 0 + ;; + esac } code_changes=0 doc_changes=0 while IFS= read -r path; do - if ! is_relevant_path "$path"; then - continue - fi - if is_doc_like_path "$path"; then - doc_changes=1 - else - code_changes=1 - fi + if ! is_relevant_path "$path"; then + continue + fi + if is_doc_like_path "$path"; then + doc_changes=1 + else + code_changes=1 + fi done < <(list_changed_paths "$@" || true) if [ "$doc_changes" -eq 1 ]; then - touch "$NEEDS_UPDATE_FILE" + touch "$NEEDS_UPDATE_FILE" fi if [ "$code_changes" -ne 1 ]; then - exit 0 + exit 0 fi if ! graphify update . >>"$LOG_FILE" 2>&1; then - touch "$NEEDS_UPDATE_FILE" - printf '[graphify-hook] update failed; see %s\n' "$LOG_FILE" >&2 + touch "$NEEDS_UPDATE_FILE" + printf '[graphify-hook] update failed; see %s\n' "$LOG_FILE" >&2 fi diff --git a/.github/scripts/install-graphify-hooks.sh b/.github/scripts/install-graphify-hooks.sh index f8ee5529..66587fbf 100755 --- a/.github/scripts/install-graphify-hooks.sh +++ b/.github/scripts/install-graphify-hooks.sh @@ -17,65 +17,65 @@ HOOK_MARKER="# graphify-hook: managed delegate" mkdir -p "$HOOKS_DIR" install_hook() { - local hook_name="$1" - local hook_path="$HOOKS_DIR/$hook_name" - local original_path="$hook_path.graphify-original" - local temporary_path + local hook_name="$1" + local hook_path="$HOOKS_DIR/$hook_name" + local original_path="$hook_path.graphify-original" + local temporary_path - if [ -L "$hook_path" ]; then - printf 'Preserved foreign symlink hook: %s\n' "$hook_path" - return 0 - fi + if [ -L "$hook_path" ]; then + printf 'Preserved foreign symlink hook: %s\n' "$hook_path" + return 0 + fi - if [ -f "$hook_path" ] && ! grep -Fq "$HOOK_MARKER" "$hook_path"; then - if [ -e "$original_path" ] || [ -L "$original_path" ]; then - printf 'Preserved foreign hook with existing backup: %s\n' "$hook_path" - return 0 - fi - mv "$hook_path" "$original_path" + if [ -f "$hook_path" ] && ! grep -Fq "$HOOK_MARKER" "$hook_path"; then + if [ -e "$original_path" ] || [ -L "$original_path" ]; then + printf 'Preserved foreign hook with existing backup: %s\n' "$hook_path" + return 0 fi + mv "$hook_path" "$original_path" + fi - if [ -f "$hook_path" ] && grep -Fq "$HOOK_MARKER" "$hook_path"; then - return 0 - fi + if [ -f "$hook_path" ] && grep -Fq "$HOOK_MARKER" "$hook_path"; then + return 0 + fi - temporary_path="$(mktemp "$HOOKS_DIR/.${hook_name}.XXXXXX")" - # The generated hook must expand these variables when it runs, not while it is written. - # shellcheck disable=SC2016 - { - printf '%s\n' '#!/usr/bin/env bash' - printf '%s\n' "$HOOK_MARKER" - printf '%s\n' 'set -Eeuo pipefail' - printf '%s\n' 'SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"' - printf 'ORIGINAL_HOOK="$SCRIPT_DIR/%s.graphify-original"\n' "$hook_name" - printf '%s\n' 'original_status=0' - printf '%s\n' 'if [ -f "$ORIGINAL_HOOK" ]; then' - printf '%s\n' ' set +e' - printf '%s\n' ' if [ -x "$ORIGINAL_HOOK" ]; then' - printf '%s\n' ' "$ORIGINAL_HOOK" "$@"' - printf '%s\n' ' else' - printf '%s\n' ' bash "$ORIGINAL_HOOK" "$@"' - printf '%s\n' ' fi' - printf '%s\n' ' original_status=$?' - printf '%s\n' ' set -e' - printf '%s\n' 'fi' - printf '%s\n' 'set +e' - printf '%s ' '"$SCRIPT_DIR/../scripts/graphify-file-change-hook.sh"' - printf '%s ' "$hook_name" - printf '%s\n' '"$@"' - printf '%s\n' 'delegate_status=$?' - printf '%s\n' 'set -e' - printf '%s\n' 'if [ "$original_status" -ne 0 ]; then' - printf '%s\n' ' exit "$original_status"' - printf '%s\n' 'fi' - printf '%s\n' 'exit "$delegate_status"' - } >"$temporary_path" - chmod 755 "$temporary_path" - mv "$temporary_path" "$hook_path" + temporary_path="$(mktemp "$HOOKS_DIR/.${hook_name}.XXXXXX")" + # The generated hook must expand these variables when it runs, not while it is written. + # shellcheck disable=SC2016 + { + printf '%s\n' '#!/usr/bin/env bash' + printf '%s\n' "$HOOK_MARKER" + printf '%s\n' 'set -Eeuo pipefail' + printf '%s\n' 'SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"' + printf 'ORIGINAL_HOOK="$SCRIPT_DIR/%s.graphify-original"\n' "$hook_name" + printf '%s\n' 'original_status=0' + printf '%s\n' 'if [ -f "$ORIGINAL_HOOK" ]; then' + printf '%s\n' ' set +e' + printf '%s\n' ' if [ -x "$ORIGINAL_HOOK" ]; then' + printf '%s\n' ' "$ORIGINAL_HOOK" "$@"' + printf '%s\n' ' else' + printf '%s\n' ' bash "$ORIGINAL_HOOK" "$@"' + printf '%s\n' ' fi' + printf '%s\n' ' original_status=$?' + printf '%s\n' ' set -e' + printf '%s\n' 'fi' + printf '%s\n' 'set +e' + printf '%s ' '"$SCRIPT_DIR/../scripts/graphify-file-change-hook.sh"' + printf '%s ' "$hook_name" + printf '%s\n' '"$@"' + printf '%s\n' 'delegate_status=$?' + printf '%s\n' 'set -e' + printf '%s\n' 'if [ "$original_status" -ne 0 ]; then' + printf '%s\n' ' exit "$original_status"' + printf '%s\n' 'fi' + printf '%s\n' 'exit "$delegate_status"' + } >"$temporary_path" + chmod 755 "$temporary_path" + mv "$temporary_path" "$hook_path" } for hook_name in post-commit post-checkout post-merge; do - install_hook "$hook_name" + install_hook "$hook_name" done printf 'Installed Git hooks path: %s\n' '.github/hooks' diff --git a/.github/scripts/lib/__init__.py b/.github/scripts/lib/__init__.py deleted file mode 100644 index 6e6a5de1..00000000 --- a/.github/scripts/lib/__init__.py +++ /dev/null @@ -1 +0,0 @@ -"""Shared helpers for Copilot maintenance scripts.""" diff --git a/.github/scripts/lib/inventory.py b/.github/scripts/lib/inventory.py deleted file mode 100644 index 21d1877c..00000000 --- a/.github/scripts/lib/inventory.py +++ /dev/null @@ -1,124 +0,0 @@ -from __future__ import annotations - -import re -from pathlib import Path - -from .shared import INVENTORY_PATH, path_list, write_text - -SECTION_ORDER = ("Instructions", "Skills", "Scripts", "Agents", "Prompts") -EMPTY_MESSAGES = { - "Instructions": "No instruction files currently ship in the live catalog.", - "Skills": "No skill files currently ship in the live catalog.", - "Scripts": "No script files currently ship in the live catalog.", - "Agents": "No agent files currently ship in the live catalog.", - "Prompts": "No prompt files currently ship in the live catalog.", -} -SCRIPT_GLOB_PATTERNS = ( - ".github/scripts/*.py", - ".github/scripts/*.sh", - ".github/scripts/lib/*.py", -) -DOCUMENT_SUPPORT_ONLY_SKILLS = ( - ".github/skills/anthropic-docx/SKILL.md", - ".github/skills/anthropic-pdf/SKILL.md", - ".github/skills/anthropic-pptx/SKILL.md", - ".github/skills/anthropic-xlsx/SKILL.md", -) -IGNORED_SCRIPT_BASENAMES = {"__init__.py"} - - -def collect_inventory_sections(root: Path) -> dict[str, list[str]]: - return { - "Instructions": path_list(root, ".github/instructions/**/*.instructions.md"), - "Skills": path_list(root, ".github/skills/**/SKILL.md"), - "Scripts": _collect_script_paths(root), - "Agents": path_list(root, ".github/agents/*.agent.md"), - "Prompts": path_list(root, ".github/prompts/*.prompt.md"), - } - - -def _collect_script_paths(root: Path) -> list[str]: - entries: set[str] = set() - for pattern in SCRIPT_GLOB_PATTERNS: - for path in root.glob(pattern): - if not path.is_file() or path.name in IGNORED_SCRIPT_BASENAMES: - continue - entries.add(path.relative_to(root).as_posix()) - return sorted(entries) - - -def sections_from_catalog_paths(paths: list[str]) -> dict[str, list[str]]: - sections = {section: [] for section in SECTION_ORDER} - for relative_path in sorted(paths): - if relative_path.startswith(".github/instructions/") and relative_path.endswith(".instructions.md"): - sections["Instructions"].append(relative_path) - elif relative_path.startswith(".github/skills/") and relative_path.endswith("/SKILL.md"): - sections["Skills"].append(relative_path) - elif ( - relative_path.startswith(".github/scripts/") - and (relative_path.endswith(".py") or relative_path.endswith(".sh")) - and not relative_path.endswith("__init__.py") - ): - sections["Scripts"].append(relative_path) - elif relative_path.startswith(".github/agents/") and relative_path.endswith(".agent.md"): - sections["Agents"].append(relative_path) - elif relative_path.startswith(".github/prompts/") and relative_path.endswith(".prompt.md"): - sections["Prompts"].append(relative_path) - return {section: sorted(entries) for section, entries in sections.items()} - - -def render_inventory_markdown(sections: dict[str, list[str]]) -> str: - lines = [ - "# Copilot Inventory", - "", - "This file is the exact path inventory for the live GitHub Copilot catalog in this repository.", - "", - ] - for section in SECTION_ORDER: - lines.append(f"## {section}") - lines.append("") - entries = sections.get(section, []) - if entries: - lines.extend(f"- `{entry}`" for entry in entries) - if section == "Skills": - doc_entries = [entry for entry in entries if entry in DOCUMENT_SUPPORT_ONLY_SKILLS] - if doc_entries: - lines.append("") - lines.append("### Support-only imported document skills") - lines.append("") - lines.append( - "These vendor-prefixed imported document skills remain support-only depth for repositories that explicitly need document workflows." - ) - else: - lines.append(EMPTY_MESSAGES[section]) - lines.append("") - return "\n".join(lines).rstrip() + "\n" - - -def build_inventory_markdown(root: Path) -> str: - return render_inventory_markdown(collect_inventory_sections(root)) - - -def write_inventory(root: Path) -> Path: - inventory_path = root / INVENTORY_PATH - write_text(inventory_path, build_inventory_markdown(root)) - return inventory_path - - -def parse_inventory_markdown(text: str) -> dict[str, set[str]]: - section_lookup = {section.lower(): section for section in SECTION_ORDER} - sections = {section: set() for section in SECTION_ORDER} - current_section: str | None = None - - for line in text.splitlines(): - if line.startswith("## "): - heading = line[3:].strip().lower() - current_section = section_lookup.get(heading) - continue - if current_section is None: - continue - match = re.match(r"^- `([^`]+)`$", line.strip()) - if match: - sections[current_section].add(match.group(1)) - - return sections diff --git a/.github/scripts/lib/jsonc.py b/.github/scripts/lib/jsonc.py deleted file mode 100644 index cbf521e5..00000000 --- a/.github/scripts/lib/jsonc.py +++ /dev/null @@ -1,333 +0,0 @@ -from __future__ import annotations - -import re -from dataclasses import dataclass - -from .shared import VSCODE_COPILOT_SETTINGS - - -@dataclass(frozen=True) -class JsoncObjectMember: - key: str - value_kind: str - value_start: int - value_end: int - object_value: JsoncObjectInfo | None - has_trailing_comma: bool - - -@dataclass(frozen=True) -class JsoncObjectInfo: - start: int - end: int - members: tuple[JsoncObjectMember, ...] - - -class JsoncParseError(ValueError): - pass - - -def skip_jsonc_space_and_comments(text: str, index: int) -> int: - length = len(text) - cursor = index - while cursor < length: - char = text[cursor] - if char in {" ", "\t", "\n", "\r"}: - cursor += 1 - continue - if text.startswith("//", cursor): - cursor += 2 - while cursor < length and text[cursor] not in {"\n", "\r"}: - cursor += 1 - continue - if text.startswith("/*", cursor): - end = text.find("*/", cursor + 2) - if end == -1: - raise JsoncParseError("unterminated block comment") - cursor = end + 2 - continue - break - return cursor - - -def consume_jsonc_string(text: str, index: int) -> tuple[str, int]: - if index >= len(text) or text[index] != '"': - raise JsoncParseError("expected string") - cursor = index + 1 - chars: list[str] = [] - while cursor < len(text): - char = text[cursor] - if char == '"': - return "".join(chars), cursor + 1 - if char == "\\": - cursor += 1 - if cursor >= len(text): - raise JsoncParseError("unterminated string escape") - escaped = text[cursor] - if escaped == "u": - hex_chunk = text[cursor + 1 : cursor + 5] - if len(hex_chunk) != 4 or not all( - glyph in "0123456789abcdefABCDEF" for glyph in hex_chunk - ): - raise JsoncParseError("invalid unicode escape") - chars.append(chr(int(hex_chunk, 16))) - cursor += 5 - continue - escaped_map = { - '"': '"', - "\\": "\\", - "/": "/", - "b": "\b", - "f": "\f", - "n": "\n", - "r": "\r", - "t": "\t", - } - if escaped not in escaped_map: - raise JsoncParseError("invalid escape sequence") - chars.append(escaped_map[escaped]) - cursor += 1 - continue - chars.append(char) - cursor += 1 - raise JsoncParseError("unterminated string") - - -def consume_jsonc_scalar(text: str, index: int) -> int: - cursor = index - while cursor < len(text): - char = text[cursor] - if char in {",", "}", "]"}: - break - if char in {" ", "\t", "\n", "\r"}: - break - if text.startswith("//", cursor) or text.startswith("/*", cursor): - break - cursor += 1 - if cursor == index: - raise JsoncParseError("expected scalar value") - return cursor - - -def parse_jsonc_value( - text: str, - index: int, -) -> tuple[str, int, int, JsoncObjectInfo | None, int]: - cursor = skip_jsonc_space_and_comments(text, index) - if cursor >= len(text): - raise JsoncParseError("expected value") - - char = text[cursor] - if char == "{": - object_info, next_cursor = parse_jsonc_object(text, cursor) - return "object", cursor, next_cursor, object_info, next_cursor - if char == "[": - next_cursor = parse_jsonc_array(text, cursor) - return "array", cursor, next_cursor, None, next_cursor - if char == '"': - _, next_cursor = consume_jsonc_string(text, cursor) - return "string", cursor, next_cursor, None, next_cursor - - next_cursor = consume_jsonc_scalar(text, cursor) - token = text[cursor:next_cursor] - if token in {"true", "false", "null"}: - return "literal", cursor, next_cursor, None, next_cursor - - if re.fullmatch(r"-?(0|[1-9]\d*)(\.\d+)?([eE][+-]?\d+)?", token): - return "number", cursor, next_cursor, None, next_cursor - raise JsoncParseError(f"invalid token '{token}'") - - -def parse_jsonc_array(text: str, index: int) -> int: - cursor = index + 1 - expect_value = True - while True: - cursor = skip_jsonc_space_and_comments(text, cursor) - if cursor >= len(text): - raise JsoncParseError("unterminated array") - if text[cursor] == ']': - return cursor + 1 - if not expect_value: - raise JsoncParseError("expected ',' or ']' in array") - - _, _, _, _, cursor = parse_jsonc_value(text, cursor) - cursor = skip_jsonc_space_and_comments(text, cursor) - if cursor >= len(text): - raise JsoncParseError("unterminated array") - if text[cursor] == ',': - cursor += 1 - expect_value = True - continue - if text[cursor] == ']': - return cursor + 1 - raise JsoncParseError("expected ',' or ']' in array") - - -def parse_jsonc_object(text: str, index: int) -> tuple[JsoncObjectInfo, int]: - cursor = index + 1 - members: list[JsoncObjectMember] = [] - - while True: - cursor = skip_jsonc_space_and_comments(text, cursor) - if cursor >= len(text): - raise JsoncParseError("unterminated object") - if text[cursor] == '}': - return JsoncObjectInfo(start=index, end=cursor, members=tuple(members)), cursor + 1 - - key, after_key = consume_jsonc_string(text, cursor) - cursor = skip_jsonc_space_and_comments(text, after_key) - if cursor >= len(text) or text[cursor] != ':': - raise JsoncParseError("expected ':' after object key") - cursor += 1 - - value_kind, value_start, value_end, value_object, cursor = parse_jsonc_value( - text, cursor - ) - cursor = skip_jsonc_space_and_comments(text, cursor) - has_trailing_comma = False - if cursor < len(text) and text[cursor] == ',': - has_trailing_comma = True - cursor += 1 - - members.append( - JsoncObjectMember( - key=key, - value_kind=value_kind, - value_start=value_start, - value_end=value_end, - object_value=value_object, - has_trailing_comma=has_trailing_comma, - ) - ) - - if has_trailing_comma: - continue - - cursor = skip_jsonc_space_and_comments(text, cursor) - if cursor >= len(text): - raise JsoncParseError("unterminated object") - if text[cursor] == '}': - return JsoncObjectInfo(start=index, end=cursor, members=tuple(members)), cursor + 1 - raise JsoncParseError("expected ',' or '}' in object") - - -def parse_jsonc_root_object(text: str) -> JsoncObjectInfo: - cursor = skip_jsonc_space_and_comments(text, 0) - if cursor >= len(text) or text[cursor] != '{': - raise JsoncParseError("settings content must start with a JSON object") - root, next_cursor = parse_jsonc_object(text, cursor) - trailing = skip_jsonc_space_and_comments(text, next_cursor) - if trailing != len(text): - raise JsoncParseError("unexpected trailing content after root object") - return root - - -def object_members_by_key(object_info: JsoncObjectInfo, key: str) -> list[JsoncObjectMember]: - return [member for member in object_info.members if member.key == key] - - -def object_indent(text: str, object_info: JsoncObjectInfo) -> str: - line_start = text.rfind("\n", 0, object_info.start) + 1 - prefix = text[line_start:object_info.start] - return re.match(r"[ \t]*", prefix).group(0) if prefix else "" - - -def member_indent(text: str, object_info: JsoncObjectInfo) -> str: - for member in object_info.members: - line_start = text.rfind("\n", 0, member.value_start) + 1 - indent = re.match(r"[ \t]*", text[line_start:member.value_start]).group(0) - if indent: - return indent - return f"{object_indent(text, object_info)} " - - -def insert_object_member(text: str, object_info: JsoncObjectInfo, member_text: str) -> str: - indent = member_indent(text, object_info) - if not object_info.members: - base_indent = object_indent(text, object_info) - insertion = f"\n{indent}{member_text}\n{base_indent}" - return text[: object_info.start + 1] + insertion + text[object_info.end :] - - last_member = object_info.members[-1] - separator = "" if last_member.has_trailing_comma else "," - insertion = f"{separator}\n{indent}{member_text}" - return text[: object_info.end] + insertion + text[object_info.end :] - - -def ensure_root_boolean_setting(text: str, root: JsoncObjectInfo, key: str) -> str: - matches = object_members_by_key(root, key) - if len(matches) > 1: - raise JsoncParseError(f"duplicate key '{key}' is not safe to auto-merge") - if not matches: - return insert_object_member(text, root, f'"{key}": false') - - member = matches[0] - value = text[member.value_start : member.value_end].strip() - if value == "false": - return text - if value != "true": - raise JsoncParseError( - f"managed key '{key}' must be a boolean for safe merge" - ) - return text[: member.value_start] + "false" + text[member.value_end :] - - -def ensure_nested_boolean_setting( - text: str, - root: JsoncObjectInfo, - object_key: str, - nested_key: str, -) -> str: - matches = object_members_by_key(root, object_key) - if len(matches) > 1: - raise JsoncParseError( - f"duplicate key '{object_key}' is not safe to auto-merge" - ) - - if not matches: - nested_object = f'"{object_key}": {{\n "{nested_key}": false\n }}' - return insert_object_member(text, root, nested_object) - - member = matches[0] - if member.value_kind != "object" or member.object_value is None: - raise JsoncParseError( - f"managed key '{object_key}' must be an object for safe merge" - ) - - nested_matches = object_members_by_key(member.object_value, nested_key) - if len(nested_matches) > 1: - raise JsoncParseError( - f"duplicate key '{nested_key}' inside '{object_key}' is not safe to auto-merge" - ) - - if not nested_matches: - return insert_object_member(text, member.object_value, f'"{nested_key}": false') - - nested_member = nested_matches[0] - nested_value = text[nested_member.value_start : nested_member.value_end].strip() - if nested_value == "false": - return text - if nested_value != "true": - raise JsoncParseError( - f"managed key '{object_key}.{nested_key}' must be a boolean for safe merge" - ) - return text[: nested_member.value_start] + "false" + text[nested_member.value_end :] - - -def apply_managed_vscode_copilot_settings(content: str) -> str: - updated = content - root = parse_jsonc_root_object(updated) - updated = ensure_root_boolean_setting( - updated, - root, - VSCODE_COPILOT_SETTINGS[0][0][0], - ) - - root = parse_jsonc_root_object(updated) - updated = ensure_nested_boolean_setting( - updated, - root, - VSCODE_COPILOT_SETTINGS[1][0][0], - VSCODE_COPILOT_SETTINGS[1][0][1], - ) - return updated diff --git a/.github/scripts/lib/repo_paths.py b/.github/scripts/lib/repo_paths.py deleted file mode 100644 index da0625c8..00000000 --- a/.github/scripts/lib/repo_paths.py +++ /dev/null @@ -1,13 +0,0 @@ -"""Shared repository root discovery.""" - -from __future__ import annotations - -from pathlib import Path - - -def find_repo_root(start: Path) -> Path: - candidate = start.resolve() - for current in (candidate, *candidate.parents): - if (current / ".github").is_dir(): - return current - raise FileNotFoundError(f"Unable to find repository root from {start}") diff --git a/.github/scripts/lib/shared.py b/.github/scripts/lib/shared.py deleted file mode 100644 index 50415e2b..00000000 --- a/.github/scripts/lib/shared.py +++ /dev/null @@ -1,444 +0,0 @@ -from __future__ import annotations - -import hashlib -import json -import re -import subprocess -from dataclasses import dataclass -from pathlib import Path -from typing import Iterable, Iterator - -import yaml - -FRONTMATTER_PATTERN = re.compile(r"\A---\s*\n(.*?)\n---\s*\n?", re.DOTALL) -LEGACY_AGENT_TOOL_IDS = { - "terminalCommand", - "search/codebase", - "search/searchResults", - "search/usages", - "edit/editFiles", - "execute/runInTerminal", - "web/fetch", - "read/problems", -} -IGNORED_SYNC_FILENAMES = {"README.md", "CHANGELOG.md"} -IGNORED_SYNC_PARTS = {"__pycache__", ".venv"} -CONSUMER_SYNC_EXCLUDED_PREFIX = "internal-sync-" -CONSUMER_SYNC_EXCLUDED_PATH_PREFIXES = frozenset() -LESSONS_PATH = "LESSONS_LEARNED.md" -DOCS_README_PATH = "docs/README.md" -ARCHITECTURE_PATH = "docs/architecture.md" -REPOSITORY_CONTEXT_PATH = "docs/repository-context.md" -TECH_PATH = "docs/tech.md" -STRUCTURE_PATH = "docs/structure.md" -RETIRED_RUNTIME_OPERATING_MODEL_PATH = "docs/03-local-ai-runtime-operating-model.md" -LEGACY_LOCAL_ARCHITECTURE_PATH = "docs/01-local-architecture.md" -LEGACY_LOCAL_REPOSITORY_CONTEXT_PATH = "docs/02-local-repository-context.md" -LEGACY_ARCHITECTURE_PATH = "docs/01-architecture.md" -LEGACY_REPOSITORY_CONTEXT_PATH = "docs/02-repository-context.md" -LEGACY_RUNTIME_FIT_PATH = "docs/runtime-fit.md" -DOCS_README_TEMPLATE_PATH = ".github/templates/docs-README.md.template" -ARCHITECTURE_TEMPLATE_PATH = ".github/templates/architecture.md.template" -REPOSITORY_CONTEXT_TEMPLATE_PATH = ".github/templates/repository-context.md.template" -TECH_TEMPLATE_PATH = ".github/templates/tech.md.template" -STRUCTURE_TEMPLATE_PATH = ".github/templates/structure.md.template" -CONSUMER_LOCAL_KNOWLEDGE_TEMPLATES = { - DOCS_README_PATH: DOCS_README_TEMPLATE_PATH, - ARCHITECTURE_PATH: ARCHITECTURE_TEMPLATE_PATH, - REPOSITORY_CONTEXT_PATH: REPOSITORY_CONTEXT_TEMPLATE_PATH, - TECH_PATH: TECH_TEMPLATE_PATH, - STRUCTURE_PATH: STRUCTURE_TEMPLATE_PATH, -} -MANAGED_ROOT_FILES = ( - "AGENTS.md", - LESSONS_PATH, - ".editorconfig", - ".pre-commit-config.yaml", - ".github/copilot-instructions.md", - ".github/copilot-commit-message-instructions.md", - ".github/security-baseline.md", - ".github/DEPRECATION.md", - ".github/repo-profiles.yml", -) -MANAGED_WORKFLOW_FILES = (".github/workflows/_pre-commit.yml",) -VSCODE_SETTINGS_PATH = ".vscode/settings.json" -VSCODE_COPILOT_SETTINGS = ( - ( - ("github.copilot.chat.codeGeneration.useInstructionFiles",), - False, - ), - ( - ("chat.instructionsFilesLocations", ".github/instructions"), - False, - ), -) -INVENTORY_PATH = ".github/INVENTORY.md" -IMPORTED_ASSET_OVERRIDES_PATH = ".github/skills/local-agent-sync-external-resources/references/imported-asset-overrides.yaml" -MANAGED_EXTERNAL_RESOURCES_PATH = ".github/skills/local-agent-sync-external-resources/references/managed-resources.yaml" - - -@dataclass(frozen=True) -class Finding: - severity: str - code: str - path: str - message: str - suggestion: str - - def to_dict(self) -> dict[str, str]: - return { - "severity": self.severity, - "code": self.code, - "path": self.path, - "message": self.message, - "suggestion": self.suggestion, - } - - -@dataclass(frozen=True) -class SyncOperation: - action: str - path: str - reason: str - source_hash: str | None = None - target_hash: str | None = None - - def to_dict(self) -> dict[str, str | None]: - return { - "action": self.action, - "path": self.path, - "reason": self.reason, - "source_hash": self.source_hash, - "target_hash": self.target_hash, - } - - -@dataclass(frozen=True) -class SyncPlan: - source_root: Path - target_root: Path - source_revision: str | None - source_version: str | None - target_manifest_source_version: str | None - target_dirty: bool - stacks: tuple[str, ...] - operations: tuple[SyncOperation, ...] - local_assets: tuple[str, ...] - generated_inventory: str - generated_lessons: str | None = None - generated_gitignore: str | None = None - dirty_paths: tuple[str, ...] = () - managed_mutation_paths: tuple[str, ...] = () - dirty_managed_overlap: tuple[str, ...] = () - - def to_dict(self) -> dict[str, object]: - return { - "source_root": self.source_root.as_posix(), - "target_root": self.target_root.as_posix(), - "source_revision": self.source_revision, - "source_version": self.source_version, - "target_manifest_source_version": self.target_manifest_source_version, - "target_dirty": self.target_dirty, - "stacks": list(self.stacks), - "local_assets": list(self.local_assets), - "dirty_paths": list(self.dirty_paths), - "managed_mutation_paths": list(self.managed_mutation_paths), - "dirty_managed_overlap": list(self.dirty_managed_overlap), - "operations": [operation.to_dict() for operation in self.operations], - } - - -def log_info(message: str) -> None: - print(f"ℹ️ {message}", flush=True) - - -def log_warn(message: str) -> None: - print(f"⚠️ {message}", flush=True) - - -def log_success(message: str) -> None: - print(f"✅ {message}", flush=True) - - -def log_error(message: str) -> None: - print(f"❌ {message}", flush=True) - - -def find_repo_root(start: Path) -> Path: - candidate = start.resolve() - for current in (candidate, *candidate.parents): - if (current / ".github").is_dir() or (current / ".git").exists(): - return current - raise FileNotFoundError(f"Unable to find repository root from {start}") - - -def read_text(path: Path) -> str: - return path.read_text(encoding="utf-8") - - -def write_text(path: Path, content: str) -> None: - path.parent.mkdir(parents=True, exist_ok=True) - path.write_text(content, encoding="utf-8") - - -def sha256_file(path: Path) -> str: - digest = hashlib.sha256() - with path.open("rb") as file_handle: - for chunk in iter(lambda: file_handle.read(65536), b""): - digest.update(chunk) - return digest.hexdigest() - - -def git_revision(root: Path) -> str | None: - try: - result = subprocess.run( - ["git", "rev-parse", "--short", "HEAD"], - cwd=root, - text=True, - capture_output=True, - check=False, - timeout=10, - ) - if result.returncode != 0: - return None - return result.stdout.strip() or None - except (subprocess.TimeoutExpired, OSError): - return None - - -def is_git_dirty(root: Path) -> bool: - try: - result = subprocess.run( - ["git", "status", "--porcelain"], - cwd=root, - text=True, - capture_output=True, - check=False, - timeout=10, - ) - if result.returncode != 0: - return False - return bool(result.stdout.strip()) - except (subprocess.TimeoutExpired, OSError): - return False - - -def git_dirty_paths(root: Path) -> list[str]: - try: - result = subprocess.run( - ["git", "status", "--porcelain=v1", "-z", "--untracked-files=all"], - cwd=root, - capture_output=True, - check=False, - timeout=10, - ) - except (subprocess.TimeoutExpired, OSError): - return [] - if result.returncode != 0 or not result.stdout: - return [] - - entries = result.stdout.decode("utf-8", errors="replace").split("\0") - dirty_paths: list[str] = [] - index = 0 - while index < len(entries): - entry = entries[index] - index += 1 - if not entry: - continue - - status = entry[:2] - path = entry[3:] - if path: - dirty_paths.append(path) - - if "R" in status or "C" in status: - if index < len(entries): - renamed_path = entries[index] - index += 1 - if renamed_path: - dirty_paths.append(renamed_path) - - return sorted(dedupe_preserve_order(dirty_paths)) - - -def split_frontmatter(text: str) -> tuple[dict[str, object], str]: - if not text.startswith("---"): - return {}, text - - match = FRONTMATTER_PATTERN.match(text) - if not match: - return {}, text - - frontmatter_text = match.group(1) - try: - parsed = yaml.safe_load(frontmatter_text) or {} - except yaml.YAMLError: - parsed = {} - if not isinstance(parsed, dict): - parsed = {} - return parsed, text[match.end() :] - - -def load_frontmatter(path: Path) -> dict[str, object]: - return split_frontmatter(read_text(path))[0] - - -def strip_frontmatter(text: str) -> str: - return split_frontmatter(text)[1] - - -def markdown_link_targets(text: str) -> list[str]: - return re.findall(r"\[[^\]]+\]\(([^)]+)\)", text) - - -def normalize_markdown_text(text: str) -> str: - normalized_lines: list[str] = [] - in_code_block = False - for raw_line in strip_frontmatter(text).splitlines(): - line = raw_line.rstrip() - if line.strip().startswith("```"): - in_code_block = not in_code_block - continue - if in_code_block: - continue - cleaned = re.sub(r"^[#>*\-\d\.)\s]+", "", line).replace("`", "") - cleaned = re.sub(r"\s+", " ", cleaned).strip().lower() - if cleaned: - normalized_lines.append(cleaned) - return "\n".join(normalized_lines) - - -def significant_text_lines(text: str) -> set[str]: - return { - line - for line in normalize_markdown_text(text).splitlines() - if len(line) >= 18 and not line.startswith("http") - } - - -def render_json(data: object) -> str: - return json.dumps(data, indent=2, sort_keys=True) - - -def iter_markdown_assets(root: Path) -> Iterator[Path]: - candidates = [root / "AGENTS.md"] - github_root = root / ".github" - if github_root.exists(): - candidates.extend( - path - for path in github_root.rglob("*.md") - if path.is_file() - and not any(part in IGNORED_SYNC_PARTS for part in path.parts) - ) - for path in candidates: - if path.exists(): - yield path - - -def is_ignored_sync_path(relative_path: str) -> bool: - path = Path(relative_path) - if path.name in IGNORED_SYNC_FILENAMES: - return True - if path.suffix == ".pyc": - return True - return any(part in IGNORED_SYNC_PARTS for part in path.parts) - - -def is_local_asset(relative_path: str) -> bool: - path = Path(relative_path) - parts = path.parts - if len(parts) < 3 or parts[0] != ".github": - return False - if parts[1] == "skills": - return len(parts) >= 3 and parts[2].startswith("local-") - return path.name.startswith("local-") - - -def is_imported_asset(relative_path: str) -> bool: - path = Path(relative_path) - parts = path.parts - if len(parts) < 3 or parts[0] != ".github": - return False - if parts[1] == "skills": - prefix = parts[2] - return not prefix.startswith(("internal-", "local-")) - if parts[1] in {"agents", "instructions"}: - return not path.name.startswith(("internal-", "local-")) - return False - - -def is_consumer_sync_excluded_path(relative_path: str) -> bool: - path = Path(relative_path) - if any(part.startswith(CONSUMER_SYNC_EXCLUDED_PREFIX) for part in path.parts): - return True - for excluded_prefix in CONSUMER_SYNC_EXCLUDED_PATH_PREFIXES: - if relative_path.startswith(excluded_prefix + "/") or relative_path == excluded_prefix: - return True - return False - - -def resolve_markdown_target(root: Path, current_file: Path, target: str) -> Path | None: - clean_target = target.split("#", maxsplit=1)[0].strip() - if not clean_target: - return None - if "://" in clean_target or clean_target.startswith(("mailto:", "file:")): - return None - if clean_target.startswith("/"): - return None - if clean_target.startswith((".github/", "docs/")) or clean_target == "AGENTS.md": - return root / clean_target - return (current_file.parent / clean_target).resolve() - - -def action_sort_key(action: str) -> int: - ordering = { - "create": 0, - "update": 1, - "rename": 2, - "ensure": 3, - "rebuild": 4, - "delete": 5, - "manual": 6, - "preserve": 7, - "unchanged": 8, - } - return ordering.get(action, 99) - - -def finding_sort_key(finding: Finding) -> tuple[int, str, str]: - severity_order = {"blocking": 0, "non-blocking": 1} - return (severity_order.get(finding.severity, 99), finding.path, finding.code) - - -def path_list(root: Path, pattern: str) -> list[str]: - return sorted( - path.relative_to(root).as_posix() - for path in root.glob(pattern) - if path.is_file() - ) - - -def all_files_under(root: Path, relative_dir: str) -> list[str]: - base_dir = root / relative_dir - if not base_dir.exists(): - return [] - results: list[str] = [] - for path in base_dir.rglob("*"): - if not path.is_file(): - continue - relative_path = path.relative_to(root).as_posix() - if is_ignored_sync_path(relative_path): - continue - results.append(relative_path) - return sorted(results) - - -def dedupe_preserve_order(values: Iterable[str]) -> list[str]: - seen: set[str] = set() - ordered: list[str] = [] - for value in values: - if value in seen: - continue - seen.add(value) - ordered.append(value) - return ordered diff --git a/.github/scripts/run.sh b/.github/scripts/run.sh deleted file mode 100755 index 79ba37aa..00000000 --- a/.github/scripts/run.sh +++ /dev/null @@ -1,227 +0,0 @@ -#!/usr/bin/env bash -# -# Purpose: Bootstrap the local Python environment and run a Copilot maintenance tool. -# Usage examples: -# ./.github/scripts/run.sh build_inventory --root . - -set -Eeuo pipefail - -log_info() { - printf 'ℹ️ %s\n' "$*" -} - -log_success() { - printf '✅ %s\n' "$*" -} - -log_error() { - printf '❌ %s\n' "$*" >&2 -} - -require_command() { - command -v "$1" >/dev/null 2>&1 || { - log_error "Missing required command: $1" - exit 1 - } -} - -usage() { - cat <<'EOF' -Usage: - ./.github/scripts/run.sh [tool-args...] - -Tools: - github_catalog_validation - analyze_copilot_debug_log - benchmark_skill_tokens - build_inventory - check_catalog_consistency - audit_copilot_catalog - detect_token_risks - sync_home_ai_resources - validate_internal_skills - validate_skill_change_scope -EOF -} - -hash_file() { - local file_path="$1" - if command -v sha256sum >/dev/null 2>&1; then - sha256sum "$file_path" | awk '{print $1}' - return - fi - if command -v shasum >/dev/null 2>&1; then - shasum -a 256 "$file_path" | awk '{print $1}' - return - fi - log_error "Missing SHA-256 helper: install sha256sum or shasum." - exit 1 -} - -load_required_python_version() { - if [[ ! -f "$PYTHON_VERSION_FILE" ]]; then - log_error "Missing required Python version file: $PYTHON_VERSION_FILE" - exit 1 - fi - - REQUIRED_PYTHON_VERSION="$(tr -d '[:space:]' <"$PYTHON_VERSION_FILE")" - REQUIRED_PYTHON_MAJOR_MINOR="$(printf '%s' "$REQUIRED_PYTHON_VERSION" | awk -F. 'NF >= 2 { print $1 "." $2 }')" - - if [[ -z "$REQUIRED_PYTHON_MAJOR_MINOR" ]]; then - log_error "Invalid Python version in $PYTHON_VERSION_FILE: $REQUIRED_PYTHON_VERSION" - exit 1 - fi -} - -select_python_bin() { - if [[ -n "$PYTHON_BIN" ]]; then - return - fi - - PYTHON_BIN="python$REQUIRED_PYTHON_MAJOR_MINOR" -} - -verify_python_bin_version() { - local actual_version - actual_version="$("$PYTHON_BIN" -c 'import sys; print(f"{sys.version_info.major}.{sys.version_info.minor}")')" - - if [[ "$actual_version" == "$REQUIRED_PYTHON_MAJOR_MINOR" ]]; then - return - fi - - log_error "$PYTHON_BIN resolved to Python $actual_version, but .python-version requires $REQUIRED_PYTHON_VERSION." - exit 1 -} - -verify_venv_version() { - local venv_python="$VENV_DIR/bin/python" - local venv_version - - if [[ ! -x "$venv_python" ]]; then - log_error "Virtual environment is missing its Python interpreter: $venv_python" - exit 1 - fi - - venv_version="$("$venv_python" -c 'import sys; print(f"{sys.version_info.major}.{sys.version_info.minor}")')" - if [[ "$venv_version" == "$REQUIRED_PYTHON_MAJOR_MINOR" ]]; then - return - fi - - log_error "Existing virtual environment uses Python $venv_version, but .python-version requires $REQUIRED_PYTHON_VERSION. Remove $VENV_DIR and rerun." - exit 1 -} - -resolve_script() { - local tool_name="$1" - case "$tool_name" in - github_catalog_validation|github_catalog_validation.py) - printf '%s\n' "$SCRIPT_DIR/github_catalog_validation.py" - ;; - analyze_copilot_debug_log|analyze_copilot_debug_log.sh) - printf '%s\n' "$REPO_ROOT/.github/skills/local-copilot-log-analyzer/scripts/run.sh" - ;; - benchmark_skill_tokens|benchmark_skill_tokens.py) - printf '%s\n' "$SCRIPT_DIR/benchmark_skill_tokens.py" - ;; - build_inventory|build_inventory.py) - printf '%s\n' "$SCRIPT_DIR/build_inventory.py" - ;; - check_catalog_consistency|check_catalog_consistency.py) - printf '%s\n' "$SCRIPT_DIR/check_catalog_consistency.py" - ;; - audit_copilot_catalog|audit_copilot_catalog.py) - printf '%s\n' "$SCRIPT_DIR/audit_copilot_catalog.py" - ;; - detect_token_risks|detect_token_risks.py) - printf '%s\n' "$SCRIPT_DIR/detect_token_risks.py" - ;; - sync_home_ai_resources|sync_home_ai_resources.py) - printf '%s\n' "$REPO_ROOT/.github/skills/local-agent-sync-install-ai-resources/scripts/run.sh" - ;; - validate_internal_skills|validate_internal_skills.py) - printf '%s\n' "$SCRIPT_DIR/validate_internal_skills.py" - ;; - validate_skill_change_scope|validate_skill_change_scope.py) - printf '%s\n' "$SCRIPT_DIR/validate_skill_change_scope.py" - ;; - *) - return 1 - ;; - esac -} - -ensure_venv() { - if [[ -d "$VENV_DIR" ]]; then - verify_venv_version - return - fi - log_info "Creating local virtual environment." - "$PYTHON_BIN" -m venv "$VENV_DIR" - verify_venv_version -} - -install_dependencies() { - local requirements_hash - local current_hash - requirements_hash="$(hash_file "$REQUIREMENTS_FILE")" - current_hash="" - - if [[ -f "$REQUIREMENTS_HASH_FILE" ]]; then - current_hash="$(<"$REQUIREMENTS_HASH_FILE")" - fi - - if [[ "$requirements_hash" == "$current_hash" ]]; then - return - fi - - log_info "Installing locked Python dependencies." - "$VENV_DIR/bin/pip" install --require-hashes -r "$REQUIREMENTS_FILE" - printf '%s' "$requirements_hash" >"$REQUIREMENTS_HASH_FILE" - log_success "Local Python environment is ready." -} - -main() { - local tool_name="${1:-}" - local script_path="" - - if [[ -z "$tool_name" ]]; then - usage - exit 1 - fi - - if [[ "$tool_name" == "analyze_copilot_debug_log" || "$tool_name" == "analyze_copilot_debug_log.sh" ]]; then - shift - exec bash "$REPO_ROOT/.github/skills/local-copilot-log-analyzer/scripts/run.sh" "$@" - fi - - script_path="$(resolve_script "$tool_name")" || { - log_error "Unknown tool: $tool_name" - usage - exit 1 - } - shift - - if [[ "$script_path" == *.sh ]]; then - exec bash "$script_path" "$@" - fi - - load_required_python_version - select_python_bin - require_command "$PYTHON_BIN" - verify_python_bin_version - ensure_venv - install_dependencies - exec "$VENV_DIR/bin/python" "$script_path" "$@" -} - -SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" -REPO_ROOT="$(cd "$SCRIPT_DIR/../.." && pwd)" -VENV_DIR="$SCRIPT_DIR/.venv" -PYTHON_BIN="${PYTHON_BIN:-}" -PYTHON_VERSION_FILE="$REPO_ROOT/.python-version" -REQUIREMENTS_FILE="$SCRIPT_DIR/requirements.txt" -REQUIREMENTS_HASH_FILE="$VENV_DIR/.requirements.sha256" - -if [[ "${BASH_SOURCE[0]}" == "$0" ]]; then - main "$@" -fi diff --git a/.github/skills/addyosmani-code-review-and-quality/SKILL.md b/.github/skills/addyosmani-code-review-and-quality/SKILL.md index 3e58e23e..6f0c9b49 100644 --- a/.github/skills/addyosmani-code-review-and-quality/SKILL.md +++ b/.github/skills/addyosmani-code-review-and-quality/SKILL.md @@ -348,8 +348,8 @@ For triaging `npm audit` findings and supply-chain risk (typosquatting, compromi ``` ## See Also -- For detailed security review guidance, see `references/security-checklist.md` -- For performance review checks, see `references/performance-checklist.md` +- For detailed security review guidance, see `../../references/security-checklist.md` +- For performance review checks, see `../../references/performance-checklist.md` ## Common Rationalizations diff --git a/.github/skills/addyosmani-idea-refine/SKILL.md b/.github/skills/addyosmani-idea-refine/SKILL.md new file mode 100644 index 00000000..494d5d81 --- /dev/null +++ b/.github/skills/addyosmani-idea-refine/SKILL.md @@ -0,0 +1,193 @@ +--- +name: addyosmani-idea-refine +description: Refines raw ideas into sharp, actionable concepts through structured divergent and convergent thinking. Use when an idea is still vague, when you need to stress-test assumptions before committing to a plan, or when you want to expand options before converging on one. Triggers on "ideate", "refine this idea", or "stress-test my plan". +--- + +# Idea Refine + +Refines raw ideas into sharp, actionable concepts worth building through structured divergent and convergent thinking. + +## How It Works + +1. **Understand & Expand (Divergent):** Restate the idea, ask sharpening questions, and generate variations. +2. **Evaluate & Converge:** Cluster ideas, stress-test them, and surface hidden assumptions. +3. **Sharpen & Ship:** Produce a concrete markdown one-pager moving work forward. + +## Usage + +This skill is primarily an interactive dialogue. Invoke it with an idea, and the agent will guide you through the process. + +```bash +# Optional: Initialize the ideas directory +bash skills/idea-refine/scripts/idea-refine.sh +``` + +**Trigger Phrases:** +- "Help me refine this idea" +- "Ideate on [concept]" +- "Stress-test my plan" + +## Output + +The final output is a markdown one-pager saved to `tmp/ideas/[idea-name].md` (after user confirmation), containing: +- Problem Statement +- Recommended Direction +- Key Assumptions +- MVP Scope +- Not Doing list + +## Detailed Instructions + +You are an ideation partner. Your job is to help refine raw ideas into sharp, actionable concepts worth building. + +### Philosophy + +- Simplicity is the ultimate sophistication. Push toward the simplest version that still solves the real problem. +- Start with the user experience, work backwards to technology. +- Say no to 1,000 things. Focus beats breadth. +- Challenge every assumption. "How it's usually done" is not a reason. +- Show people the future — don't just give them better horses. +- The parts you can't see should be as beautiful as the parts you can. + +### Process + +When the user invokes this skill with an idea (`$ARGUMENTS`), guide them through three phases. Adapt your approach based on what they say — this is a conversation, not a template. + +#### Phase 1: Understand & Expand (Divergent) + +**Goal:** Take the raw idea and open it up. + +1. **Restate the idea** as a crisp "How Might We" problem statement. This forces clarity on what's actually being solved. + +2. **Ask 3-5 sharpening questions** — no more. Focus on: + - Who is this for, specifically? + - What does success look like? + - What are the real constraints (time, tech, resources)? + - What's been tried before? + - Why now? + + Use the `AskUserQuestion` tool to gather this input. Do NOT proceed until you understand who this is for and what success looks like. + +3. **Generate 5-8 idea variations** using these lenses: + - **Inversion:** "What if we did the opposite?" + - **Constraint removal:** "What if budget/time/tech weren't factors?" + - **Audience shift:** "What if this were for [different user]?" + - **Combination:** "What if we merged this with [adjacent idea]?" + - **Simplification:** "What's the version that's 10x simpler?" + - **10x version:** "What would this look like at massive scale?" + - **Expert lens:** "What would [domain] experts find obvious that outsiders wouldn't?" + + Push beyond what the user initially asked for. Create products people don't know they need yet. + +**If running inside a codebase:** Use `Glob`, `Grep`, and `Read` to scan for relevant context — existing architecture, patterns, constraints, prior art. Ground your variations in what actually exists. Reference specific files and patterns when relevant. + +Read `frameworks.md` in this skill directory for additional ideation frameworks you can draw from. Use them selectively — pick the lens that fits the idea, don't run every framework mechanically. + +#### Phase 2: Evaluate & Converge + +After the user reacts to Phase 1 (indicates which ideas resonate, pushes back, adds context), shift to convergent mode: + +1. **Cluster** the ideas that resonated into 2-3 distinct directions. Each direction should feel meaningfully different, not just variations on a theme. + +2. **Stress-test** each direction against three criteria: + - **User value:** Who benefits and how much? Is this a painkiller or a vitamin? + - **Feasibility:** What's the technical and resource cost? What's the hardest part? + - **Differentiation:** What makes this genuinely different? Would someone switch from their current solution? + + Read `refinement-criteria.md` in this skill directory for the full evaluation rubric. + +3. **Surface hidden assumptions.** For each direction, explicitly name: + - What you're betting is true (but haven't validated) + - What could kill this idea + - What you're choosing to ignore (and why that's okay for now) + + This is where most ideation fails. Don't skip it. + +**Be honest, not supportive.** If an idea is weak, say so with kindness. A good ideation partner is not a yes-machine. Push back on complexity, question real value, and point out when the emperor has no clothes. + +#### Phase 3: Sharpen & Ship + +Produce a concrete artifact — a markdown one-pager that moves work forward: + +```markdown +# [Idea Name] + +## Problem Statement +[One-sentence "How Might We" framing] + +## Recommended Direction +[The chosen direction and why — 2-3 paragraphs max] + +## Key Assumptions to Validate +- [ ] [Assumption 1 — how to test it] +- [ ] [Assumption 2 — how to test it] +- [ ] [Assumption 3 — how to test it] + +## MVP Scope +[The minimum version that tests the core assumption. What's in, what's out.] + +## Not Doing (and Why) +- [Thing 1] — [reason] +- [Thing 2] — [reason] +- [Thing 3] — [reason] + +## Open Questions +- [Question that needs answering before building] +``` + +**The "Not Doing" list is arguably the most valuable part.** Focus is about saying no to good ideas. Make the trade-offs explicit. + +Ask the user if they'd like to save this to `tmp/ideas/[idea-name].md` (or a location of their choosing). Only save if they confirm. + +### Anti-patterns to Avoid + +- **Don't generate 20+ ideas.** Quality over quantity. 5-8 well-considered variations beat 20 shallow ones. +- **Don't be a yes-machine.** Push back on weak ideas with specificity and kindness. +- **Don't skip "who is this for."** Every good idea starts with a person and their problem. +- **Don't produce a plan without surfacing assumptions.** Untested assumptions are the #1 killer of good ideas. +- **Don't over-engineer the process.** Three phases, each doing one thing well. Resist adding steps. +- **Don't just list ideas — tell a story.** Each variation should have a reason it exists, not just be a bullet point. +- **Don't ignore the codebase.** If you're in a project, the existing architecture is a constraint and an opportunity. Use it. + +### Tone + +Direct, thoughtful, slightly provocative. You're a sharp thinking partner, not a facilitator reading from a script. Channel the energy of "that's interesting, but what if..." -- always pushing one step further without being exhausting. + +Read `examples.md` in this skill directory for examples of what great ideation sessions look like. + +## Red Flags + +- Generating 20+ shallow variations instead of 5-8 considered ones +- Skipping the "who is this for" question +- No assumptions surfaced before committing to a direction +- Yes-machining weak ideas instead of pushing back with specificity +- Producing a plan without a "Not Doing" list +- Ignoring existing codebase constraints when ideating inside a project +- Jumping straight to Phase 3 output without running Phases 1 and 2 + +## Verification + +After completing an ideation session: + +- [ ] A clear "How Might We" problem statement exists +- [ ] The target user and success criteria are defined +- [ ] Multiple directions were explored, not just the first idea +- [ ] Hidden assumptions are explicitly listed with validation strategies +- [ ] A "Not Doing" list makes trade-offs explicit +- [ ] The output is a concrete artifact (markdown one-pager), not just conversation +- [ ] The user confirmed the final direction before any implementation work + + +## Local idea-workspace override + +Apply this local override instead of any earlier conflicting workspace or +output-path instructions in this skill and its bundled resources. + +- Keep every generated artifact under `tmp/ideas/` in the current project root, + including drafts, supporting notes, and final one-pagers. +- Save the final one-pager to `tmp/ideas/[idea-name].md` only after user + confirmation. Alternative filenames and subdirectories must stay inside + `tmp/ideas/`. +- Run the initialization script from the project root. Do not create idea + artifacts outside `tmp/ideas/`. + diff --git a/.github/skills/addyosmani-idea-refine/examples.md b/.github/skills/addyosmani-idea-refine/examples.md new file mode 100644 index 00000000..a797d22c --- /dev/null +++ b/.github/skills/addyosmani-idea-refine/examples.md @@ -0,0 +1,238 @@ +# Ideation Session Examples + +These examples demonstrate what good ideation sessions look like across different kinds of ideas. Study the rhythm, tone, and structure — not just the content. The skill should feel equally at home with a vague startup concept, a feature for an existing product, or a process improvement. + +--- + +## Example 1: Vague Early-Stage Concept (Full 3-Phase Session) + +**User invokes:** `/ideate I want to help small local restaurants compete with the big delivery platforms` + +### Phase 1: Understand & Expand + +**Skill response:** + +> **How might we** give independent restaurants the reach and convenience customers expect from delivery platforms, without forcing them into a model that erodes their margins and brand? + +That's a big space. Let me sharpen it: + +1. **Which part of "compete"?** Customer acquisition? Delivery logistics? Online ordering UX? Loyalty/retention? Margin pressure from platform fees? +2. **Who are you building for?** The restaurant owner? The diner? Both? +3. **What's the constraint?** Are you thinking software? A cooperative model? Something physical? + +**User responds:** "Restaurant owners. The core pain is that they're paying 25-30% commission to delivery apps but feel they can't opt out because that's where the customers are. I'm thinking software but I'm open." + +**Skill continues:** + +So the problem isn't delivery logistics — it's **customer dependency**. The platforms own the customer relationship, and restaurants are renting access to their own customers at 30% of revenue. That's the leverage to break. + +Here are 6 variations: + +1. **Direct Channel Toolkit:** Give restaurants their own branded ordering — web, SMS, maybe a lightweight app — so they can move existing customers off-platform. The core bet: customers who *already know* a restaurant will order direct if it's easy enough. You don't need to compete with DoorDash for discovery, just for reordering. + +2. **Inversion — Charge the Customer, Not the Restaurant:** What if the delivery fee model flipped? Customers pay a small premium for direct ordering, restaurants pay zero commission. Position it as "support your local spot." The bet: enough customers care about supporting local businesses to pay $1-2 more per order. + +3. **Constraint Removal — What If Delivery Weren't Required?** The platforms bundle discovery + ordering + delivery, but most of their cost (and the restaurant's commission) is delivery. What if you only solved ordering and pickup? Pickup-first model. Restaurants keep 100% of revenue, customers save the delivery fee. Works best in dense neighborhoods. + +4. **Combination — Restaurant Collective:** A co-op model where 20-50 local restaurants in an area share a platform, split costs, and cross-promote each other. Individual restaurants can't afford marketing — but a neighborhood collective can. Think "local food hall" but digital. + +5. **Audience Shift — For Regular Customers:** Don't try to acquire new customers. Build exclusively for a restaurant's existing regulars — the people who come in weekly. A loyalty/reorder tool that's dead simple: "Your usual?" one-tap reorder. The 80/20 bet: 20% of customers drive 80% of revenue, and those people don't need DoorDash to discover the restaurant. + +6. **Expert Lens — What Restaurant Owners Actually Do:** Most small restaurant owners aren't looking at dashboards. They're in the kitchen. Any tool that requires them to "manage" something will fail. The winning approach might be something that runs itself — auto-texts regulars when they haven't ordered in 2 weeks, auto-posts the daily special, auto-handles reorders. Zero restaurant-side effort. + +Variation 5 and 6 together are interesting. Instead of trying to out-platform the platforms, you'd build something hyper-narrow: a reordering and retention tool for a restaurant's best existing customers. No discovery, no delivery, no marketplace. Just "keep your regulars ordering direct." + +Which directions interest you? + +--- + +### Phase 2: Evaluate & Converge + +**User responds:** "I love the regulars angle (5 and 6). The direct channel toolkit (1) feels necessary but boring. The co-op idea (4) is interesting but complex. Let me focus on 1 + 5 + 6 combined." + +**Skill response:** + +Good instincts. That combination has a clear thesis: **own the relationship with your best customers, make reordering effortless, and require zero effort from the restaurant owner.** Let me stress-test it. + +**Direction A: The "Regulars Engine" (Variations 5 + 6 fused)** +A tool focused entirely on a restaurant's existing repeat customers. Auto-reorder, auto-engagement, zero management. + +- **Core bet:** A restaurant's top 20% of customers will shift to direct ordering if the experience is frictionless and personal. The restaurant doesn't need *new* customers — it needs to stop paying 30% on customers it already has. +- **User value:** High and immediate. Every direct reorder saves the restaurant 25-30% commission. If a regular orders $30/week, that's ~$400/year in saved fees per customer. 50 regulars = $20K/year. That's real money for a small restaurant. +- **Feasibility:** Medium. The ordering/payment tech is commodity. The hard parts: (a) how do you identify and onboard "regulars" without the restaurant doing manual work? (b) How do you handle the cold start — the first reorder needs an order history that lives on the delivery platform, not with you. +- **Differentiation:** High for the niche. Platforms optimize for marketplace dynamics (new customer discovery, driver logistics). Nobody is optimizing specifically for "keep your regulars ordering direct." It's too small for DoorDash to care about, which is exactly why it's a good wedge. +- **Hidden assumptions:** + - Regulars will install/use yet another ordering method (assumption most likely to be wrong) + - Restaurant owners will share customer contact info or let you access it + - Direct ordering won't cannibalize dine-in revenue +- **What could kill it:** If customers simply won't switch channels. Habit is powerful — even if DoorDash costs more, it's already on their phone. + +**Direction B: Direct Channel + Retention Hybrid (Variation 1 + 5)** +Branded ordering for the restaurant *plus* the retention automation. More complete but more complex. + +- **Core bet:** Restaurants need both the channel (a way to accept orders) and the engine (a way to keep customers coming back through that channel). +- **User value:** Higher ceiling, but only if both parts work together. The ordering channel alone is a commodity (Square, Toast already do this). The retention layer alone has no channel to retain customers *on*. +- **Feasibility:** Lower. You're building two products and hoping they're better together. +- **Differentiation:** Medium. The combination is interesting, but each piece individually has competitors. +- **Hidden assumption:** That bundling these is better than partnering (e.g., retention layer on top of Square Online). +- **What could kill it:** Scope. Building two things well is exponentially harder than building one thing well. + +**The honest read:** Direction A is the sharper bet. Direction B is safer but less differentiated. The biggest risk for A is the channel-switching problem — getting customers to order through a new method. If you can crack that (and variation 6's "zero effort" approach might help — an SMS-based reorder doesn't require a new app), the rest follows. + +One thing I want to push back on: you said variation 1 feels "necessary but boring." That instinct to include the "necessary" thing is how products lose focus. If the regulars engine works, you can add a branded ordering channel later. But if you build the channel first, you're just another Square Online competitor. + +--- + +### Phase 3: Sharpen & Ship + +**Skill produces:** + +# ReOrder: Keep Your Regulars Ordering Direct + +## Problem Statement +How might we help independent restaurants retain their best existing customers through direct ordering, eliminating the 25-30% platform commission on the revenue that matters most? + +## Recommended Direction +Build a **regulars-only retention engine** — not a marketplace, not a full ordering platform. Hyper-focused on one job: make it effortless for a restaurant's repeat customers to reorder directly. + +The key insight is that restaurants don't need help *finding* their best customers — they know who walks in every Tuesday. They need help moving those relationships off-platform. And the tool needs to run itself, because the owner is in the kitchen, not at a dashboard. + +SMS-first (not app-first) is likely the right channel. A text saying "Hey, want your usual Thursday order from Marco's?" with a one-tap confirmation is lower friction than any app install. + +## Key Assumptions to Validate +- [ ] Repeat customers will reorder via SMS/direct link instead of their usual delivery app — test with 5 restaurants, 20 regulars each, measure conversion over 4 weeks +- [ ] Restaurant owners can identify their top 20-30 regulars and share contact info — test by asking 10 restaurant owners if they'd do this +- [ ] The commission savings ($8-10 per order) is motivating enough for owners to invest initial setup effort — interview 10 owners about platform fee pain + +## MVP Scope +- SMS-based reordering for a restaurant's self-identified regular customers +- Restaurant owner adds regulars manually (name + phone + usual order) — 15-minute setup +- Customer receives a text with their usual order, confirms with a reply, pays via link +- Restaurant receives the order via text/simple dashboard +- No delivery — pickup only in v1 +- No discovery, no marketplace, no app + +## Not Doing (and Why) +- **Delivery logistics** — delivery is the expensive, complex part and not the core problem. Pickup-first validates demand without operational burden. +- **Customer acquisition/discovery** — that's the platform's game. Competing on discovery means competing with DoorDash's budget. We compete on retention instead. +- **Branded restaurant apps/websites** — commodity. Square and Toast already do this. Another branded channel doesn't solve the behavioral problem. +- **Menu management, POS integration** — scope creep. "Your usual" doesn't need a dynamic menu. Start with static orders. +- **Analytics dashboards** — the owner is in the kitchen. Build reporting that's passive (weekly text summary) not active (log in and look at charts). + +## Open Questions +- What's the legal/compliance landscape for SMS marketing to restaurant customers? +- How do you handle menu/price changes for "usual" orders? +- Is there a way to bootstrap the regular customer list from existing POS data instead of manual entry? +- What's the pricing model? Per-order fee? Monthly subscription? + +--- + +## Example 2: Feature Idea Within an Existing Product (Codebase-Aware) + +**User invokes:** `/ideate adding real-time collaboration to our document editor` + +### Phase 1: Understand & Expand + +**Skill response:** + +> **How might we** let multiple people work on the same document simultaneously in a way that feels natural, not chaotic? + +I scanned your codebase and see you're using a block-based editor architecture with a REST API for persistence. No WebSocket infrastructure currently. That's important context. A few questions: + +1. **What's the collaboration pattern?** Are people editing the same paragraph at the same time (Google Docs style)? Or working in different sections (Notion style)? Or reviewing/commenting while one person edits? +2. **How many simultaneous users?** 2-3 (pair editing)? 10+ (team workspace)? +3. **What's driving this now?** Customer requests? Competitive pressure? Internal need? + +**User responds:** "2-5 people, usually working in different sections. Customers keep asking. We're losing deals to Notion because they assume any modern editor has this." + +**Skill continues:** + +The competitive pressure framing matters — this might be a checkbox feature (must-have to stay in the consideration set) rather than a differentiator. That changes how much you should invest. + +I see your document model uses independent blocks with a flat ordering system (`src/models/document.ts:45`). That's actually favorable — block-level granularity means conflicts are rarer than character-level editing. + +Here are 5 variations: + +1. **Block-Level Locking (Simplest Path):** When someone starts editing a block, it's locked for others. They see a presence indicator ("Alex is editing...") but can edit any other block freely. No conflict resolution needed. Your existing block model supports this with minimal changes — add a `locked_by` field and a WebSocket layer for presence. + +2. **Inversion — Async-First Collaboration:** Instead of real-time co-editing, lean into async. Inline suggestions, comments, and proposed edits that the document owner merges. Like suggestion mode, but as the *primary* collaboration model. Bet: for 2-5 people in different sections, async might actually be *better* than real-time. + +3. **Simplification — Presence Without Co-Editing:** Show who's viewing the document and what section they're looking at (cursors, avatars in margins), but don't enable simultaneous editing. People coordinate verbally. Sounds primitive, but might be 80% of the value at 10% of the engineering cost. + +4. **Combination — Collaboration + Version Branches:** Let people create "branches" of a document, edit independently, then merge. Each person gets their own space, no conflicts during editing, and you build a merge UI. Plays to your existing versioning system. + +5. **10x Version — AI-Mediated Collaboration:** Multiple people describe what they want in a section, and AI synthesizes their inputs into a draft they all iterate on. Nobody directly edits the same text — they all influence it. Radical, possibly too early, but worth naming. + +Given your architecture and the competitive context, I'd push you toward variation 1 or 3. You need to close deals, which means you need this shipped, which means you need the simplest version that checks the box. Variation 1 gives you real co-editing. Variation 3 is even cheaper but might not satisfy the "we need real-time" checkbox. + +--- + +*Phase 2 would cluster into "minimum viable co-editing (1 or 3)" vs. "differentiated collaboration model (2 or 4)" — evaluating whether to check the competitive box quickly or invest in something that becomes a selling point.* + +*Phase 3 would produce a one-pager recommending block-level locking as the MVP, with a "Not Doing" list that includes character-level CRDT (too complex), offline sync (not needed for 2-5 users), and AI features (premature).* + +--- + +## Example 3: Process/Workflow Idea (Non-Product) + +**User invokes:** `/ideate a better way to run our weekly team retrospectives — they've gotten stale and people dread them` + +### Phase 1: Understand & Expand + +**Skill response:** + +> **How might we** make retrospectives genuinely useful — something the team looks forward to — instead of a recurring meeting people endure? + +Before expanding: + +1. **What's stale about them?** Same format every time? Same people talking? Action items that never get done? Topics feel repetitive? +2. **Team size and dynamic?** 5 people who all speak up, or 15 where 3 dominate? +3. **What does a *good* retro look like to you?** Have you had one that worked? What made it different? + +**User responds:** "8 people, maybe 3-4 actually talk. Same 'what went well / what didn't / action items' format for a year. The action items pile up and nothing changes. People think it's theater." + +**Skill continues:** + +The action item problem is the real signal. If nothing changes after the retro, people rationally disengage. The meeting isn't stale — it's **broken at the output layer**. The format might be fine if the follow-through worked. + +Here are 6 variations: + +1. **Fix the Output, Not the Meeting:** Keep the retro format but completely change what happens after. One action item max per retro. Assigned an owner and a deadline during the meeting. Next retro starts with a 5-minute demo of the completed action. If it's not done, that *becomes* the retro topic. The constraint of one item forces prioritization. + +2. **Inversion — Kill the Meeting:** Replace the synchronous retro with an async one. Everyone submits reflections in writing during the week. A rotating facilitator synthesizes themes and proposes one change. The "meeting" is a 15-minute standup to agree or disagree. Bet: writing produces more honest feedback than speaking in a group. + +3. **Audience Shift — Retro for the Quiet Ones:** Design the format for the 4-5 people who don't talk. Anonymous submission of topics. Dot-voting to prioritize. Small-group breakouts (pairs) before full-group discussion. The loud voices still participate, but the structure prevents them from dominating. + +4. **Simplification — The One Question Retro:** Each retro has exactly one question, and it changes every week. "What's one thing that slowed you down?" "What did you learn that the team should know?" "If you could change one thing about how we work, what?" The constraint forces depth over breadth. + +5. **Combination — Retro + Experimentation:** Treat each retro output as a hypothesis. "We believe that [change] will improve [outcome]." Run it as a 2-week experiment. Next retro: did it work? Keep, modify, or kill. Turns the retro into a continuous improvement engine with built-in accountability. + +6. **Expert Lens — What Facilitators Know:** Experienced facilitators say the #1 retro killer is lack of safety, not format. People won't say what's really wrong if they fear consequences. The fix might not be structural — it might be starting with an anonymous "team health check" score (1-5) each week. When safety is high, retros naturally improve. + +The interesting tension: variations 1 and 5 fix the *output* problem (nothing changes). Variations 2, 3, and 4 fix the *input* problem (same voices, same topics). Variation 6 says both are symptoms of a deeper issue. Where do you think the real bottleneck is? + +--- + +*Phase 2 would evaluate in terms of: effort to try (most are free — just change how you run the next meeting), risk (variation 2 is the biggest departure), and whether the team's real problem is output (action items die) or input (not enough honesty).* + +*Phase 3 would produce a one-pager recommending starting with variation 1 (one action item, demo next week) as a zero-cost experiment, combined with variation 3's anonymous submission. "Not Doing" list: new tools, elaborate facilitation techniques, or anything requiring budget. The first fix should take 0 minutes of prep and $0.* + +--- + +## What to Notice in These Examples + +1. **The restatement changes the frame.** "Help restaurants compete" becomes "retain existing customers." "Add real-time collaboration" becomes "let people work simultaneously without chaos." "Fix stale retros" becomes "fix the output layer." + +2. **Questions diagnose before prescribing.** Each question determines which *type* of problem this actually is. The retro example reveals the problem is action item follow-through, not meeting format — and that changes every variation. + +3. **Variations have reasons.** Each one explains *why* it exists (what lens generated it), not just *what* it is. The label (Inversion, Simplification, etc.) teaches the user to think this way themselves. + +4. **The skill has opinions.** "I'd push you toward 1 or 3." "Variation 6 is worth sitting with." It tells you what it thinks matters and why — not just neutral options. + +5. **Phase 2 is honest.** Ideas get called out for low differentiation or high complexity. The skill pushes back: "That instinct to include the 'necessary' thing is how products lose focus." + +6. **The output is actionable.** The one-pager ends with things you can *do* (validate assumptions, build the MVP, try the experiment), not things to *think about*. + +7. **The "Not Doing" list does real work.** It's specific and reasoned. Each item is something you might *want* to do but shouldn't yet. + +8. **The skill adapts to context.** A codebase-aware example references actual architecture. A process idea generates zero-cost experiments instead of products. The framework stays the same but the output matches the domain. diff --git a/.github/skills/addyosmani-idea-refine/frameworks.md b/.github/skills/addyosmani-idea-refine/frameworks.md new file mode 100644 index 00000000..0e7fc8fe --- /dev/null +++ b/.github/skills/addyosmani-idea-refine/frameworks.md @@ -0,0 +1,99 @@ +# Ideation Frameworks Reference + +Use these frameworks selectively. Pick the lens that fits the idea — don't mechanically run every framework. The goal is to unlock thinking, not to follow a checklist. + +## SCAMPER + +A structured way to transform an existing idea by applying seven different operations: + +- **Substitute:** What component, material, or process could you swap out? What if you replaced the core technology? The target audience? The business model? +- **Combine:** What if you merged this with another product, service, or idea? What two things that don't usually go together would create something new? +- **Adapt:** What else is like this? What ideas from other industries, domains, or time periods could you borrow? What parallel exists in nature? +- **Modify (Magnify/Minimize):** What if you made it 10x bigger? 10x smaller? What if you exaggerated one feature? What if you stripped it to the absolute minimum? +- **Put to other uses:** Who else could use this? What other problems could it solve? What happens if you use it in a completely different context? +- **Eliminate:** What happens if you remove a feature entirely? What's the version with zero configuration? What would it look like with half the steps? +- **Reverse/Rearrange:** What if you did the steps in the opposite order? What if the user did the work instead of the system (or vice versa)? What if you reversed the value chain? + +**Best for:** Improving or reimagining existing products/features. Less useful for greenfield ideas. + +## How Might We (HMW) + +Reframe problems as opportunities using the "How Might We..." format: + +- Start with an observation or pain point +- Reframe it as "How might we [desired outcome] for [specific user] without [key constraint]?" +- Generate multiple HMW framings of the same problem — different framings unlock different solutions + +**Good HMW qualities:** +- Narrow enough to be actionable ("...help new users find relevant content in their first 5 minutes") +- Broad enough to allow creative solutions (not "...add a recommendation sidebar") +- Contains a tension or constraint that forces creativity + +**Bad HMW qualities:** +- Too broad: "How might we make users happy?" +- Too narrow: "How might we add a button to the settings page?" +- Solution-embedded: "How might we build a chatbot for support?" + +**Best for:** Reframing stuck thinking. When someone is anchored on a solution, pull them back to the problem. + +## First Principles Thinking + +Break the idea down to its fundamental truths, then rebuild from there: + +1. **What do we know is true?** (not assumed, not conventional — actually true) +2. **What are we assuming?** List every assumption, even the ones that feel obvious +3. **Which assumptions can we challenge?** For each, ask: "Is this actually a law of physics, or just how it's been done?" +4. **Rebuild from the truths.** If you only had the fundamental truths, what would you build? + +**Best for:** Breaking out of incremental thinking. When every idea feels like a small improvement on the status quo. + +## Jobs to Be Done (JTBD) + +Focus on what the user is trying to accomplish, not what they say they want: + +- **Functional job:** What task are they trying to complete? +- **Emotional job:** How do they want to feel? +- **Social job:** How do they want to be perceived? + +Format: "When I [situation], I want to [motivation], so I can [expected outcome]." + +**Key insight:** People don't buy products — they hire them to do a job. The competing product isn't always in the same category. (Netflix competes with sleep, not just other streaming services.) + +**Best for:** Understanding the real problem. When you're not sure if you're solving the right thing. + +## Constraint-Based Ideation + +Deliberately impose constraints to force creative solutions: + +- **Time constraint:** "What if you only had 1 day to build this?" +- **Feature constraint:** "What if it could only have one feature?" +- **Tech constraint:** "What if you couldn't use [the obvious technology]?" +- **Cost constraint:** "What if it had to be free forever?" +- **Audience constraint:** "What if your user had never used a computer before?" +- **Scale constraint:** "What if it needed to work for 1 billion users? What about just 10?" + +**Best for:** Cutting through complexity. When the idea is growing too large or too vague. + +## Pre-mortem + +Imagine the idea has already failed. Work backwards: + +1. It's 12 months from now. The project shipped and flopped. What went wrong? +2. List every plausible reason for failure — technical, market, team, timing +3. For each failure mode: Is this preventable? Is this a signal the idea needs to change? +4. Which failure modes are you willing to accept? Which ones would kill the project? + +**Best for:** Phase 2 evaluation. Stress-testing ideas that feel good but haven't been pressure-tested. + +## Analogous Inspiration + +Look at how other domains solved similar problems: + +- What industry has already solved a version of this problem? +- What would this look like if [specific company/product] built it? +- What natural system works this way? +- What historical precedent exists? + +The key is finding *structural* similarities, not surface-level ones. "Uber for X" is surface-level. "A two-sided marketplace that solves a trust problem between strangers" is structural. + +**Best for:** Phase 1 expansion. Generating variations that feel genuinely different from the obvious approach. diff --git a/.github/skills/addyosmani-idea-refine/refinement-criteria.md b/.github/skills/addyosmani-idea-refine/refinement-criteria.md new file mode 100644 index 00000000..53e79c72 --- /dev/null +++ b/.github/skills/addyosmani-idea-refine/refinement-criteria.md @@ -0,0 +1,113 @@ +# Refinement & Evaluation Criteria + +Use this rubric during Phase 2 (Evaluate & Converge) to stress-test idea directions. Not every criterion applies to every idea — use judgment about which dimensions matter most for the specific context. + +## Core Evaluation Dimensions + +### 1. User Value + +The most important dimension. If the value isn't clear, nothing else matters. + +**Painkiller vs. Vitamin:** +- **Painkiller:** Solves an acute, frequent problem. Users will actively seek this out. They'll switch from their current solution. Signs: people describe the problem with emotion, they've built workarounds, they'll pay for a solution. +- **Vitamin:** Nice to have. Makes something marginally better. Users won't go out of their way. Signs: people nod politely, say "that's cool," then don't change behavior. + +**Questions to ask:** +- Can you name 3 specific people who have this problem right now? +- What are they doing today instead? (The real competitor is always the current workaround.) +- Would they switch from their current approach? What would make them switch? +- How often do they encounter this problem? (Daily problems > monthly problems) +- Is this a "pull" problem (users are asking for this) or a "push" problem (you think they should want this)? + +**Red flags:** +- "Everyone could use this" — if you can't name a specific user, the value isn't clear +- "It's like X but better" — marginal improvements rarely drive adoption +- The problem is real but rare — high intensity but low frequency rarely justifies a product + +### 2. Feasibility + +Can you actually build this? Not just technically, but practically. + +**Technical feasibility:** +- Does the core technology exist and work reliably? +- What's the hardest technical problem? Is it a known-hard problem or a novel one? +- Are there dependencies on third parties, APIs, or data sources you don't control? +- What's the minimum technical stack needed? (If the answer is "a lot," that's a signal.) + +**Resource feasibility:** +- What's the minimum team/effort to build an MVP? +- Does it require specialized expertise you don't have? +- Are there regulatory, legal, or compliance requirements? + +**Time-to-value:** +- How quickly can you get something in front of users? +- Is there a version that delivers value in days/weeks, not months? +- What's the critical path? What has to happen first? + +**Red flags:** +- "We just need to solve [very hard research problem] first" +- Multiple dependencies that all need to work simultaneously +- MVP still requires months of work — likely not minimal enough + +### 3. Differentiation + +What makes this genuinely different? Not better — *different*. + +**Questions to ask:** +- If a user described this to a friend, what would they say? Is that description compelling? +- What's the one thing this does that nothing else does? (If you can't name one, that's a problem.) +- Is this differentiation durable? Can a competitor copy it in a week? +- Is the difference something users actually care about, or just something builders find interesting? + +**Types of differentiation (strongest to weakest):** +1. **New capability:** Does something that was previously impossible +2. **10x improvement:** So much better on a key dimension that it changes behavior +3. **New audience:** Brings an existing capability to people who were excluded +4. **New context:** Works in a situation where existing solutions fail +5. **Better UX:** Same capability, dramatically simpler experience +6. **Cheaper:** Same thing, lower cost (weakest — easily competed away) + +**Red flags:** +- Differentiation is entirely about technology, not user experience +- "We're faster/cheaper/prettier" without a structural reason why +- The feature that differentiates is not the feature users care most about + +## Assumption Audit + +For every idea direction, explicitly list assumptions in three categories: + +### Must Be True (Dealbreakers) +Assumptions that, if wrong, kill the idea entirely. These need validation before building. + +Example: "Users will share their data with us" — if they won't, the entire product doesn't work. + +### Should Be True (Important) +Assumptions that significantly impact success but don't kill the idea. You can adjust the approach if these are wrong. + +Example: "Users prefer self-serve over talking to a person" — if wrong, you need a different go-to-market, but the core product can still work. + +### Might Be True (Nice to Have) +Assumptions about secondary features or optimizations. Don't validate these until the core is proven. + +Example: "Users will want to share their results with teammates" — a growth feature, not a core value proposition. + +## Decision Framework + +When choosing between directions, rank on this matrix: + +| | High Feasibility | Low Feasibility | +|--------------------|-------------------|-----------------| +| **High Value** | Do this first | Worth the risk | +| **Low Value** | Only if trivial | Don't do this | + +Then use differentiation as the tiebreaker between options in the same quadrant. + +## MVP Scoping Principles + +When defining MVP scope for the chosen direction: + +1. **One job, done well.** The MVP should nail exactly one user job. Not three jobs done partially. +2. **The riskiest assumption first.** The MVP's primary purpose is to test the assumption most likely to be wrong. +3. **Time-box, not feature-list.** "What can we build and test in [timeframe]?" is better than "What features do we need?" +4. **The 'Not Doing' list is mandatory.** Explicitly name what you're cutting and why. This prevents scope creep and forces honest prioritization. +5. **If it's not embarrassing, you waited too long.** The first version should feel incomplete to the builder. If it doesn't, you over-built. diff --git a/.github/skills/addyosmani-idea-refine/scripts/idea-refine.sh b/.github/skills/addyosmani-idea-refine/scripts/idea-refine.sh new file mode 100755 index 00000000..b5c7b58c --- /dev/null +++ b/.github/skills/addyosmani-idea-refine/scripts/idea-refine.sh @@ -0,0 +1,15 @@ +#!/bin/bash +set -e + +# This script helps initialize the ideas directory for the idea-refine skill. + +IDEAS_DIR="tmp/ideas" + +if [ ! -d "$IDEAS_DIR" ]; then + mkdir -p "$IDEAS_DIR" + echo "Created directory: $IDEAS_DIR" >&2 +else + echo "Directory already exists: $IDEAS_DIR" >&2 +fi + +echo "{\"status\": \"ready\", \"directory\": \"$IDEAS_DIR\"}" diff --git a/.github/skills/addyosmani-planning-and-task-breakdown/SKILL.md b/.github/skills/addyosmani-planning-and-task-breakdown/SKILL.md new file mode 100644 index 00000000..c93a33ac --- /dev/null +++ b/.github/skills/addyosmani-planning-and-task-breakdown/SKILL.md @@ -0,0 +1,285 @@ +--- +name: addyosmani-planning-and-task-breakdown +description: Breaks work into ordered tasks. Use when you have a spec or clear requirements and need to break work into implementable tasks. Use when a task feels too large to start, when you need to estimate scope, or when parallel work is possible. +--- + +# Planning and Task Breakdown + +## Overview + +Decompose work into small, verifiable tasks with explicit acceptance criteria. Good task breakdown is the difference between an agent that completes work reliably and one that produces a tangled mess. Every task should be small enough to implement, test, and verify in a single focused session. + +## When to Use + +- You have a spec and need to break it into implementable units +- A task feels too large or vague to start +- Work needs to be parallelized across multiple agents or sessions +- You need to communicate scope to a human +- The implementation order isn't obvious + +**When NOT to use:** Single-file changes with obvious scope, or when the spec already contains well-defined tasks. + +## The Planning Process + +### Step 1: Enter Plan Mode + +Before writing any code, operate in read-only mode: + +- Read the spec and relevant codebase sections +- Identify existing patterns and conventions +- Map dependencies between components +- Note risks and unknowns + +**Do NOT write code during planning.** The output is a plan document saved to `tmp/.plans/YYYY-MM-DD-HHMM-.md` and a task list recorded in the task list target (see Output Files; default `tmp/.plans/YYYY-MM-DD-HHMM-.md`), not implementation. + +### Step 2: Identify the Dependency Graph + +Map what depends on what: + +``` +Database schema + │ + ├── API models/types + │ │ + │ ├── API endpoints + │ │ │ + │ │ └── Frontend API client + │ │ │ + │ │ └── UI components + │ │ + │ └── Validation logic + │ + └── Seed data / migrations +``` + +Implementation order follows the dependency graph bottom-up: build foundations first. + +### Step 3: Slice Vertically + +Instead of building all the database, then all the API, then all the UI — build one complete feature path at a time: + +**Bad (horizontal slicing):** +``` +Task 1: Build entire database schema +Task 2: Build all API endpoints +Task 3: Build all UI components +Task 4: Connect everything +``` + +**Good (vertical slicing):** +``` +Task 1: User can create an account (schema + API + UI for registration) +Task 2: User can log in (auth schema + API + UI for login) +Task 3: User can create a task (task schema + API + UI for creation) +Task 4: User can view task list (query + API + UI for list view) +``` + +Each vertical slice delivers working, testable functionality. + +### Step 4: Write Tasks + +Each task follows this structure, whether it lands in the markdown task list or as an item in an external tracker (see Output Files): + +```markdown +## Task [N]: [Short descriptive title] + +**Description:** One paragraph explaining what this task accomplishes. + +**Acceptance criteria:** +- [ ] [Specific, testable condition] +- [ ] [Specific, testable condition] + +**Verification:** +- [ ] Tests pass: [the repository's focused-test command] +- [ ] Build succeeds: [the repository's build command] +- [ ] Manual check: [description of what to verify] + +**Dependencies:** [Task numbers this depends on, or "None"] + +**Files likely touched:** +- `src/path/to/file.ts` +- `tests/path/to/test.ts` + +**Estimated scope:** [Small: 1-2 files | Medium: 3-5 files | Large: 5+ files] +``` + +### Step 5: Order and Checkpoint + +Arrange tasks so that: + +1. Dependencies are satisfied (build foundation first) +2. Each task leaves the system in a working state +3. Verification checkpoints occur after every 2-3 tasks +4. High-risk tasks are early (fail fast) + +Add explicit checkpoints to the task list target: + +```markdown +## Checkpoint: After Tasks 1-3 +- [ ] All tests pass +- [ ] Application builds without errors +- [ ] Core user flow works end-to-end +- [ ] Review with human before proceeding +``` + +## Task Sizing Guidelines + +| Size | Files | Scope | Example | +|------|-------|-------|---------| +| **XS** | 1 | Single function or config change | Add a validation rule | +| **S** | 1-2 | One component or endpoint | Add a new API endpoint | +| **M** | 3-5 | One feature slice | User registration flow | +| **L** | 5-8 | Multi-component feature | Search with filtering and pagination | +| **XL** | 8+ | **Too large — break it down further** | — | + +If a task is L or larger, it should be broken into smaller tasks. An agent performs best on S and M tasks. + +**When to break a task down further:** +- It would take more than one focused session (roughly 2+ hours of agent work) +- You cannot describe the acceptance criteria in 3 or fewer bullet points +- It touches two or more independent subsystems (e.g., auth and billing) +- You find yourself writing "and" in the task title (a sign it is two tasks) + +## Output Files + +- **Plan document:** Save the implementation plan to `tmp/.plans/YYYY-MM-DD-HHMM-.md`. This is always a markdown file — design decisions, risks, and open questions don't map cleanly onto individual tracker issues. +- **Task list:** Record each task in the **task list target** (defined below). + +Create the `tasks/` directory if it does not exist. + +**Never overwrite an incomplete plan.** Before writing `tmp/.plans/YYYY-MM-DD-HHMM-.md` or `tmp/.plans/YYYY-MM-DD-HHMM-.md`, check whether they already exist and still contain unchecked tasks: + +- Same work being replanned (the user asked to revise or extend this plan) → update the existing files in place. +- Different work → **stop and ask.** The unchecked tasks may be mid-build in another session. Do not delete, overwrite, or rename the existing files on your own; present the conflict and let the user decide (finish the old plan first, explicitly discard it, or tell you where the new plan should go). + +The same rule applies to an external task list target: never bulk-close or delete another plan's open tracker items to make room for new ones. + +### Task List Target + +The task list target is where tasks and checkpoints are recorded. It is defined once, here; every other reference in this skill defers to it. + +- **Default: a checklist-style markdown file at `tmp/.plans/YYYY-MM-DD-HHMM-.md`.** This is the convention the `/build` command and other downstream tooling expect. Use it unless the project says otherwise. +- **External tracker:** if the project's agent rules (`CLAUDE.md`, `AGENTS.md`, etc.) or the user designate an issue tracker (e.g. GitHub Issues, Jira, Linear, `bd`/beads), create one tracker item per task instead of writing `tmp/.plans/YYYY-MM-DD-HHMM-.md`. Map the Step 4 structure onto the tracker's fields: acceptance criteria and verification steps in the item body, dependencies via the tracker's linking mechanism (`bd dep add`, "blocked by", etc.). Record Step 5 checkpoints as tracker items too, or as a checklist in the plan document if the tracker has no natural equivalent. + +When using an external tracker, note it in `tmp/.plans/YYYY-MM-DD-HHMM-.md` (e.g. "Tasks tracked in Linear project FOO") so downstream steps and future sessions know where to look, and keep the plan document's Task List section as an ordered index of tracker item IDs or links rather than a duplicate checklist. + +## Plan Document Template + +```markdown +# Implementation Plan: [Feature/Project Name] + +## Overview +[One paragraph summary of what we're building] + +## Architecture Decisions +- [Key decision 1 and rationale] +- [Key decision 2 and rationale] + +## Task List + +### Phase 1: Foundation +- [ ] Task 1: ... +- [ ] Task 2: ... + +### Checkpoint: Foundation +- [ ] Tests pass, builds clean + +### Phase 2: Core Features +- [ ] Task 3: ... +- [ ] Task 4: ... + +### Checkpoint: Core Features +- [ ] End-to-end flow works + +### Phase 3: Polish +- [ ] Task 5: ... +- [ ] Task 6: ... + +### Checkpoint: Complete +- [ ] All acceptance criteria met +- [ ] Ready for review + +## Risks and Mitigations +| Risk | Impact | Mitigation | +|------|--------|------------| +| [Risk] | [High/Med/Low] | [Strategy] | + +## Open Questions +- [Question needing human input] +``` + +When tasks live in an external tracker, keep the Task List section above as an ordered index of tracker item IDs or links instead of a duplicate checklist. + +## Parallelization Opportunities + +When multiple agents or sessions are available: + +- **Safe to parallelize:** Independent feature slices, tests for already-implemented features, documentation +- **Must be sequential:** Database migrations, shared state changes, dependency chains +- **Needs coordination:** Features that share an API contract (define the contract first, then parallelize) + +## Common Rationalizations + +| Rationalization | Reality | +|---|---| +| "I'll figure it out as I go" | That's how you end up with a tangled mess and rework. 10 minutes of planning saves hours. | +| "The tasks are obvious" | Write them down anyway. Explicit tasks surface hidden dependencies and forgotten edge cases. | +| "Planning is overhead" | Planning is the task. Implementation without a plan is just typing. | +| "I can hold it all in my head" | Context windows are finite. Written plans survive session boundaries and compaction. | +| "The old `tmp/.plans/YYYY-MM-DD-HHMM-.md` is stale, I'll just replace it" | Unchecked tasks may be mid-build in another session. Overwriting them destroys work state that exists nowhere else. Stop and ask. | + +## Red Flags + +- Starting implementation without a written task list +- Overwriting a `tmp/.plans/YYYY-MM-DD-HHMM-.md` or `tmp/.plans/YYYY-MM-DD-HHMM-.md` that still has unchecked tasks for different work, without asking +- Writing `tmp/.plans/YYYY-MM-DD-HHMM-.md` when the project has designated an external tracker (or scattering tasks across both) +- Tasks that say "implement the feature" without acceptance criteria +- No verification steps in the plan +- All tasks are XL-sized +- No checkpoints between tasks +- Dependency order isn't considered + +## Verification + +Before starting implementation, confirm: + +- [ ] Every task has acceptance criteria +- [ ] Every task has a verification step +- [ ] Task dependencies are identified and ordered correctly +- [ ] Tasks are recorded in the task list target (default `tmp/.plans/YYYY-MM-DD-HHMM-.md`) +- [ ] No pre-existing incomplete plan was overwritten without explicit user confirmation +- [ ] No task touches more than ~5 files +- [ ] Checkpoints exist between major phases +- [ ] The human has reviewed and approved the plan + +## See Also + +Acceptance criteria are per-task and answer "did we build the right thing?". They sit on top of the project-wide Definition of Done, the standing bar every task clears before it counts as done. See `references/definition-of-done.md`. + + +## Local single-plan override + +This contract supersedes all conflicting output, task-list, tracker, directory, +and human-checkpoint instructions above and in this bundle's resources. + +- Create one implementation plan at `tmp/.plans/YYYY-MM-DD-HHMM-.md` in the project + root. Use the creation date and time and a stable dash-case topic. If that + filename already exists for different work, choose a distinct topic suffix; + never overwrite another incomplete plan. Revise the same plan in place only + when explicitly requested. +- Put the overview, decisions, risks, detailed tasks, checklist, acceptance + criteria, dependencies, exact authorized writable paths, and concrete checks + in that one file. Do not create a separate todo file or tracker items unless + the user explicitly requests them. Even then, keep the plan under `tmp/.plans/`. +- Create only `tmp/.plans/` for plan output, not `tasks/`. Reading this skill is + not permission to write a plan or implement it. +- Separate stable task contracts from mutable progress. Permit checkbox and + concise verification or blocker updates only in the progress section; they + do not change approval. Requirement or writable-scope changes need approval. +- Checkpoints are automatic checks, not human confirmation gates. Require a + human checkpoint only when the user explicitly requested that gate. Preserve + stops for missing decisions, unsafe actions, conflicts, or out-of-scope work. +- Planning alone does not start implementation. The user's invocation of + `/internal-gateway-execute-plans` approves the identified plan for execution; + an internal skill call does not manufacture that authorization. + diff --git a/.github/skills/addyosmani-planning-and-task-breakdown/references/definition-of-done.md b/.github/skills/addyosmani-planning-and-task-breakdown/references/definition-of-done.md new file mode 100644 index 00000000..35e39f9e --- /dev/null +++ b/.github/skills/addyosmani-planning-and-task-breakdown/references/definition-of-done.md @@ -0,0 +1,67 @@ +# Definition of Done + +A standing, project-wide bar that every change must clear before it counts as done. Unlike acceptance criteria, which vary per task and answer "did we build the right thing?", the Definition of Done is the same every time and answers "is this finished to our standard?". Use it as the final gate in `planning-and-task-breakdown`, `incremental-implementation`, and `shipping-and-launch`. + +## Definition of Done vs. Acceptance Criteria + +| | Acceptance Criteria | Definition of Done | +|---|---|---| +| Scope | Specific to one task or spec | Applies to every increment | +| Changes | Different for each item | Fixed and reused | +| Answers | "Did we build *this thing*?" | "Is it *ready*?" | +| Owner | Defined when planning the task | Defined once for the project | +| Example | "User can reset password via email link" | "Tests pass, no regressions, docs updated" | + +The two are complementary. A task is done only when **its** acceptance criteria are met **and** the standing Definition of Done is satisfied. Skipping either leaves work that looks finished but is not. + +## The Standing Checklist + +Apply this to every change before declaring it done. + +### Correctness +- [ ] All acceptance criteria for the task are met +- [ ] Code runs and behaves as intended, verified at runtime, not just compiled or typechecked +- [ ] New behavior is covered by tests that fail without the change and pass with it +- [ ] Existing tests still pass; no regressions introduced +- [ ] Edge cases and error paths are handled, not just the happy path + +### Quality +- [ ] Code reveals intent through naming and structure; no comments needed to explain *what* it does +- [ ] No duplicated business logic +- [ ] No dead code, debug output, or commented-out blocks left behind +- [ ] Changes are scoped to the task; no unrelated refactors snuck in +- [ ] Linting and formatting pass + +The depth behind these items lives in `code-review-and-quality` (the five-axis review) and `code-simplification` (reducing complexity without changing behavior). + +### Integration +- [ ] Change works with the rest of the system, not just in isolation +- [ ] Database migrations, config changes, and feature flags are accounted for +- [ ] Backward compatibility considered for any public interface or API change + +### Documentation +- [ ] Public interfaces, APIs, and user-facing behavior are documented +- [ ] Architectural decisions worth preserving are recorded (see `documentation-and-adrs`) +- [ ] Documentation describes the current state in timeless language, not the change history + +### Ship-readiness +- [ ] Security implications reviewed for any untrusted input, auth, or data handling (see `security-and-hardening`) +- [ ] Observability in place for new critical paths (logs, metrics, traces) (see `observability-and-instrumentation`) +- [ ] Rollback path exists for anything risky (see `shipping-and-launch`) +- [ ] The human has reviewed and approved before merge or deploy + +## How to Apply + +- **Per task**: confirm the Correctness and Quality sections before checking the task off. +- **Per feature**: confirm Integration and Documentation before considering the feature complete. +- **Per release**: the full checklist is the floor; `shipping-and-launch` adds the deploy-specific gates on top. + +Tailor the list to the project once, then reuse it unchanged. A Definition of Done that is renegotiated every sprint is not a Definition of Done. + +## Red Flags + +- "It's done, I just haven't run it yet": unverified work is not done. +- "Tests pass" used as a synonym for done while docs, regressions, or runtime verification are skipped. +- A different bar applied depending on deadline pressure. +- Acceptance criteria treated as the whole bar, with no standing quality floor. +- "Done" declared before human review on changes that need it. diff --git a/.github/skills/antigravity-cloudformation-best-practices/SKILL.md b/.github/skills/antigravity-cloudformation-best-practices/SKILL.md index 4e5ce451..cc11f2fd 100644 --- a/.github/skills/antigravity-cloudformation-best-practices/SKILL.md +++ b/.github/skills/antigravity-cloudformation-best-practices/SKILL.md @@ -1,7 +1,7 @@ --- name: antigravity-cloudformation-best-practices description: "CloudFormation template optimization, nested stacks, drift detection, and production-ready patterns. Use when writing or reviewing CF templates." -risk: unknown +risk: critical source: community date_added: "2026-02-27" --- diff --git a/.github/skills/antigravity-golang-pro/SKILL.md b/.github/skills/antigravity-golang-pro/SKILL.md index 8be7565b..dd6412d2 100644 --- a/.github/skills/antigravity-golang-pro/SKILL.md +++ b/.github/skills/antigravity-golang-pro/SKILL.md @@ -1,7 +1,7 @@ --- name: antigravity-golang-pro description: Master Go 1.21+ with modern patterns, advanced concurrency, performance optimization, and production-ready microservices. -risk: unknown +risk: critical source: community date_added: '2026-02-27' --- diff --git a/.github/skills/antigravity-grafana-dashboards/SKILL.md b/.github/skills/antigravity-grafana-dashboards/SKILL.md index 83e79c4e..1bde7015 100644 --- a/.github/skills/antigravity-grafana-dashboards/SKILL.md +++ b/.github/skills/antigravity-grafana-dashboards/SKILL.md @@ -1,7 +1,7 @@ --- name: antigravity-grafana-dashboards description: "Create and manage production-ready Grafana dashboards for comprehensive system observability." -risk: unknown +risk: critical source: community date_added: "2026-02-27" --- @@ -117,7 +117,7 @@ Design effective Grafana dashboards for monitoring applications, infrastructure, } ``` -**Reference:** See `assets/api-dashboard.json` +**Reference:** See [inline example](#api-monitoring-dashboard) ## Panel Types @@ -306,7 +306,7 @@ providers: - Pod count by namespace - Node status -**Reference:** See `assets/infrastructure-dashboard.json` +**Reference:** See [inline example](#infrastructure-dashboard) ### Database Dashboard @@ -319,7 +319,7 @@ providers: - Replication lag - Slow queries -**Reference:** See `assets/database-dashboard.json` +**Reference:** See [inline example](#database-dashboard) ### Application Dashboard @@ -373,9 +373,9 @@ resource "grafana_folder" "monitoring" { ## Reference Files -- `assets/api-dashboard.json` - API monitoring dashboard -- `assets/infrastructure-dashboard.json` - Infrastructure dashboard -- `assets/database-dashboard.json` - Database monitoring dashboard +- [inline example](#api-monitoring-dashboard) - API monitoring dashboard +- [inline example](#infrastructure-dashboard) - Infrastructure dashboard +- [inline example](#database-dashboard) - Database monitoring dashboard - `references/dashboard-design.md` - Dashboard design guide ## Related Skills diff --git a/.github/skills/antigravity-grafana-dashboards/references/dashboard-design.md b/.github/skills/antigravity-grafana-dashboards/references/dashboard-design.md new file mode 100644 index 00000000..2550e3b6 --- /dev/null +++ b/.github/skills/antigravity-grafana-dashboards/references/dashboard-design.md @@ -0,0 +1,13 @@ +# Dashboard design review + +## Inputs + +Name the operational question, service filter, datasource and expected units for each panel. + +## Procedure and verification + +Place impact and trend first, then diagnostic detail. Keep ratios on the same population and time range. Include descriptions, refresh time and a no-data state. Test a known incident, zero traffic, an empty variable selection and multiple service selections. Compare a hand-calculated fixture to the displayed value. + +## Limitations + +Dashboard JSON depends on the Grafana version. Export from the installed instance and test import in a staging folder; panel colors and labels alone do not establish an alert policy. diff --git a/.github/skills/antigravity-grafana-dashboards/resources/implementation-playbook.md b/.github/skills/antigravity-grafana-dashboards/resources/implementation-playbook.md new file mode 100644 index 00000000..6c386c51 --- /dev/null +++ b/.github/skills/antigravity-grafana-dashboards/resources/implementation-playbook.md @@ -0,0 +1,23 @@ +# Grafana dashboard implementation + +## Inputs + +Installed Grafana version, datasource UID, metric names/labels, viewer role and the operational question. + +## Procedure + +1. Inspect real series and their units before drafting panels. Choose a service filter and time window; keep numerator and denominator populations identical. +2. Build one representative panel in the target Grafana version and export its supported schema. Keep secrets out of exported JSON and use stable datasource references. +3. Validate normal traffic, no traffic, missing data and a known incident. Check variables, units, thresholds and query cost, then import into a staging folder and inspect the rendered result. + +## Worked example + +For an API error panel, compare a known 5-error/100-request sample with the displayed 5%. An empty source must show no data, not healthy zero. + +## Verification and handoff + +Report the actual files or configuration changed, checks performed, observed results and any untested environment. Keep the original inputs and evidence sufficient to reproduce the conclusion. + +## Limitations + +Legacy graph and embedded-alert JSON is not universally importable. Generate schema from the installed version and configure alert rules through its supported interface. diff --git a/.github/skills/antigravity-kubernetes-architect/SKILL.md b/.github/skills/antigravity-kubernetes-architect/SKILL.md index 9fb904f6..3a72a45e 100644 --- a/.github/skills/antigravity-kubernetes-architect/SKILL.md +++ b/.github/skills/antigravity-kubernetes-architect/SKILL.md @@ -1,7 +1,7 @@ --- name: antigravity-kubernetes-architect description: Expert Kubernetes architect specializing in cloud-native infrastructure, advanced GitOps workflows (ArgoCD/Flux), and enterprise container orchestration. -risk: unknown +risk: critical source: community date_added: '2026-02-27' --- diff --git a/.github/skills/antonbabenko-terraform-skill/references/ci-cd-workflows.md b/.github/skills/antonbabenko-terraform-skill/references/ci-cd-workflows.md index 65cc0ccf..ca33a380 100644 --- a/.github/skills/antonbabenko-terraform-skill/references/ci-cd-workflows.md +++ b/.github/skills/antonbabenko-terraform-skill/references/ci-cd-workflows.md @@ -23,7 +23,7 @@ This document provides detailed CI/CD workflow templates and optimization strate ```yaml # .github/workflows/terraform.yml -name: terraform-terraform-skill +name: antonbabenko-terraform-skill on: push: diff --git a/.github/skills/grill-me/SKILL.md b/.github/skills/grill-me/SKILL.md index 207e836e..22e68478 100644 --- a/.github/skills/grill-me/SKILL.md +++ b/.github/skills/grill-me/SKILL.md @@ -3,50 +3,51 @@ name: grill-me description: A relentless interview to sharpen a plan or design. --- -# Grill Me +Interview the user relentlessly until you reach a shared understanding. Map this as a **design tree**: every decision branches into the decisions that hang off it. -## Referenced skills +Work the tree in **rounds**. The **frontier** is every decision whose prerequisites are already settled: the questions you can ask _now_ without guessing at answers you haven't heard yet. Ask the whole frontier in one round: number each question and give your recommended answer. Then wait for the user's answers before the next round. -- None. +Format a round like so: -## Interview Approach +``` +❓ **Q1** - ****: -Interview me relentlessly about every aspect of this plan, design, or action -context until we reach a shared understanding. Walk down each branch of the -decision tree, resolving dependencies between decisions one-by-one. +➡️ -Before asking, inspect the repository, codebase, documentation, or local files for answers that can be recovered from evidence. - -By default, ask the full initial question set in one numbered list. - -Structure the list by decision branch and dependency order. Start with goal and scope, then constraints, architecture or options, risks and failure modes, rollout, and validation as relevant. - -For each numbered question, use this format: Question, Recommendation, Why, and Default if accepted. Make the recommendation detailed enough to explain what you want to decide, why it matters, and what answer you would choose by default. - -Treat your recommendations as accepted unless the user says otherwise. The user may override any recommendation by referencing the question number or giving different direction. - -Do not treat accepted recommendations as the end of the grilling process. If accepted defaults create contradictions, weak assumptions, unresolved risks, or dependent decisions, surface them explicitly. - -After the initial numbered list, ask one question at a time only for unresolved ambiguity, dependent follow-up decisions, or branches that cannot be settled from the user's bulk response. - -A caller may override the follow-up pacing with iterative numbered blocks. When a caller declares that override, replace the default one-at-a-time follow-up with the caller's pacing. - -Do not ask questions that can be answered by exploring the codebase, documentation, or local files. - -End by summarizing the resolved decisions, explicit assumptions, and any -unresolved questions the user chose to accept or defer. - - -## Local guided-question contract - -This repository-owned contract overrides any earlier instruction to ask one question at a time. +--- -- Ask all currently known questions in numbered bulk question blocks. -- Use `Question`, `Recommendation`, `Why`, and `Default if accepted` for every - numbered question. -- Make `Recommendation` the suggested answer and `Why` its concrete rationale. -- Keep each question, recommendation, and reason brief, clear, and - decision-ready. -- Put unresolved follow-ups in another numbered block. If only one blocking - question remains, present it as a numbered one-item block. - +❓ **Q2** - ****: + +➡️ +``` + +Each round the user answers reshapes the tree: settled decisions push the frontier outward and unblock questions that depended on them. Recompute the frontier and ask the next round. A question whose answer depends on another question still open in this round belongs to a _later_ round, not this one. + +Finding _facts_ is your job, never the user's. When a frontier question needs a fact from the environment (filesystem, tools, etc.), dispatch a sub-agent to find it; don't ask the user for anything you could look up yourself. Don't block on it: a running exploration is an unsettled prerequisite, so only the questions downstream of it wait for the sub-agent to report; ask the rest of the frontier now. The _decisions_ are the user's: put each to them and wait. + +The session is done when the frontier is empty: every branch of the design tree visited, nothing left silently assumed. Do not act on it until the user confirms you have reached a shared understanding. + + +## Scope and convergence guardrail + +Narrow tree expansion without replacing the existing rounds, frontier traversal, +fact recovery, or final user-confirmation mechanics. + +- Establish an interview envelope from caller context and the user's request: + subject, desired outcome, scope, anti-scope, and requested level of detail. +- Infer the envelope when it is already clear. Ask for clarification only when + an ambiguity would materially change the interview. +- Before every later round, admit a candidate question only when it maps to an + unresolved user decision with settled prerequisites and its answer could + materially change the outcome, recommendation, acceptance criteria, or a + material risk inside the envelope. +- Prune duplicate, cosmetic, speculative, premature implementation, and + adjacent-improvement branches. +- Treat each answer as input to the existing decisions, not implicit permission + to broaden the subject. Park useful out-of-scope items until the user + explicitly accepts a scope change. +- Treat the frontier as empty when no material in-scope decision remains, even + when conceivable downstream questions exist. +- Interpret “every branch” as every material, decision-relevant branch inside + the agreed interview envelope. + diff --git a/.github/skills/grill-me/agents/openai.yaml b/.github/skills/grill-me/agents/openai.yaml index 79d19f2f..faa7479f 100644 --- a/.github/skills/grill-me/agents/openai.yaml +++ b/.github/skills/grill-me/agents/openai.yaml @@ -1,3 +1,3 @@ interface: - display_name: "grill-me" - short_description: "Relentless interview to sharpen a plan or design" + display_name: Grill Me + short_description: Sharpen a plan through interview diff --git a/.github/skills/grilling/SKILL.md b/.github/skills/grilling/SKILL.md new file mode 100644 index 00000000..46fbd748 --- /dev/null +++ b/.github/skills/grilling/SKILL.md @@ -0,0 +1,6 @@ +--- +disable-model-invocation: true +name: grilling +description: Grill the user relentlessly about a plan, decision, or idea. Use when the user wants to stress-test their thinking, or uses any 'grill' trigger phrases. +--- +Run a `/grill-me` session. diff --git a/.github/skills/grilling/agents/openai.yaml b/.github/skills/grilling/agents/openai.yaml new file mode 100644 index 00000000..4ebba8aa --- /dev/null +++ b/.github/skills/grilling/agents/openai.yaml @@ -0,0 +1,5 @@ +interface: + display_name: Grilling + short_description: Stress-test thinking a round of questions at a time +policy: + allow_implicit_invocation: false diff --git a/.github/skills/internal-agent-creator/SKILL.md b/.github/skills/internal-agent-creator/SKILL.md index 1a3efb3d..7049a452 100644 --- a/.github/skills/internal-agent-creator/SKILL.md +++ b/.github/skills/internal-agent-creator/SKILL.md @@ -131,7 +131,7 @@ kept clear without a new reusable owner, stop and surface that boundary. ## Authoring Workflow 1. Run a proportional requirements gate. - Resolve purpose, route, inputs, actions, invocation boundary, risk, output, and validation. Ask only questions that materially change the contract. + Resolve purpose, route, inputs, actions, invocation boundary, risk, output, and validation. Route every question that materially changes the contract through `/grill-me`; do not ask material questions in an ad-hoc route. 2. Define the operating role and stance in one sentence. Translate persona language into observable behavior. Do not invent credentials or prestige. 3. Validate a new name and target path. diff --git a/.github/skills/internal-agent-creator/agents/openai.yaml b/.github/skills/internal-agent-creator/agents/openai.yaml index bc14451a..867359a2 100644 --- a/.github/skills/internal-agent-creator/agents/openai.yaml +++ b/.github/skills/internal-agent-creator/agents/openai.yaml @@ -1,4 +1,4 @@ interface: display_name: "internal-agent-creator" short_description: "Create, refine, or realign repository-owned Copilot agents" - default_prompt: "Use $internal-agent-creator to create or update an internal agent in .github/agents/ following the repository contract." + default_prompt: "Use $internal-agent-creator to create, review, or update an internal agent in .github/agents/ following the repository contract." diff --git a/.github/skills/internal-agent-creator/references/requirements-and-persona.md b/.github/skills/internal-agent-creator/references/requirements-and-persona.md index 087514d1..eba7d973 100644 --- a/.github/skills/internal-agent-creator/references/requirements-and-persona.md +++ b/.github/skills/internal-agent-creator/references/requirements-and-persona.md @@ -16,9 +16,9 @@ Resolve only the inputs that materially change the agent contract: - role-specific output and validation expectations Inspect repository evidence before asking the user for facts that are already -available. Ask one focused question at a time when an unresolved choice would -materially change the result. Do not force an interview when the request and -local contract already make the answer deterministic. +available. Route an unresolved choice that would materially change the result +through `/grill-me` rather than an ad-hoc question. Do not force an interview +when the request and local contract already make the answer deterministic. End the gate with one sentence that states the target role, its main boundary, and the observable result it must produce. diff --git a/.github/skills/internal-aws-governance/SKILL.md b/.github/skills/internal-aws-governance/SKILL.md deleted file mode 100644 index a2bd6b6b..00000000 --- a/.github/skills/internal-aws-governance/SKILL.md +++ /dev/null @@ -1,47 +0,0 @@ ---- -name: internal-aws-governance -description: Use when /internal-aws selects the AWS governance lane for IAM operating models, trust policies, federation, permission boundaries, SCPs, tag policies, exception controls, or access guardrails. ---- - -# Internal AWS Governance - -Own AWS IAM, trust, SCP, federation, permission-boundary, and access-guardrail -decisions. Separate org-level guardrails from account-level grants and keep -permission decisions auditable. - -## When to use - -- IAM operating-model guidance across accounts, including role and group strategy. -- SCP, trust, federation, permission-boundary, or session-constraint guidance, including role assumption. -- Separating preventive controls from granted permissions. -- Guardrail design, exception handling, or access-governance review, including identity and access security guardrails. - -## Core rules - -- Keep org-level guardrails distinct from account-level grants. -- Treat SCPs as limits on maximum permission, not as grants. -- Prefer roles and federation over long-lived IAM users unless a proven reason exists. -- Make scope explicit: root, OU, account set, or single account. -- Make exception handling explicit when a control is not universal. -- On `AccessDenied` with "explicit deny in a service control policy", treat the SCP as the blocking control; target-account IAM grants cannot override an explicit SCP deny. - -Load `references/guardrail-map.md` when the governance surface is ambiguous or a deeper split between IAM, trust, SCP, and boundary controls is needed. - -## Common mistakes - -| Mistake | Why it matters | Instead | -|---|---|---| -| Using SCPs as if they grant access | Preventive controls mistaken for execution permissions | Pair SCP guidance with the required IAM grant path | -| Answering without naming scope | Root, OU, and account controls behave differently | State the exact governance scope before recommending a mechanism | -| Mixing org-wide guardrails and in-account authorization into one vague recommendation | Reviewers cannot see which control prevents versus grants | Separate the org-level mechanism from the account-level design | -| Proposing break-glass access without boundaries or audit expectations | Emergency access becomes standing privilege with weak accountability | Define who can invoke it, how it is bounded, and what audit evidence must exist | -| Recommending rollout without simulation when blast radius is high | A wide deny or trust failure can interrupt platform operations | Use simulation, targeted rollout, and explicit rollback triggers before widening | -| Treating permission boundaries as a replacement for trust design | Delegation stays too broad even if identity policies are constrained | Use boundaries to limit delegated builders and trust policies to control who assumes the role | - -## Completion contract - -- Governance scope is explicit: root, OU, account set, or single account. -- Recommended mechanism is clear about whether it prevents, grants, or constrains. -- Trust boundaries and exception paths are explicit when access crosses account boundaries. -- Staged validation or simulation is named before high-blast-radius rollout. -- Expected effects and evidence requirements are explicit for every proposed control. diff --git a/.github/skills/internal-aws-governance/agents/openai.yaml b/.github/skills/internal-aws-governance/agents/openai.yaml deleted file mode 100644 index be39687a..00000000 --- a/.github/skills/internal-aws-governance/agents/openai.yaml +++ /dev/null @@ -1,6 +0,0 @@ -policy: - allow_implicit_invocation: false -interface: - display_name: "internal-aws-governance" - short_description: "AWS IAM, SCP, and guardrail guidance" - default_prompt: "Use $internal-aws-governance for this routed AWS governance task." diff --git a/.github/skills/internal-aws-governance/references/guardrail-map.md b/.github/skills/internal-aws-governance/references/guardrail-map.md deleted file mode 100644 index a36855ad..00000000 --- a/.github/skills/internal-aws-governance/references/guardrail-map.md +++ /dev/null @@ -1,50 +0,0 @@ -# AWS Governance Guardrail Map - -Use this reference when the user needs a clearer split between AWS governance surfaces. - -## Quick split - -| Need | Use first | Why | -| --- | --- | --- | -| Limit what principals can ever do across many accounts | SCP | Org-level preventive guardrail | -| Define what a role or workload can do in one account | IAM policy plus trust policy | Execution-level authorization | -| Constrain delegated builders or automation | Permission boundary or session policy | Limits delegated execution | -| Standardize metadata expectations | Tag policy plus local enforcement | Governance consistency | -| Permit emergency access | Break-glass role design plus audit path | Exceptional access with visibility | - -## Control design checklist - -1. State the scope: root, OU, account set, account, principal, or session. -2. State whether the mechanism prevents, grants, constrains, delegates, or creates an exception. -3. Record the trust boundary and exception path. -4. Name the simulation or staged validation required before rollout. - -## Control placement - -- Keep placement decisions separate from permission and guardrail effects. -- Record the expected observable effect for each control. - -## Placement patterns - -| Scenario | Prefer | Why | -| --- | --- | --- | -| Block risky services or regions across many accounts | SCP at root or OU | Central preventive control with explicit blast radius | -| Allow a workload role to use only the resources in one account | IAM policy plus trust policy | Grants the action at the execution boundary | -| Let platform engineers create roles without giving them unrestricted permissions | Permission boundary plus scoped creation role | Limits delegated builders without replacing trust design | -| Standardize required tags for cost or ownership workflows | Tag policy plus enforcement in deployment paths | Keeps metadata expectations central but still operationally enforceable | - -## Trust-boundary examples - -| Need | Primary control | Review note | -| --- | --- | --- | -| Human access from an external IdP into AWS accounts | Federation plus tightly scoped assume-role paths | Keep identity-source trust separate from account authorization | -| CI or automation assuming deployment roles across accounts | Trust policy constrained to named principals and conditions | Make environment scope and break-glass path explicit | -| Shared security tooling reading logs across accounts | Resource policy or role assumption with read-only scope | Prefer narrow data-access roles over broad admin trust | - -## Exception patterns with audit expectations - -| Exception type | Pattern | Audit expectation | -| --- | --- | --- | -| Temporary break-glass for incident response | Time-bounded role path with explicit approver and logging | Record who approved, who assumed the role, and when access ended | -| Service team needs an OU-level SCP carve-out | Targeted OU or account exception with expiry review | Record the business reason, compensating controls, and review date | -| Automation cannot yet satisfy a required tag or policy condition | Narrow deployment exception with compensating report | Track affected accounts or resources and the closure plan | diff --git a/.github/skills/internal-aws-lambda/SKILL.md b/.github/skills/internal-aws-lambda/SKILL.md index 6f68277d..a447cba7 100644 --- a/.github/skills/internal-aws-lambda/SKILL.md +++ b/.github/skills/internal-aws-lambda/SKILL.md @@ -1,47 +1,95 @@ --- name: internal-aws-lambda -description: Use when /internal-aws selects the AWS Lambda lane for handlers, event sources, runtimes, packaging, retries, concurrency, cold starts, or Lambda-specific configuration. +description: Use when implementing or reviewing an AWS Lambda function contract, including the handler boundary, event source, retries and idempotency, partial batch failure, concurrency, timeouts, packaging, cold starts, or current Lambda limits and runtime support. Route execution-role trust, account, or guardrail design to /internal-aws, and language-only refactoring to the language skill. --- # Internal AWS Lambda -Own AWS Lambda handler, event-source, runtime, packaging, retry, concurrency, -cold-start, and configuration behavior. +Own the AWS Lambda function contract: handler boundary, event source, +runtime, packaging, retry, concurrency, cold-start, and configuration +behavior. ## When to use - Implementing or reviewing Lambda handlers. -- Designing API Gateway, Function URL, or other HTTP-triggered request/response handling. -- Designing SQS-triggered batch processing, retry, DLQ, or partial-batch-failure flows. -- Packaging, dependency, cold-start, VPC, or runtime-configuration choices specific to Lambda. +- Designing API Gateway, Function URL, or other HTTP-triggered + request/response handling. +- Designing SQS-triggered batch processing, retry, DLQ, or + partial-batch-failure flows. +- Packaging, dependency, cold-start, VPC, or runtime-configuration choices + specific to Lambda. -## Core guidance +Hand off other work: -- Keep the handler as a transport adapter; move business logic to testable helpers. +- execution-role trust, cross-account permissions, SCPs, or account + placement: `/internal-aws`; +- language structure and tests that do not depend on the Lambda contract: + the language skill, such as `/internal-python-project` or + `/internal-nodejs-project`; +- Terraform or CloudFormation authoring: `/internal-terraform` or + `/antigravity-cloudformation-best-practices`. + +## Core rules + +- Keep the handler as a transport adapter; move business logic to testable + helpers. - Code to one event-source contract at a time: HTTP, queue, schedule, or async. -- Parse and validate inputs at the boundary; normalize data passed to business logic. -- Initialize AWS clients outside the handler when reuse is safe; keep imports small. -- Prefer modular SDK clients and narrow dependencies over broad convenience packages. -- Size timeout, memory, concurrency, batch size, and queue visibility timeout as one operating profile. -- Treat duplicate delivery, retries, and idempotency as normal for async triggers. -- Use environment variables for configuration; fetch secrets from managed secret stores. -- Log stable identifiers (request IDs, message IDs); do not log raw sensitive payloads by default. +- Parse and validate inputs at the boundary; normalize data passed to business + logic. +- Initialize AWS clients outside the handler when reuse is safe; keep imports + small. +- Prefer modular SDK clients and narrow dependencies over broad convenience + packages. +- Size timeout, memory, concurrency, batch size, and queue visibility timeout + as one operating profile. +- Treat duplicate delivery, retries, and idempotency as normal for async + triggers. +- Use environment variables for configuration; fetch secrets from managed + secret stores. +- Log stable identifiers (request IDs, message IDs); do not log raw sensitive + payloads by default. +- Attach a function to a VPC only for a concrete private dependency; then + name its DNS and egress path. ## Event-source guidance -- **HTTP**: normalize body, path, and query once; return transport-compatible JSON with explicit headers; keep CORS intentional. -- **SQS**: process records independently; handle poison messages explicitly; return only failed item identifiers when partial batch retry is enabled. -- **Scheduled**: make time-window assumptions explicit; guard against duplicate or overlapping execution. -- **File-driven**: avoid recursive triggers by separating input and output prefixes or buckets. +- **HTTP**: normalize body, path, and query once; return transport-compatible + JSON with explicit headers; keep CORS intentional. +- **SQS**: process records independently; handle poison messages explicitly; + return only failed item identifiers when partial batch retry is enabled. +- **Scheduled**: make time-window assumptions explicit; guard against duplicate + or overlapping execution. +- **File-driven**: avoid recursive triggers by separating input and output + prefixes or buckets. + +## Freshness + +Timeouts, payload and `/tmp` limits, runtime support, and newer features +change. Verify a number against current AWS documentation before stating +it; otherwise mark it unverified. + +Optional enrichment: when AWS Knowledge MCP is available, AWS publishes an +`aws-serverless` agent skill for SnapStart, Powertools, event source +mappings, and current limits. Discover the exact `skill_name` with +`search_documentation` topic `agent_skills`, then load it with +`retrieve_skill`. This skill stays complete without it. + +## References -Load `references/examples.md` for minimal handler patterns and event-source checklists. -Load `references/sharp-edges.md` when diagnosing cold starts, VPC latency, retry storms, response-shape mismatches, or file-ingest recursion. -Load `references/common-mistakes.md` for the full mistake table. +- [references/examples.md](references/examples.md): load for minimal handler + patterns and event-source checklists. +- [references/sharp-edges.md](references/sharp-edges.md): load when diagnosing + cold starts, VPC latency, retry storms, duplicate side effects, + response-shape mismatches, file-ingest recursion, or secret exposure. -## Completion contract +## Completion criteria - Unit tests run outside the Lambda runtime; AWS boundaries are mocked. - HTTP handlers: malformed body, path, query, and error-response cases tested. -- Queue consumers: duplicate delivery, poison messages, timeout pressure, and partial batch failure tested. -- Code assumptions validated together with deployed timeout, memory, event-source, and queue configuration. -- Event contract, boundary normalization, retry behavior, and idempotency are explicit. +- Queue consumers: duplicate delivery, poison messages, timeout pressure, and + partial batch failure tested. +- Code assumptions validated together with deployed timeout, memory, + event-source, and queue configuration. +- Event contract, boundary normalization, retry behavior, and idempotency are + explicit. +- Numeric limits are sourced or marked unverified. diff --git a/.github/skills/internal-aws-lambda/agents/openai.yaml b/.github/skills/internal-aws-lambda/agents/openai.yaml index e8c96fdc..9fef2fe4 100644 --- a/.github/skills/internal-aws-lambda/agents/openai.yaml +++ b/.github/skills/internal-aws-lambda/agents/openai.yaml @@ -1,6 +1,6 @@ policy: - allow_implicit_invocation: false + allow_implicit_invocation: true interface: display_name: "internal-aws-lambda" short_description: "AWS Lambda runtime and event-source patterns" - default_prompt: "Use $internal-aws-lambda for this routed AWS Lambda task." + default_prompt: "Use $internal-aws-lambda for AWS Lambda handler, event-source, runtime, packaging, retry, concurrency, or cold-start work." diff --git a/.github/skills/internal-aws-lambda/references/common-mistakes.md b/.github/skills/internal-aws-lambda/references/common-mistakes.md deleted file mode 100644 index ac231616..00000000 --- a/.github/skills/internal-aws-lambda/references/common-mistakes.md +++ /dev/null @@ -1,12 +0,0 @@ -# Common mistakes - -| Mistake | Why it matters | Instead | -| --- | --- | --- | -| Treating the handler as the business layer | Hard to test, high coupling to AWS events | Keep the handler thin and call testable helpers | -| Parsing every event inline differently | Inconsistent behavior and brittle error handling | Normalize one event contract per trigger type | -| Returning whole-batch failure for one bad SQS record | Causes replay storms and blocks healthy messages | Catch per-record failures and return failed IDs only when supported | -| Ignoring idempotency on async triggers | Duplicate deliveries create data corruption or repeated side effects | Use idempotent writes, dedupe keys, or safe upserts | -| Shipping large shared bundles to every function | Slower cold starts and harder ownership boundaries | Split by function responsibility and keep dependencies narrow | -| Putting secrets directly in environment variables | Rotation and exposure risks | Use Secrets Manager or SSM-backed retrieval patterns | -| Attaching Lambda to a VPC by default | Adds latency and networking failure modes | Keep it out of a VPC unless there is a concrete dependency | -| Mixing HTTP response shaping with core logic | Hard to reuse and easy to break integrations | Keep request mapping and response mapping at the edge | diff --git a/.github/skills/internal-aws-lambda/references/examples.md b/.github/skills/internal-aws-lambda/references/examples.md index ffcb92c4..8b57524f 100644 --- a/.github/skills/internal-aws-lambda/references/examples.md +++ b/.github/skills/internal-aws-lambda/references/examples.md @@ -38,7 +38,8 @@ function json(statusCode, payload) { Node.js note: -- Set `context.callbackWaitsForEmptyEventLoop = false` only when open handles are intentional and understood. +- Set `context.callbackWaitsForEmptyEventLoop = false` only when open handles + are intentional and understood. ## SQS batch consumer shape @@ -81,11 +82,13 @@ Use this pattern when the function runs from EventBridge or another scheduler: - Accept that schedules can arrive late, retry, or overlap. - Make the target time window explicit in logs and downstream calls. -- Protect side effects with idempotent markers or a coordination lock when overlap would be harmful. +- Protect side effects with idempotent markers or a coordination lock when + overlap would be harmful. ## Packaging checklist - Keep one function package focused on one responsibility. - Prefer modular SDK clients over broad imports. - Measure init duration before assuming the bottleneck is invocation logic. -- If multiple functions share the same large dependency, consider a versioned layer only after confirming the tradeoff is worth it. +- If multiple functions share the same large dependency, consider a versioned + layer only after confirming the tradeoff is worth it. diff --git a/.github/skills/internal-aws-lambda/references/sharp-edges.md b/.github/skills/internal-aws-lambda/references/sharp-edges.md index ce81b00c..f24fcfe6 100644 --- a/.github/skills/internal-aws-lambda/references/sharp-edges.md +++ b/.github/skills/internal-aws-lambda/references/sharp-edges.md @@ -11,3 +11,6 @@ | File workflows recurse forever | Input and output use the same trigger surface | Bucket notification scope, prefixes, output location | Separate input and output prefixes or buckets | | Node.js invocations hang after work is done | Open sockets, timers, or connection pools keep the event loop alive | Client lifecycle, timers, runtime settings | Close unused handles or intentionally set `callbackWaitsForEmptyEventLoop = false` | | Secrets leak into logs or configs | Secrets are treated like normal configuration values | Environment variables, startup logs, error serialization | Use managed secret stores and scrub logs by default | +| Duplicate side effects after retries | Async and queue triggers deliver at least once | Write paths, dedupe keys, idempotency markers | Use idempotent writes, dedupe keys, or safe upserts | +| One large shared bundle slows every function | Unrelated dependencies ship to every function | Package size per function, shared imports | Split by function responsibility and keep dependencies narrow | +| Integrations break when core logic changes | HTTP request and response mapping is mixed into business code | Where parsing and response shaping live | Keep request and response mapping at the edge | diff --git a/.github/skills/internal-aws-lambda/tests/evaluation/evals.json b/.github/skills/internal-aws-lambda/tests/evaluation/evals.json new file mode 100644 index 00000000..e860ca7a --- /dev/null +++ b/.github/skills/internal-aws-lambda/tests/evaluation/evals.json @@ -0,0 +1,443 @@ +{ + "schema": "skill-eval-pack/v1", + "skill": "internal-aws-lambda", + "requirements": [ + { + "id": "R-THIN-HANDLER", + "text": "Keep the handler a transport adapter over testable business logic.", + "source": "internal-aws-lambda SKILL.md: Core rules" + }, + { + "id": "R-EVENT-CONTRACT", + "text": "Code to one event-source contract and validate input at the boundary.", + "source": "internal-aws-lambda SKILL.md: Core rules" + }, + { + "id": "R-IDEMPOTENCY", + "text": "Treat duplicate delivery as normal for async and queue triggers.", + "source": "internal-aws-lambda SKILL.md: Core rules" + }, + { + "id": "R-PARTIAL-BATCH", + "text": "Return only failed item identifiers when partial batch response is enabled.", + "source": "internal-aws-lambda SKILL.md: Event-source guidance" + }, + { + "id": "R-OPERATING-PROFILE", + "text": "Size timeout, memory, concurrency, batch size, and queue visibility timeout together.", + "source": "internal-aws-lambda SKILL.md: Core rules" + }, + { + "id": "R-RECURSION", + "text": "Separate input and output surfaces for file-driven triggers.", + "source": "internal-aws-lambda SKILL.md: Event-source guidance" + }, + { + "id": "R-SECRETS", + "text": "Fetch secrets from managed stores and do not log raw sensitive payloads by default.", + "source": "internal-aws-lambda SKILL.md: Core rules" + }, + { + "id": "R-VPC-DEFAULT", + "text": "Attach a function to a VPC only for a concrete private dependency.", + "source": "internal-aws-lambda SKILL.md: Core rules" + }, + { + "id": "R-LIMIT-FRESHNESS", + "text": "Verify numeric Lambda limits and runtime support, or mark them unverified.", + "source": "internal-aws-lambda SKILL.md: Freshness" + }, + { + "id": "R-HANDOFF", + "text": "Hand off execution-role trust and guardrails to /internal-aws and language-only work to the language skill.", + "source": "internal-aws-lambda SKILL.md: Boundaries" + } + ], + "cases": [ + { + "id": "C-SQS-WHOLE-BATCH", + "family": "defective-artifact", + "kind": "deterministic", + "requirement_ids": [ + "R-PARTIAL-BATCH", + "R-IDEMPOTENCY" + ], + "prompt": "Review this SQS consumer. One malformed message keeps the whole batch retrying.", + "initial_state": "Partial batch response is enabled on the event source mapping.", + "expected_output": "The review replaces the whole-batch failure with per-record handling that returns only failed message IDs and notes idempotency.", + "files": [ + "tests/evaluation/fixtures/sqs-whole-batch-failure.py" + ], + "assertions": [ + { + "id": "A-PER-RECORD", + "text": "The review returns batchItemFailures with only the failed message IDs.", + "critical": true + }, + { + "id": "A-IDEMPOTENT", + "text": "The review requires the processing step to be idempotent.", + "critical": false + } + ], + "forbidden_actions": [ + "Keep raising on the first bad record." + ], + "status": "not-run", + "held_out": false, + "defective_fixture": "tests/evaluation/fixtures/sqs-whole-batch-failure.py" + }, + { + "id": "C-S3-RECURSION", + "family": "defective-artifact", + "kind": "deterministic", + "requirement_ids": [ + "R-RECURSION" + ], + "prompt": "Review this SAM template. Costs spiked after the thumbnail function went live.", + "initial_state": "The function writes output to the prefix that triggers it.", + "expected_output": "The review identifies the recursive trigger and separates input and output prefixes or buckets.", + "files": [ + "tests/evaluation/fixtures/s3-recursive-trigger.yaml" + ], + "assertions": [ + { + "id": "A-RECURSION", + "text": "The review names the same-prefix trigger loop as the cause.", + "critical": true + } + ], + "forbidden_actions": [ + "Recommend only a concurrency limit as the fix." + ], + "status": "not-run", + "held_out": false, + "defective_fixture": "tests/evaluation/fixtures/s3-recursive-trigger.yaml" + }, + { + "id": "C-TIMEOUT-MISMATCH", + "family": "defective-artifact", + "kind": "deterministic", + "requirement_ids": [ + "R-OPERATING-PROFILE" + ], + "prompt": "Review this queue and function configuration; messages are processed twice under load.", + "initial_state": "The queue visibility timeout is lower than the function timeout.", + "expected_output": "The review raises the visibility timeout above the function timeout and sizes batch and concurrency together.", + "files": [ + "tests/evaluation/fixtures/sqs-timeout-mismatch.yaml" + ], + "assertions": [ + { + "id": "A-VISIBILITY", + "text": "The review states that the visibility timeout must exceed the function timeout.", + "critical": true + } + ], + "forbidden_actions": [ + "Only lower the batch size." + ], + "status": "not-run", + "held_out": true, + "defective_fixture": "tests/evaluation/fixtures/sqs-timeout-mismatch.yaml" + }, + { + "id": "C-THIN-HANDLER", + "family": "design", + "kind": "rubric", + "requirement_ids": [ + "R-THIN-HANDLER", + "R-EVENT-CONTRACT" + ], + "prompt": "Implement an API Gateway handler that validates an order and stores it in DynamoDB.", + "initial_state": "HTTP API proxy integration.", + "expected_output": "A thin handler that parses and validates at the boundary, calls a testable helper, and returns a normalized response.", + "files": [], + "assertions": [ + { + "id": "A-THIN", + "text": "Business logic lives outside the handler in a testable function.", + "critical": true + }, + { + "id": "A-VALIDATE", + "text": "Malformed body and missing fields return a 4xx response.", + "critical": true + } + ], + "forbidden_actions": [], + "status": "not-run", + "held_out": false, + "rubric": { + "pass": [ + "Thin adapter with boundary validation." + ], + "fail": [ + "Business logic inline in the handler." + ] + } + }, + { + "id": "C-SECRET-ENV", + "family": "security", + "kind": "rubric", + "requirement_ids": [ + "R-SECRETS" + ], + "prompt": "Store the payment API key in the function environment variables and log the request for debugging.", + "initial_state": "The key is a production secret.", + "expected_output": "The response uses Secrets Manager or Parameter Store retrieval and logs identifiers, not raw payloads.", + "files": [], + "assertions": [ + { + "id": "A-SECRET-STORE", + "text": "The response fetches the secret from a managed store.", + "critical": true + }, + { + "id": "A-NO-RAW-LOG", + "text": "The response avoids logging the raw payload or secret.", + "critical": true + } + ], + "forbidden_actions": [ + "Put the secret in a plaintext environment variable." + ], + "status": "not-run", + "held_out": false, + "rubric": { + "pass": [ + "Managed secret retrieval and safe logging." + ], + "fail": [ + "Secret in env var or raw payload logged." + ] + } + }, + { + "id": "C-VPC-DEFAULT", + "family": "configuration", + "kind": "rubric", + "requirement_ids": [ + "R-VPC-DEFAULT" + ], + "prompt": "Attach all our Lambda functions to the shared VPC for security.", + "initial_state": "Most functions call only public AWS APIs.", + "expected_output": "The response attaches only functions with a private dependency and names DNS and egress needs for those.", + "files": [], + "assertions": [ + { + "id": "A-VPC-NEED", + "text": "The response ties VPC attachment to a concrete private dependency.", + "critical": true + } + ], + "forbidden_actions": [], + "status": "not-run", + "held_out": true, + "rubric": { + "pass": [ + "VPC use justified per function." + ], + "fail": [ + "Blanket VPC attachment." + ] + } + }, + { + "id": "C-LIMIT-FRESHNESS", + "family": "freshness", + "kind": "rubric", + "requirement_ids": [ + "R-LIMIT-FRESHNESS" + ], + "prompt": "What is the maximum Lambda timeout for an SQS event source mapping today?", + "initial_state": "Limits have changed recently; no source is attached.", + "expected_output": "The response verifies the number against current AWS documentation or marks it unverified.", + "files": [], + "assertions": [ + { + "id": "A-SOURCED", + "text": "The numeric limit is sourced or explicitly marked unverified.", + "critical": true + } + ], + "forbidden_actions": [ + "State a number from memory as settled fact." + ], + "status": "not-run", + "held_out": false, + "rubric": { + "pass": [ + "Sourced or flagged number." + ], + "fail": [ + "Unsourced number." + ] + } + }, + { + "id": "C-TRUST-HANDOFF", + "family": "competing-owner", + "kind": "rubric", + "requirement_ids": [ + "R-HANDOFF" + ], + "prompt": "Review the execution role trust policy and the cross-account permissions of our Lambda deployment pipeline.", + "initial_state": "The request is about trust and permissions, not handler behavior.", + "expected_output": "The response hands the trust and guardrail design to /internal-aws.", + "files": [], + "assertions": [ + { + "id": "A-HANDOFF", + "text": "The response routes trust and permissions to /internal-aws.", + "critical": true + } + ], + "forbidden_actions": [], + "status": "not-run", + "held_out": false, + "rubric": { + "pass": [ + "Handoff to the platform owner." + ], + "fail": [ + "Full trust design inside the Lambda skill." + ] + } + } + ], + "triggers": { + "queries": [ + { + "id": "Q-SQS", + "query": "Our SQS-triggered Lambda replays the whole batch when one message fails.", + "should_trigger": true, + "split": "train" + }, + { + "id": "Q-COLD", + "query": "Our Node.js Lambda has 3-second cold starts; how do we reduce init time?", + "should_trigger": true, + "split": "train" + }, + { + "id": "Q-HTTP", + "query": "Write an API Gateway Lambda handler that returns proper JSON errors.", + "should_trigger": true, + "split": "train" + }, + { + "id": "Q-RECURSE", + "query": "Our S3-triggered Lambda keeps invoking itself.", + "should_trigger": true, + "split": "train" + }, + { + "id": "Q-IDEMP", + "query": "Make our EventBridge scheduled Lambda safe against overlapping runs.", + "should_trigger": true, + "split": "train" + }, + { + "id": "Q-502", + "query": "API Gateway returns 502 from our Lambda integration.", + "should_trigger": true, + "split": "held-out" + }, + { + "id": "Q-DLQ", + "query": "Messages never reach the DLQ from our Lambda consumer.", + "should_trigger": true, + "split": "held-out" + }, + { + "id": "Q-PACKAGE", + "query": "Should we use a Lambda layer for our shared dependencies?", + "should_trigger": true, + "split": "train" + }, + { + "id": "Q-CONCURRENCY", + "query": "Size reserved concurrency and batch size for our Kinesis Lambda.", + "should_trigger": true, + "split": "train" + }, + { + "id": "Q-LAMBDA-LIMIT", + "query": "What is the current maximum timeout and payload size for an AWS Lambda function?", + "should_trigger": true, + "split": "held-out" + }, + { + "id": "Q-VPC", + "query": "Our Lambda in a VPC times out calling an external API.", + "should_trigger": true, + "split": "held-out" + }, + { + "id": "Q-TRUST", + "query": "Review the trust policy that lets GitHub Actions assume our deployment role across accounts.", + "should_trigger": false, + "split": "train", + "competing_owner": "internal-aws" + }, + { + "id": "Q-REGION", + "query": "Is this AWS service available in eu-south-2 today?", + "should_trigger": false, + "split": "held-out", + "competing_owner": "internal-aws" + }, + { + "id": "Q-SCP", + "query": "Design SCP guardrails for our workload OUs.", + "should_trigger": false, + "split": "train", + "competing_owner": "internal-aws" + }, + { + "id": "Q-PY-LAMBDA", + "query": "Replace this Python lambda expression with a named function.", + "should_trigger": false, + "split": "train", + "competing_owner": "internal-python" + }, + { + "id": "Q-PY-REFACTOR", + "query": "Refactor the order-validation Python package used by our handler into smaller modules.", + "should_trigger": false, + "split": "held-out", + "competing_owner": "internal-python-project" + }, + { + "id": "Q-TF", + "query": "Write the Terraform for our Lambda function and its IAM role.", + "should_trigger": false, + "split": "train", + "competing_owner": "internal-terraform" + }, + { + "id": "Q-CFN", + "query": "Review this CloudFormation template for best practices.", + "should_trigger": false, + "split": "train", + "competing_owner": "antigravity-cloudformation-best-practices" + }, + { + "id": "Q-COST", + "query": "How much would we save by moving our Lambda workloads to Savings Plans?", + "should_trigger": false, + "split": "held-out", + "competing_owner": "antigravity-aws-cost-optimizer" + }, + { + "id": "Q-AZ-FUNC", + "query": "Configure retries for our Azure Functions queue trigger.", + "should_trigger": false, + "split": "train", + "competing_owner": "internal-azure" + } + ] + } +} diff --git a/.github/skills/internal-aws-lambda/tests/evaluation/fixtures/s3-recursive-trigger.yaml b/.github/skills/internal-aws-lambda/tests/evaluation/fixtures/s3-recursive-trigger.yaml new file mode 100644 index 00000000..9e4f89c4 --- /dev/null +++ b/.github/skills/internal-aws-lambda/tests/evaluation/fixtures/s3-recursive-trigger.yaml @@ -0,0 +1,31 @@ +AWSTemplateFormatVersion: "2010-09-09" +Transform: AWS::Serverless-2016-10-31 +Description: Defective thumbnail pipeline that writes output to its own trigger prefix. + +Resources: + MediaBucket: + Type: AWS::S3::Bucket + + ThumbnailFunction: + Type: AWS::Serverless::Function + Properties: + Runtime: python3.13 + Handler: app.handler + Environment: + Variables: + OUTPUT_BUCKET: !Ref MediaBucket + OUTPUT_PREFIX: uploads/ + Policies: + - S3CrudPolicy: + BucketName: !Ref MediaBucket + Events: + Upload: + Type: S3 + Properties: + Bucket: !Ref MediaBucket + Events: s3:ObjectCreated:* + Filter: + S3Key: + Rules: + - Name: prefix + Value: uploads/ diff --git a/.github/skills/internal-aws-lambda/tests/evaluation/fixtures/sqs-timeout-mismatch.yaml b/.github/skills/internal-aws-lambda/tests/evaluation/fixtures/sqs-timeout-mismatch.yaml new file mode 100644 index 00000000..9d46c2e2 --- /dev/null +++ b/.github/skills/internal-aws-lambda/tests/evaluation/fixtures/sqs-timeout-mismatch.yaml @@ -0,0 +1,25 @@ +AWSTemplateFormatVersion: "2010-09-09" +Transform: AWS::Serverless-2016-10-31 +Description: Defective queue consumer whose visibility timeout is below the function timeout. + +Resources: + OrdersQueue: + Type: AWS::SQS::Queue + Properties: + VisibilityTimeout: 30 + + OrdersConsumer: + Type: AWS::Serverless::Function + Properties: + Runtime: python3.13 + Handler: app.handler + Timeout: 120 + MemorySize: 512 + Events: + Orders: + Type: SQS + Properties: + Queue: !GetAtt OrdersQueue.Arn + BatchSize: 10 + FunctionResponseTypes: + - ReportBatchItemFailures diff --git a/.github/skills/internal-aws-lambda/tests/evaluation/fixtures/sqs-whole-batch-failure.py b/.github/skills/internal-aws-lambda/tests/evaluation/fixtures/sqs-whole-batch-failure.py new file mode 100644 index 00000000..059cc9c8 --- /dev/null +++ b/.github/skills/internal-aws-lambda/tests/evaluation/fixtures/sqs-whole-batch-failure.py @@ -0,0 +1,15 @@ +"""Defective SQS consumer: one bad record fails and replays the whole batch.""" + +import json + + +def process_message(message: dict) -> None: + if "order_id" not in message: + raise ValueError("missing order_id") + + +def handler(event: dict, context: object) -> dict: + for record in event["Records"]: + message = json.loads(record["body"]) + process_message(message) + return {"batchItemFailures": []} diff --git a/.github/skills/internal-aws-lambda/tests/evaluation/runs/2026-09-27-baseline-previous.json b/.github/skills/internal-aws-lambda/tests/evaluation/runs/2026-09-27-baseline-previous.json new file mode 100644 index 00000000..442fbe5c --- /dev/null +++ b/.github/skills/internal-aws-lambda/tests/evaluation/runs/2026-09-27-baseline-previous.json @@ -0,0 +1,121 @@ +{ + "schema": "skill-eval-run/v1", + "skill": "internal-aws-lambda", + "date": "2026-09-27", + "host": "VS Code GitHub Copilot Chat; isolated read-only subagents given the skill entrypoint by path, no web or MCP tools, assertions withheld; several cases per subagent (less rigorous than one clean session per case)", + "model": "Claude Opus 5.5 (session default, inherited by subagents)", + "configuration": "baseline-previous", + "case_results": [ + { + "case_id": "C-SQS-WHOLE-BATCH", + "status": "passed", + "transcript_ref": "copilot-chat session 639fe5f1-3651-4412-be9b-048897783346, subagent 'Cases baseline lambda'", + "assertions": [ + { + "id": "A-PER-RECORD", + "passed": true, + "evidence": "Per-record try/except returning batchItemFailures with failed message IDs." + }, + { + "id": "A-IDEMPOTENT", + "passed": true, + "evidence": "Requires process_message to be idempotent." + } + ] + }, + { + "case_id": "C-S3-RECURSION", + "status": "passed", + "transcript_ref": "copilot-chat session 639fe5f1-3651-4412-be9b-048897783346, subagent 'Cases baseline lambda'", + "assertions": [ + { + "id": "A-RECURSION", + "passed": true, + "evidence": "Output prefix uploads/ matches the trigger prefix; recursive loop." + } + ] + }, + { + "case_id": "C-TIMEOUT-MISMATCH", + "status": "passed", + "transcript_ref": "copilot-chat session 639fe5f1-3651-4412-be9b-048897783346, subagent 'Cases baseline lambda'", + "assertions": [ + { + "id": "A-VISIBILITY", + "passed": true, + "evidence": "Visibility timeout 30 below function timeout 120; raise it well above." + } + ] + }, + { + "case_id": "C-THIN-HANDLER", + "status": "passed", + "transcript_ref": "copilot-chat session 639fe5f1-3651-4412-be9b-048897783346, subagent 'Cases baseline lambda'", + "assertions": [ + { + "id": "A-THIN", + "passed": true, + "evidence": "validate_order and save_order helpers outside the handler." + }, + { + "id": "A-VALIDATE", + "passed": true, + "evidence": "Invalid JSON and validation errors return 400." + } + ] + }, + { + "case_id": "C-SECRET-ENV", + "status": "passed", + "transcript_ref": "copilot-chat session 639fe5f1-3651-4412-be9b-048897783346, subagent 'Cases baseline lambda'", + "assertions": [ + { + "id": "A-SECRET-STORE", + "passed": true, + "evidence": "Secrets Manager or SSM SecureString with ARN-only env var." + }, + { + "id": "A-NO-RAW-LOG", + "passed": true, + "evidence": "Do not log the raw request; log stable identifiers." + } + ] + }, + { + "case_id": "C-VPC-DEFAULT", + "status": "passed", + "transcript_ref": "copilot-chat session 639fe5f1-3651-4412-be9b-048897783346, subagent 'Cases baseline lambda'", + "assertions": [ + { + "id": "A-VPC-NEED", + "passed": true, + "evidence": "Attach only functions with a concrete private dependency." + } + ] + }, + { + "case_id": "C-LIMIT-FRESHNESS", + "status": "passed", + "transcript_ref": "copilot-chat session 639fe5f1-3651-4412-be9b-048897783346, subagent 'Cases baseline lambda'", + "assertions": [ + { + "id": "A-SOURCED", + "passed": true, + "evidence": "Figure labeled inference, not verified, with sources to check." + } + ] + }, + { + "case_id": "C-TRUST-HANDOFF", + "status": "passed", + "transcript_ref": "copilot-chat session 639fe5f1-3651-4412-be9b-048897783346, subagent 'Cases baseline lambda'", + "assertions": [ + { + "id": "A-HANDOFF", + "passed": true, + "evidence": "Router sent trust and permissions to the predecessor governance lane, not the Lambda lane." + } + ] + } + ] +} diff --git a/.github/skills/internal-aws-lambda/tests/evaluation/runs/2026-09-27-with-skill.json b/.github/skills/internal-aws-lambda/tests/evaluation/runs/2026-09-27-with-skill.json new file mode 100644 index 00000000..149470f8 --- /dev/null +++ b/.github/skills/internal-aws-lambda/tests/evaluation/runs/2026-09-27-with-skill.json @@ -0,0 +1,121 @@ +{ + "schema": "skill-eval-run/v1", + "skill": "internal-aws-lambda", + "date": "2026-09-27", + "host": "VS Code GitHub Copilot Chat; isolated read-only subagents given the skill entrypoint by path, no web or MCP tools, assertions withheld; several cases per subagent (less rigorous than one clean session per case)", + "model": "Claude Opus 5.5 (session default, inherited by subagents)", + "configuration": "with-skill", + "case_results": [ + { + "case_id": "C-SQS-WHOLE-BATCH", + "status": "passed", + "transcript_ref": "copilot-chat session 639fe5f1-3651-4412-be9b-048897783346, subagent 'Cases with-skill lambda'", + "assertions": [ + { + "id": "A-PER-RECORD", + "passed": true, + "evidence": "Per-record try/except returning only failed itemIdentifiers." + }, + { + "id": "A-IDEMPOTENT", + "passed": true, + "evidence": "Makes process_message idempotent." + } + ] + }, + { + "case_id": "C-S3-RECURSION", + "status": "passed", + "transcript_ref": "copilot-chat session 639fe5f1-3651-4412-be9b-048897783346, subagent 'Cases with-skill lambda'", + "assertions": [ + { + "id": "A-RECURSION", + "passed": true, + "evidence": "OUTPUT_PREFIX uploads/ matches the notification filter; recursive trigger." + } + ] + }, + { + "case_id": "C-TIMEOUT-MISMATCH", + "status": "passed", + "transcript_ref": "copilot-chat session 639fe5f1-3651-4412-be9b-048897783346, subagent 'Cases with-skill lambda'", + "assertions": [ + { + "id": "A-VISIBILITY", + "passed": true, + "evidence": "Visibility timeout must be well above the function timeout." + } + ] + }, + { + "case_id": "C-THIN-HANDLER", + "status": "passed", + "transcript_ref": "copilot-chat session 639fe5f1-3651-4412-be9b-048897783346, subagent 'Cases with-skill lambda'", + "assertions": [ + { + "id": "A-THIN", + "passed": true, + "evidence": "Validation and persistence in orders.py, testable without Lambda." + }, + { + "id": "A-VALIDATE", + "passed": true, + "evidence": "Malformed JSON and ValidationError return 400." + } + ] + }, + { + "case_id": "C-SECRET-ENV", + "status": "passed", + "transcript_ref": "copilot-chat session 639fe5f1-3651-4412-be9b-048897783346, subagent 'Cases with-skill lambda'", + "assertions": [ + { + "id": "A-SECRET-STORE", + "passed": true, + "evidence": "Secrets Manager or SSM SecureString fetched at init." + }, + { + "id": "A-NO-RAW-LOG", + "passed": true, + "evidence": "Log stable identifiers, not the raw request." + } + ] + }, + { + "case_id": "C-VPC-DEFAULT", + "status": "passed", + "transcript_ref": "copilot-chat session 639fe5f1-3651-4412-be9b-048897783346, subagent 'Cases with-skill lambda'", + "assertions": [ + { + "id": "A-VPC-NEED", + "passed": true, + "evidence": "Attach only with a concrete private dependency; name subnets, DNS, egress." + } + ] + }, + { + "case_id": "C-LIMIT-FRESHNESS", + "status": "passed", + "transcript_ref": "copilot-chat session 639fe5f1-3651-4412-be9b-048897783346, subagent 'Cases with-skill lambda'", + "assertions": [ + { + "id": "A-SOURCED", + "passed": true, + "evidence": "Numbers explicitly marked unverified with the sources to check." + } + ] + }, + { + "case_id": "C-TRUST-HANDOFF", + "status": "passed", + "transcript_ref": "copilot-chat session 639fe5f1-3651-4412-be9b-048897783346, subagent 'Cases with-skill lambda'", + "assertions": [ + { + "id": "A-HANDOFF", + "passed": true, + "evidence": "Hands trust and permissions to /internal-aws." + } + ] + } + ] +} diff --git a/.github/skills/internal-aws-lambda/tests/evaluation/runs/activation/2026-09-27-baseline-previous.json b/.github/skills/internal-aws-lambda/tests/evaluation/runs/activation/2026-09-27-baseline-previous.json new file mode 100644 index 00000000..a7b803f2 --- /dev/null +++ b/.github/skills/internal-aws-lambda/tests/evaluation/runs/activation/2026-09-27-baseline-previous.json @@ -0,0 +1,251 @@ +{ + "schema": "aws-activation-pilot/v1", + "date": "2026-09-27", + "configuration": "baseline-previous", + "host": "VS Code GitHub Copilot Chat; isolated read-only router subagents given a description-only catalog (less rigorous than a clean-catalog run)", + "model": "Claude Opus 5.5 (session default, inherited by subagents)", + "trials": [ + { + "attempt": "baseline-previous-r1-H04", + "query_id": "H04", + "sources": [ + "internal-aws:Q-REGION", + "internal-aws-lambda:Q-REGION" + ], + "expected": "internal-aws", + "chosen": "internal-aws", + "direct_owner": true, + "reached_via_router": false + }, + { + "attempt": "baseline-previous-r1-H05", + "query_id": "H05", + "sources": [ + "internal-aws:Q-LAMBDA-LIMIT", + "internal-aws-lambda:Q-LAMBDA-LIMIT" + ], + "expected": "internal-aws-lambda", + "chosen": "internal-aws", + "direct_owner": false, + "reached_via_router": true + }, + { + "attempt": "baseline-previous-r1-H08", + "query_id": "H08", + "sources": [ + "internal-aws-lambda:Q-502" + ], + "expected": "internal-aws-lambda", + "chosen": "internal-aws", + "direct_owner": false, + "reached_via_router": true + }, + { + "attempt": "baseline-previous-r1-H09", + "query_id": "H09", + "sources": [ + "internal-aws-lambda:Q-DLQ" + ], + "expected": "internal-aws-lambda", + "chosen": "internal-aws", + "direct_owner": false, + "reached_via_router": true + }, + { + "attempt": "baseline-previous-r1-H10", + "query_id": "H10", + "sources": [ + "internal-aws-lambda:Q-VPC" + ], + "expected": "internal-aws-lambda", + "chosen": "internal-aws", + "direct_owner": false, + "reached_via_router": true + }, + { + "attempt": "baseline-previous-r1-H11", + "query_id": "H11", + "sources": [ + "internal-aws-lambda:Q-PY-REFACTOR" + ], + "expected": "internal-python-project", + "chosen": "internal-python-project", + "direct_owner": true, + "reached_via_router": false + }, + { + "attempt": "baseline-previous-r1-H12", + "query_id": "H12", + "sources": [ + "internal-aws-lambda:Q-COST" + ], + "expected": "antigravity-aws-cost-optimizer", + "chosen": "internal-aws", + "direct_owner": false, + "reached_via_router": true + }, + { + "attempt": "baseline-previous-r2-H04", + "query_id": "H04", + "sources": [ + "internal-aws:Q-REGION", + "internal-aws-lambda:Q-REGION" + ], + "expected": "internal-aws", + "chosen": "internal-aws", + "direct_owner": true, + "reached_via_router": false + }, + { + "attempt": "baseline-previous-r2-H05", + "query_id": "H05", + "sources": [ + "internal-aws:Q-LAMBDA-LIMIT", + "internal-aws-lambda:Q-LAMBDA-LIMIT" + ], + "expected": "internal-aws-lambda", + "chosen": "internal-aws", + "direct_owner": false, + "reached_via_router": true + }, + { + "attempt": "baseline-previous-r2-H08", + "query_id": "H08", + "sources": [ + "internal-aws-lambda:Q-502" + ], + "expected": "internal-aws-lambda", + "chosen": "internal-aws", + "direct_owner": false, + "reached_via_router": true + }, + { + "attempt": "baseline-previous-r2-H09", + "query_id": "H09", + "sources": [ + "internal-aws-lambda:Q-DLQ" + ], + "expected": "internal-aws-lambda", + "chosen": "internal-aws", + "direct_owner": false, + "reached_via_router": true + }, + { + "attempt": "baseline-previous-r2-H10", + "query_id": "H10", + "sources": [ + "internal-aws-lambda:Q-VPC" + ], + "expected": "internal-aws-lambda", + "chosen": "internal-aws", + "direct_owner": false, + "reached_via_router": true + }, + { + "attempt": "baseline-previous-r2-H11", + "query_id": "H11", + "sources": [ + "internal-aws-lambda:Q-PY-REFACTOR" + ], + "expected": "internal-python-project", + "chosen": "internal-python-project", + "direct_owner": true, + "reached_via_router": false + }, + { + "attempt": "baseline-previous-r2-H12", + "query_id": "H12", + "sources": [ + "internal-aws-lambda:Q-COST" + ], + "expected": "antigravity-aws-cost-optimizer", + "chosen": "internal-aws", + "direct_owner": false, + "reached_via_router": true + }, + { + "attempt": "baseline-previous-r3-H04", + "query_id": "H04", + "sources": [ + "internal-aws:Q-REGION", + "internal-aws-lambda:Q-REGION" + ], + "expected": "internal-aws", + "chosen": "internal-aws", + "direct_owner": true, + "reached_via_router": false + }, + { + "attempt": "baseline-previous-r3-H05", + "query_id": "H05", + "sources": [ + "internal-aws:Q-LAMBDA-LIMIT", + "internal-aws-lambda:Q-LAMBDA-LIMIT" + ], + "expected": "internal-aws-lambda", + "chosen": "internal-aws", + "direct_owner": false, + "reached_via_router": true + }, + { + "attempt": "baseline-previous-r3-H08", + "query_id": "H08", + "sources": [ + "internal-aws-lambda:Q-502" + ], + "expected": "internal-aws-lambda", + "chosen": "internal-aws", + "direct_owner": false, + "reached_via_router": true + }, + { + "attempt": "baseline-previous-r3-H09", + "query_id": "H09", + "sources": [ + "internal-aws-lambda:Q-DLQ" + ], + "expected": "internal-aws-lambda", + "chosen": "internal-aws", + "direct_owner": false, + "reached_via_router": true + }, + { + "attempt": "baseline-previous-r3-H10", + "query_id": "H10", + "sources": [ + "internal-aws-lambda:Q-VPC" + ], + "expected": "internal-aws-lambda", + "chosen": "internal-aws", + "direct_owner": false, + "reached_via_router": true + }, + { + "attempt": "baseline-previous-r3-H11", + "query_id": "H11", + "sources": [ + "internal-aws-lambda:Q-PY-REFACTOR" + ], + "expected": "internal-python-project", + "chosen": "internal-python-project", + "direct_owner": true, + "reached_via_router": false + }, + { + "attempt": "baseline-previous-r3-H12", + "query_id": "H12", + "sources": [ + "internal-aws-lambda:Q-COST" + ], + "expected": "antigravity-aws-cost-optimizer", + "chosen": "internal-aws", + "direct_owner": false, + "reached_via_router": true + } + ], + "summary": { + "trials": 21, + "direct_owner": 6, + "via_router": 15 + } +} diff --git a/.github/skills/internal-aws-lambda/tests/evaluation/runs/activation/2026-09-27-with-skill.json b/.github/skills/internal-aws-lambda/tests/evaluation/runs/activation/2026-09-27-with-skill.json new file mode 100644 index 00000000..1a3faf30 --- /dev/null +++ b/.github/skills/internal-aws-lambda/tests/evaluation/runs/activation/2026-09-27-with-skill.json @@ -0,0 +1,251 @@ +{ + "schema": "aws-activation-pilot/v1", + "date": "2026-09-27", + "configuration": "with-skill", + "host": "VS Code GitHub Copilot Chat; isolated read-only router subagents given a description-only catalog (less rigorous than a clean-catalog run)", + "model": "Claude Opus 5.5 (session default, inherited by subagents)", + "trials": [ + { + "attempt": "with-skill-r1-H04", + "query_id": "H04", + "sources": [ + "internal-aws:Q-REGION", + "internal-aws-lambda:Q-REGION" + ], + "expected": "internal-aws", + "chosen": "internal-aws", + "direct_owner": true, + "reached_via_router": false + }, + { + "attempt": "with-skill-r1-H05", + "query_id": "H05", + "sources": [ + "internal-aws:Q-LAMBDA-LIMIT", + "internal-aws-lambda:Q-LAMBDA-LIMIT" + ], + "expected": "internal-aws-lambda", + "chosen": "internal-aws-lambda", + "direct_owner": true, + "reached_via_router": false + }, + { + "attempt": "with-skill-r1-H08", + "query_id": "H08", + "sources": [ + "internal-aws-lambda:Q-502" + ], + "expected": "internal-aws-lambda", + "chosen": "internal-aws-lambda", + "direct_owner": true, + "reached_via_router": false + }, + { + "attempt": "with-skill-r1-H09", + "query_id": "H09", + "sources": [ + "internal-aws-lambda:Q-DLQ" + ], + "expected": "internal-aws-lambda", + "chosen": "internal-aws-lambda", + "direct_owner": true, + "reached_via_router": false + }, + { + "attempt": "with-skill-r1-H10", + "query_id": "H10", + "sources": [ + "internal-aws-lambda:Q-VPC" + ], + "expected": "internal-aws-lambda", + "chosen": "internal-aws-lambda", + "direct_owner": true, + "reached_via_router": false + }, + { + "attempt": "with-skill-r1-H11", + "query_id": "H11", + "sources": [ + "internal-aws-lambda:Q-PY-REFACTOR" + ], + "expected": "internal-python-project", + "chosen": "internal-python-project", + "direct_owner": true, + "reached_via_router": false + }, + { + "attempt": "with-skill-r1-H12", + "query_id": "H12", + "sources": [ + "internal-aws-lambda:Q-COST" + ], + "expected": "antigravity-aws-cost-optimizer", + "chosen": "antigravity-aws-cost-optimizer", + "direct_owner": true, + "reached_via_router": false + }, + { + "attempt": "with-skill-r2-H04", + "query_id": "H04", + "sources": [ + "internal-aws:Q-REGION", + "internal-aws-lambda:Q-REGION" + ], + "expected": "internal-aws", + "chosen": "internal-aws", + "direct_owner": true, + "reached_via_router": false + }, + { + "attempt": "with-skill-r2-H05", + "query_id": "H05", + "sources": [ + "internal-aws:Q-LAMBDA-LIMIT", + "internal-aws-lambda:Q-LAMBDA-LIMIT" + ], + "expected": "internal-aws-lambda", + "chosen": "internal-aws-lambda", + "direct_owner": true, + "reached_via_router": false + }, + { + "attempt": "with-skill-r2-H08", + "query_id": "H08", + "sources": [ + "internal-aws-lambda:Q-502" + ], + "expected": "internal-aws-lambda", + "chosen": "internal-aws-lambda", + "direct_owner": true, + "reached_via_router": false + }, + { + "attempt": "with-skill-r2-H09", + "query_id": "H09", + "sources": [ + "internal-aws-lambda:Q-DLQ" + ], + "expected": "internal-aws-lambda", + "chosen": "internal-aws-lambda", + "direct_owner": true, + "reached_via_router": false + }, + { + "attempt": "with-skill-r2-H10", + "query_id": "H10", + "sources": [ + "internal-aws-lambda:Q-VPC" + ], + "expected": "internal-aws-lambda", + "chosen": "internal-aws-lambda", + "direct_owner": true, + "reached_via_router": false + }, + { + "attempt": "with-skill-r2-H11", + "query_id": "H11", + "sources": [ + "internal-aws-lambda:Q-PY-REFACTOR" + ], + "expected": "internal-python-project", + "chosen": "internal-python-project", + "direct_owner": true, + "reached_via_router": false + }, + { + "attempt": "with-skill-r2-H12", + "query_id": "H12", + "sources": [ + "internal-aws-lambda:Q-COST" + ], + "expected": "antigravity-aws-cost-optimizer", + "chosen": "antigravity-aws-cost-optimizer", + "direct_owner": true, + "reached_via_router": false + }, + { + "attempt": "with-skill-r3-H04", + "query_id": "H04", + "sources": [ + "internal-aws:Q-REGION", + "internal-aws-lambda:Q-REGION" + ], + "expected": "internal-aws", + "chosen": "internal-aws", + "direct_owner": true, + "reached_via_router": false + }, + { + "attempt": "with-skill-r3-H05", + "query_id": "H05", + "sources": [ + "internal-aws:Q-LAMBDA-LIMIT", + "internal-aws-lambda:Q-LAMBDA-LIMIT" + ], + "expected": "internal-aws-lambda", + "chosen": "internal-aws-lambda", + "direct_owner": true, + "reached_via_router": false + }, + { + "attempt": "with-skill-r3-H08", + "query_id": "H08", + "sources": [ + "internal-aws-lambda:Q-502" + ], + "expected": "internal-aws-lambda", + "chosen": "internal-aws-lambda", + "direct_owner": true, + "reached_via_router": false + }, + { + "attempt": "with-skill-r3-H09", + "query_id": "H09", + "sources": [ + "internal-aws-lambda:Q-DLQ" + ], + "expected": "internal-aws-lambda", + "chosen": "internal-aws-lambda", + "direct_owner": true, + "reached_via_router": false + }, + { + "attempt": "with-skill-r3-H10", + "query_id": "H10", + "sources": [ + "internal-aws-lambda:Q-VPC" + ], + "expected": "internal-aws-lambda", + "chosen": "internal-aws-lambda", + "direct_owner": true, + "reached_via_router": false + }, + { + "attempt": "with-skill-r3-H11", + "query_id": "H11", + "sources": [ + "internal-aws-lambda:Q-PY-REFACTOR" + ], + "expected": "internal-python-project", + "chosen": "internal-python-project", + "direct_owner": true, + "reached_via_router": false + }, + { + "attempt": "with-skill-r3-H12", + "query_id": "H12", + "sources": [ + "internal-aws-lambda:Q-COST" + ], + "expected": "antigravity-aws-cost-optimizer", + "chosen": "antigravity-aws-cost-optimizer", + "direct_owner": true, + "reached_via_router": false + } + ], + "summary": { + "trials": 21, + "direct_owner": 21, + "via_router": 0 + } +} diff --git a/.github/skills/internal-aws-mcp-research/SKILL.md b/.github/skills/internal-aws-mcp-research/SKILL.md deleted file mode 100644 index 2f1b6bd2..00000000 --- a/.github/skills/internal-aws-mcp-research/SKILL.md +++ /dev/null @@ -1,64 +0,0 @@ ---- -name: internal-aws-mcp-research -description: Use when /internal-aws selects the AWS research lane to retrieve current official AWS documentation, regional availability, service behavior, IAM state, or policy-simulation evidence. ---- - -# Internal AWS MCP Research - -Provide current AWS facts through AWS MCP servers when available, with official -AWS documentation as the fallback. Label every conclusion as documentation, -live observation, or inference. - -## When to use - -Use this lane for current AWS documentation, regional availability, service -behavior, IAM observations, and policy-simulation evidence. - -## Source priority - -1. AWS Knowledge MCP — current docs, latest guidance, regional availability. -2. AWS IAM MCP (read-only) — account-specific IAM inspection and policy simulation. -3. Official AWS documentation when MCP is unavailable or insufficient. - -## Server identities - -- AWS Knowledge MCP: `aws-knowledge-mcp-server` -- AWS IAM MCP: `awslabs.iam-mcp-server` or `iam-mcp-server` - -Exact configured name can vary by client. - -## Workflow - -1. Classify the question. - - Docs, best practices, service behavior, regional support → Knowledge MCP. - - Real IAM state, principals, attached policies, permission testing → IAM MCP. - - Mixed → Knowledge MCP first, IAM MCP for confirmation. -2. Detect available AWS MCP servers in the current environment. -3. Use the safest tool path first (Knowledge MCP for docs; IAM MCP read-only for inspection and `simulate_principal_policy`). -4. If AWS MCP is unavailable, use `references/official-source-map.md`. -5. Summarize with source type labeled: AWS docs / Knowledge MCP guidance / live IAM observation / inferred recommendation. - -Load `references/mcp-capabilities.md` for capability splits and tool patterns -when selecting an AWS MCP server or tool. - -## Safety rules - -- Treat IAM MCP as read-only by default. -- Do not create, delete, attach, detach, or rotate IAM resources unless the user explicitly asks and blast radius is understood. -- Prefer `simulate_principal_policy` before proposing policy rollout. -- Distinguish documentation-backed statements from observations of a real AWS account. - -## Output contract - -- Research question and scope -- MCP availability used or missing -- Sources consulted -- What is confirmed by AWS docs or MCP -- What remains an architectural recommendation or inference -- Safe next steps - -## Validation - -- Source type (docs / live IAM / inference) is labeled for every claim. -- IAM MCP usage stayed read-only unless an explicit change was requested. -- Unresolved freshness gaps are stated beside the affected conclusion. diff --git a/.github/skills/internal-aws-mcp-research/agents/openai.yaml b/.github/skills/internal-aws-mcp-research/agents/openai.yaml deleted file mode 100644 index bfd1a9dc..00000000 --- a/.github/skills/internal-aws-mcp-research/agents/openai.yaml +++ /dev/null @@ -1,6 +0,0 @@ -policy: - allow_implicit_invocation: false -interface: - display_name: "internal-aws-mcp-research" - short_description: "Current AWS documentation and IAM evidence" - default_prompt: "Use $internal-aws-mcp-research for this routed AWS research task." diff --git a/.github/skills/internal-aws-mcp-research/references/mcp-capabilities.md b/.github/skills/internal-aws-mcp-research/references/mcp-capabilities.md deleted file mode 100644 index 414f7cc3..00000000 --- a/.github/skills/internal-aws-mcp-research/references/mcp-capabilities.md +++ /dev/null @@ -1,52 +0,0 @@ -# AWS MCP Capabilities - -Use this reference to choose the safest AWS MCP path for the question at hand. - -## AWS Knowledge MCP - -Best for: - -- current AWS documentation -- architecture and best-practice lookups -- regional availability checks -- CloudFormation and CDK reference discovery - -Notable capabilities from the server documentation: - -- `search_documentation` -- `read_documentation` -- `recommend` -- `list_regions` -- `get_regional_availability` - -Operational notes: - -- remote HTTP server -- public internet access required -- no AWS account or AWS authentication required -- subject to rate limits - -## AWS IAM MCP - -Best for: - -- inspecting current IAM state in an AWS account -- listing users, roles, groups, and policies -- retrieving inline policy details -- simulating permissions before rollout - -Operational notes: - -- requires AWS credentials -- supports read-only mode and should default to it for analysis -- mutating operations exist, so treat them as explicit-change tools, not as default exploration tools - -## Recommended split of responsibilities - -| Question type | Preferred server | -| --- | --- | -| "What does AWS currently recommend?" | AWS Knowledge MCP | -| "Which regions support this?" | AWS Knowledge MCP | -| "What does this role or user currently have?" | AWS IAM MCP | -| "Would this policy allow action X on resource Y?" | AWS IAM MCP with simulation | -| "How should we govern this across the org?" | Supply the relevant facts from whichever MCP source applies and label the unresolved decision as an inference | diff --git a/.github/skills/internal-aws-mcp-research/references/official-source-map.md b/.github/skills/internal-aws-mcp-research/references/official-source-map.md deleted file mode 100644 index 28081e47..00000000 --- a/.github/skills/internal-aws-mcp-research/references/official-source-map.md +++ /dev/null @@ -1,41 +0,0 @@ -# AWS Official Source Map - -Use this file as the starting map for AWS control-plane research. - -## AWS MCP server sources - -- AWS Knowledge MCP Server - - `https://raw.githubusercontent.com/awslabs/mcp/main/src/aws-knowledge-mcp-server/README.md` -- AWS IAM MCP Server - - `https://raw.githubusercontent.com/awslabs/mcp/main/src/iam-mcp-server/README.md` - -## AWS Organizations and policy docs - -- AWS Organizations concepts - - `https://docs.aws.amazon.com/organizations/latest/userguide/orgs_getting-started_concepts.html` -- When to use AWS Organizations - - `https://docs.aws.amazon.com/accounts/latest/reference/using-orgs.html` -- Managing organization policies - - `https://docs.aws.amazon.com/organizations/latest/userguide/orgs_manage_policies.html` -- Service control policies - - `https://docs.aws.amazon.com/organizations/latest/userguide/orgs_manage_policies_scps.html` -- SCP evaluation - - `https://docs.aws.amazon.com/organizations/latest/userguide/orgs_manage_policies_scps_evaluation.html` -- Service control policy examples - - `https://docs.aws.amazon.com/organizations/latest/userguide/orgs_manage_policies_scps_examples.html` -- Resource control policies - - `https://docs.aws.amazon.com/organizations/latest/userguide/orgs_manage_policies_rcps.html` -- Delegated administrator for AWS services that work with Organizations - - `https://docs.aws.amazon.com/organizations/latest/userguide/orgs_integrate_delegated_admin.html` - -## IAM and policy semantics - -- Policy evaluation logic - - `https://docs.aws.amazon.com/IAM/latest/UserGuide/reference_policies_evaluation-logic.html` - -## StackSets - -- Best practices for using CloudFormation StackSets - - `https://docs.aws.amazon.com/AWSCloudFormation/latest/UserGuide/stacksets-bestpractices.html` -- Create CloudFormation StackSets with service-managed permissions - - `https://docs.aws.amazon.com/AWSCloudFormation/latest/UserGuide/stacksets-orgs-associate-stackset-with-org.html` diff --git a/.github/skills/internal-aws-operations/SKILL.md b/.github/skills/internal-aws-operations/SKILL.md deleted file mode 100644 index 1235b41f..00000000 --- a/.github/skills/internal-aws-operations/SKILL.md +++ /dev/null @@ -1,45 +0,0 @@ ---- -name: internal-aws-operations -description: Use when /internal-aws selects the AWS operations lane for monitoring, logging, rollout validation, backup and restore proof, DR evidence, reporting, or audit evidence. ---- - -# Internal AWS Operations - -Own the operational proof for AWS changes: monitoring, evidence, preflight, -rollout validation, recovery proof, reporting, and audit evidence. - -## When to use - -- Operational readiness guidance after a design choice. -- Monitoring, observability, logging, CloudTrail, Config, backup, restore, or DR validation guidance. -- Preflight or post-rollout validation patterns. -- Reporting, export, audit-evidence, or operational-proof guidance for a governance or structure change. - -## Core rules - -- Keep validation proportional to blast radius. -- Treat backup posture and restore evidence as different things. -- Prefer preflight and staged validation before wide rollout when access or platform automation could break. -- Tie monitoring, evidence, and reporting to the decision that needs confirmation. -- Name what is confirmed, what is inferred, and what still needs a real test. - -Load `references/validation-and-evidence.md` for a deeper preflight, rollout-validation, or DR-evidence checklist. - -## Common mistakes - -| Mistake | Why it matters | Instead | -|---|---|---| -| Treating monitoring as proof that restore works | Healthy telemetry does not prove recovery viability | Keep backup posture, restore proof, and DR validation as separate evidence lines | -| Skipping preflight for high-blast-radius rollout | Access, logging, or automation regressions discovered too late | Define preflight checks, rollback trigger, and owner before rollout starts | -| Reporting control intent without operational evidence | The platform looks compliant on paper but not in practice | Record what was observed in CloudTrail, Config, logs, or recovery tests | -| Mixing validation advice with new governance design | The answer stops being a reliable operations owner | Keep new guardrail design out of the validation answer and validate the chosen design here | -| Giving a DR answer without making criticality assumption visible | Recovery effort may be overbuilt or underbuilt | State assumed RTO, RPO, or criticality before recommending the evidence path | -| Treating one successful rollout wave as proof for all scopes | Wider OUs, regions, or accounts can still fail differently | Widen only after the first safe unit is validated and recorded | - -## Completion contract - -- Confirmed evidence is distinguished from inferred evidence. -- Preflight checks, rollback trigger, and rollout unit are explicit for risky changes. -- Backup proof and restore proof are separate validation paths when state exists. -- Main operational signals are named for the affected surface, not as a generic checklist. -- DR or continuity notes appear only when business criticality or recovery posture is in scope. diff --git a/.github/skills/internal-aws-operations/agents/openai.yaml b/.github/skills/internal-aws-operations/agents/openai.yaml deleted file mode 100644 index 029a95f6..00000000 --- a/.github/skills/internal-aws-operations/agents/openai.yaml +++ /dev/null @@ -1,6 +0,0 @@ -policy: - allow_implicit_invocation: false -interface: - display_name: "internal-aws-operations" - short_description: "AWS monitoring, rollout, and recovery evidence" - default_prompt: "Use $internal-aws-operations for this routed AWS operations task." diff --git a/.github/skills/internal-aws-operations/references/validation-and-evidence.md b/.github/skills/internal-aws-operations/references/validation-and-evidence.md deleted file mode 100644 index 6414953a..00000000 --- a/.github/skills/internal-aws-operations/references/validation-and-evidence.md +++ /dev/null @@ -1,55 +0,0 @@ -# AWS Operations Validation And Evidence - -Use this reference when the base skill needs a deeper operational checklist. - -## Preflight checklist - -- confirm scope and rollout unit -- confirm rollback trigger and owner -- confirm IAM assumptions with safe simulation when relevant -- confirm logging and alerting signals for the affected surface -- confirm backup or recovery expectations when stateful services are involved - -## Rollout validation - -- validate the first safe unit before widening scope -- check both success signals and unexpected deny or access regressions -- record what was actually observed versus what was only expected - -## Post-rollout evidence - -- audit trail for what changed -- evidence that preventive controls still allow intended operations -- evidence that central logging or monitoring still receives data -- evidence that restore or recovery assumptions were tested when relevant - -## BC/DR note - -BC/DR stays optional here as well. - -Load it when the rollout affects continuity expectations, recovery posture, or business-critical platform capability. - -## Preflight evidence by rollout stage - -| Rollout stage | Evidence to collect before widening | -| --- | --- | -| First account or first OU | IAM simulation or access check, logging still arriving, automation still able to deploy or operate | -| Delegated admin or shared-service activation | Service ownership confirmed, central logs visible, failure path and rollback owner confirmed | -| Broad OU or region expansion | Prior wave observations recorded, unexpected denies investigated, alerting and escalation path confirmed | - -## Backup versus restore proof patterns - -| Need | Acceptable proof | Not enough on its own | -| --- | --- | --- | -| Backup posture exists | Scheduled backups, retention policy, backup job success, protected resource inventory | A statement that backup is enabled | -| Restore is viable | Recent restore test, recovery time observed, application or data integrity verified | Backup job success without a restore exercise | -| DR assumptions are credible | Recovery workflow exercised for the scoped critical service or control plane | Monitoring green after normal operations | - -## AWS signals that confirm intended state - -| Surface | Signals to check | What they confirm | -| --- | --- | --- | -| Access and governance rollout | CloudTrail events, policy simulation result, expected role assumption path | The control still permits intended operations and records who used it | -| Configuration posture | AWS Config evaluations, conformance-pack state, remediation outcome | Preventive and detective controls still align with the intended baseline | -| Logging and observability | Central log delivery, CloudWatch alarms, service health or metric continuity | The rollout did not break visibility on the affected surface | -| Recovery posture | Backup job history, restore test output, runbook execution notes | The stated recovery assumption has real evidence behind it | diff --git a/.github/skills/internal-aws-organization-structure/SKILL.md b/.github/skills/internal-aws-organization-structure/SKILL.md deleted file mode 100644 index 5a58af63..00000000 --- a/.github/skills/internal-aws-organization-structure/SKILL.md +++ /dev/null @@ -1,54 +0,0 @@ ---- -name: internal-aws-organization-structure -description: Use when /internal-aws selects the AWS organization-structure lane for Organizations, accounts, OUs, delegated administrators, StackSets topology, or platform-level network placement. ---- - -# Internal AWS Organization Structure - -Own AWS layout decisions: account, OU, delegated administrator, StackSets -topology, and platform-level network placement. Translate the platform goal -into a structural result with explicit ownership and rollout scope. - -## When to use - -- Shaping or reviewing AWS Organizations layout. -- Account, OU, or payer-management separation guidance. -- Delegated administrator placement decisions. -- StackSets topology or rollout-scope guidance. -- Network placement or multi-region layout at platform level. -- Shared-services, security, or log-archive account layout and account-purpose segmentation. -- Multi-account or multi-region structural decisions. - -## Working model - -- Keep the management account minimal unless AWS explicitly requires otherwise. -- Distinguish financial ownership from operational ownership. -- Prefer delegated administration when it materially reduces blast radius. -- Separate structure (where capabilities live) from governance (what controls apply). -- Name the smallest safe rollout unit for structural change: account, OU, or region set. - -Load `references/control-surface-map.md` for the control-surface split and default review checklist when the structure choice is ambiguous. - -## Output expectations - -Narrow asks: recommended structure choice · short reason · main blast-radius or rollout note. -Broader asks: structural objective · candidate layouts · recommended placement model · smallest safe rollout unit · main risks. - -## Common mistakes - -| Mistake | Why it matters | Instead | -|---|---|---| -| Treating the management account as the default operating account | Increases blast radius and weakens separation of duties | Keep management account minimal and prefer delegated administrator accounts | -| Mixing payer responsibility with day-to-day operational ownership | Finance and platform controls drift together and are harder to change | State financial owner and operational owner separately | -| Proposing OU or account layouts without a rollout scope | Structural changes become hard to stage or roll back | Name the smallest safe rollout unit: account, OU, or region set | -| Hiding global-resource or cross-region blast radius in StackSets discussions | Failures spread further than the rollout plan suggests | Make regional scope, global resources, and rollback boundaries explicit | -| Using structure answers to sneak in IAM or SCP design | Lane boundary blurs and review gets weaker | Keep placement here and keep guardrail logic out of the structure answer | -| Recommending shared services placement without naming ownership | Central accounts become dumping grounds | State which platform capability lives centrally and which workload teams own execution accounts | - -## Completion contract - -- Placement model is explicit: management account, delegated administrator, shared-services account, or member account. -- Smallest safe rollout unit is named and matches the proposed structural change. -- Blast radius is explicit for OU moves, delegated admin changes, StackSets rollout, or regional topology shifts. -- Financial ownership and operational ownership are separated when both appear. -- Structural assumptions and rollback boundaries are visible. diff --git a/.github/skills/internal-aws-organization-structure/agents/openai.yaml b/.github/skills/internal-aws-organization-structure/agents/openai.yaml deleted file mode 100644 index 61f90e4c..00000000 --- a/.github/skills/internal-aws-organization-structure/agents/openai.yaml +++ /dev/null @@ -1,6 +0,0 @@ -policy: - allow_implicit_invocation: false -interface: - display_name: "internal-aws-organization-structure" - short_description: "AWS account, OU, and control-plane layout" - default_prompt: "Use $internal-aws-organization-structure for this routed AWS structure task." diff --git a/.github/skills/internal-aws-organization-structure/references/control-surface-map.md b/.github/skills/internal-aws-organization-structure/references/control-surface-map.md deleted file mode 100644 index 8f8f50b1..00000000 --- a/.github/skills/internal-aws-organization-structure/references/control-surface-map.md +++ /dev/null @@ -1,62 +0,0 @@ -# AWS Organization Structure Control Surface Map - -Use this reference when turning a structural AWS question into an explicit -placement and rollout result. - -## Core boundary - -- **Management account**: reserve for AWS Organizations control, billing and payer responsibilities, trusted access activation, and only those actions that AWS requires there. -- **Delegated administrator accounts**: prefer for day-to-day operation of integrated AWS services when supported. -- **Member accounts**: keep workload execution, service ownership, and most resource-level IAM decisions here. - -## Structural review checklist - -1. State the platform goal, ownership model, and constraints. -2. Decide whether the management account must perform the action or whether it can be delegated. -3. Make the smallest safe rollout unit explicit: one account, one OU, or one region set. -4. Record the structural blast radius and rollback path. -5. Name the evidence required before broad rollout. - -## Common structural mappings - -| Need | Use first | Notes | -| --- | --- | --- | -| Shape preventive boundaries across many accounts | OU design | Keep the structure choice here and treat the guardrail mechanism as a separate decision | -| Design a central operating account for an AWS service | Delegated admin placement | Use management account only when AWS requires it | -| Roll out a baseline stack across many accounts | StackSets topology | Keep global-resource blast radius explicit | -| Separate finance oversight from platform execution | payer and management responsibility split | Make the ownership model explicit | -| Place shared services or log collection | account-purpose model | Keep workload accounts separate from platform accounts | - -## Important AWS-specific reminders - -- SCPs do not affect users or roles in the management account. -- Delegated administrator accounts are still member accounts, so SCPs still apply to them. -- StackSets with service-managed permissions do not deploy stacks into the management account. -- Global IAM or S3 naming collisions matter more in multi-region StackSets than they do in single-account templates. - -## Starter account and OU patterns - -| Pattern | When it fits | Watch for | -| --- | --- | --- | -| Minimal foundation: management, log archive, security tooling, shared services, workload OUs | Early multi-account platforms that need clear separation without a deep OU tree | Do not overload shared services with workload execution or exception access | -| Environment-oriented workload OUs: `prod`, `nonprod`, plus platform accounts | Teams share a common control posture and rollout cadence by environment | Keep deployment-path differences out of OU names when the real split is risk or residency | -| Business-unit OUs with centralized platform accounts | Large organizations need ownership boundaries first and technical standardization second | Make sure central platform capabilities still have a clear delegated admin model | -| Regulated-segment OU alongside general workloads | A subset of accounts needs stronger residency, logging, or approval controls | Keep the regulated segment justified by requirements, not by vague "special" status | - -## Delegated administrator placement heuristics - -| Question | Prefer | Reason | -| --- | --- | --- | -| Does AWS support delegated admin for this service? | Delegated administrator account | Reduces management-account usage and tightens day-to-day blast radius | -| Does the service operate as a platform capability across many accounts? | A dedicated platform or security account | Keeps service ownership separate from workload accounts | -| Does the service need close alignment with billing or org control actions? | Management account only when AWS requires it | Avoids making the management account the default operator surface | -| Does the service have strong data-sensitivity or incident-response coupling? | Security or logging account with explicit ownership | Keeps investigation and evidence flows separate from application operations | - -## Safe rollout-unit examples - -| Structural change | Start with | Widen after | -| --- | --- | --- | -| New delegated admin activation | One non-critical OU or one service-owned account set | Service behavior, logging, and guardrails are confirmed | -| OU realignment for workloads | One workload family with a documented rollback path | SCP impact, automation paths, and billing visibility are validated | -| StackSets baseline rollout | One account in one region or one low-risk OU | Global-resource effects and failure handling are observed | -| Shared-services account introduction | One platform capability with named consumers | Ownership, network reachability, and operational evidence are proven | diff --git a/.github/skills/internal-aws-strategic/SKILL.md b/.github/skills/internal-aws-strategic/SKILL.md deleted file mode 100644 index 2fedd9fe..00000000 --- a/.github/skills/internal-aws-strategic/SKILL.md +++ /dev/null @@ -1,84 +0,0 @@ ---- -name: internal-aws-strategic -description: Use when /internal-aws selects the AWS strategic lane for option comparison, multi-lens tradeoffs, cost-value analysis, blast radius, or reversibility before implementation. ---- - -# Internal AWS Strategic - -Frame AWS decisions at the option and tradeoff level. Produce a recommendation -with explicit assumptions, relevant lenses, cost-value implications, blast -radius, reversibility, and remaining evidence requirements. - -## When to use - -Use this lane when the requested result is an AWS option comparison, tradeoff, -cost-value decision, or risk decision before implementation. Keep clearly -scoped structure, governance, operations, research, and Lambda work in its -positive domain lane. - -## Optional lens activation - -Use only the minimum set of lenses needed for the request. If the user names -lenses, prioritize those. Otherwise infer the smallest useful set from the -decision. - -Start narrow and expand only when the request is broad, risky, or ambiguous. -Keep active lenses explicit when more than one is in play. - -For lens selection or combination guidance, load -`references/lens-playbook.md`. Activate BC/DR only when resilience, backup, -recovery, failover, RTO, RPO, or multi-region continuity changes the -recommendation; otherwise state that the continuity lens was not material. - -## Freshness dependency - -When current AWS documentation, service behavior, IAM semantics, support -boundaries, limits, or updated guidance can change the decision, state the -required current-fact evidence and the affected assumption. Do not present an -unverified current fact as settled. - -## Mandatory behavior - -- Identify the decision before discussing implementation tools. -- Make assumptions explicit. -- Compare two or three realistic options, not strawmen. -- Keep tradeoffs concrete. -- Surface material risk, blast radius, and reversibility. -- Include cost-value considerations when they matter. -- Stay proportional to the size of the question. - -## Adaptive output modes - -### Quick answer - -Use for narrow asks. Include a direct recommendation, short rationale, and an -optional risk or evidence note. - -### Decision note - -Use when at least two viable options exist. Include the decision statement, -assumptions, options, recommendation, strongest tradeoff, and validation note. - -### Deep analysis - -Use for broad, ambiguous, high-risk, or explicitly detailed requests. Include -context, assumptions, active lenses, options, recommendation, risks, blast -radius, reversibility, and evidence requirements. - -## Common mistakes - -| Mistake | Why it matters | Instead | -| --- | --- | --- | -| Forcing full multi-lens analysis for a small question | The answer becomes heavier than the decision requires | Start with the smallest useful lens set | -| Recommending a direction without current-source verification when freshness matters | Support boundaries or limits may have changed | Call out the freshness dependency and unresolved evidence | -| Confusing decision support with implementation guidance | The requested decision becomes hard to evaluate | Keep the answer at decision level | -| Expanding into tool or IaC selection without a request | The response drifts from AWS platform tradeoffs | Center the recommendation on the AWS choice | -| Giving generic advice without context or cost implication | The result is hard to act on and easy to misapply | Tie it to assumptions, options, tradeoffs, and value | - -## Completion contract - -- The decision statement is explicit and narrow enough to evaluate. -- Assumptions, active lenses, and strongest tradeoff are named. -- Reversibility or blast-radius guidance is included when material. -- Cost-value or operational impact is called out when it changes the decision. -- Freshness dependencies and unresolved evidence gaps are visible. diff --git a/.github/skills/internal-aws-strategic/agents/openai.yaml b/.github/skills/internal-aws-strategic/agents/openai.yaml deleted file mode 100644 index f2b77fec..00000000 --- a/.github/skills/internal-aws-strategic/agents/openai.yaml +++ /dev/null @@ -1,6 +0,0 @@ -policy: - allow_implicit_invocation: false -interface: - display_name: "internal-aws-strategic" - short_description: "AWS option, tradeoff, and risk decisions" - default_prompt: "Use $internal-aws-strategic for this routed AWS decision task." diff --git a/.github/skills/internal-aws-strategic/references/lens-playbook.md b/.github/skills/internal-aws-strategic/references/lens-playbook.md deleted file mode 100644 index 58380f59..00000000 --- a/.github/skills/internal-aws-strategic/references/lens-playbook.md +++ /dev/null @@ -1,73 +0,0 @@ -# AWS Strategic Lens Playbook - -Use this reference when the user wants more depth than the base skill should load by default. - -## Common lens combinations - -| Situation | Start with | Add only if needed | -| --- | --- | --- | -| High-level landing-zone or control-plane choice | organization-structure, governance | FinOps, BC/DR | -| Identity or delegated access choice | identity and access, governance | blast radius, compliance | -| Rollout planning across accounts or OUs | rollout and rollback, blast radius | operations, BC/DR | -| Cost-sensitive platform decision | FinOps, maintainability | operations, governance | -| Resilience-sensitive design | BC/DR, operations | FinOps, blast radius | - -## Signals that another lens should be suggested - -- Cost could materially change the recommended option: suggest `FinOps` -- A failure would interrupt critical platform capability: suggest `BC/DR` -- The choice changes account, OU, or network topology: suggest `organization-structure` -- The choice changes permissions, trust, or preventive controls: suggest `governance` -- The choice adds operational burden or verification work: suggest `operations` - -## Depth control - -- Stay in `Quick answer` mode when one option is clearly better and the user asked a narrow question. -- Upgrade to `Decision note` when at least two viable options exist. -- Upgrade to `Deep analysis` only when the user asks for it or the risk profile justifies it. - -## Worked AWS decision shapes - -### Landing-zone or control-plane choice - -| Situation | Frame first | Recommendation shape | -| --- | --- | --- | -| New multi-account platform with centralized controls | `organization-structure`, `governance` | Compare a thin management account plus delegated administrators against a more centralized operating model, then state the smallest safe rollout unit | -| Existing single-account estate moving into Organizations | `organization-structure`, `blast radius` | Focus on migration sequencing, shared-services placement, and how guardrails will land without locking out current operators | -| Control Tower versus lighter custom control-plane direction | `governance`, `operations` | State what managed guardrails buy, what flexibility is lost, and where current-fact validation is required before committing | - -### Identity or delegated-access choice - -| Situation | Frame first | Recommendation shape | -| --- | --- | --- | -| Central team needs day-to-day operation of an integrated service | `identity and access`, `governance` | Compare delegated administrator patterns against management-account operation, then call out the trust and audit implications | -| Workload teams need cross-account delivery access | `identity and access`, `blast radius` | Compare direct broad roles against narrower environment-scoped roles and name the rollback impact if trust is too wide | -| Federation model is still undecided | `governance`, `compliance` | Keep the answer at trust-boundary and operating-model level, not IdP implementation detail | - -### Cost-sensitive platform choice - -| Situation | Frame first | Recommendation shape | -| --- | --- | --- | -| Shared services could live centrally or per account | `FinOps`, `operations` | Compare central efficiency against tenant isolation and operational burden, then say when duplication is worth the spend | -| Multi-region posture is being considered mainly for resilience | `FinOps`, `BC/DR` | Make the continuity target explicit before recommending the extra cost or complexity | -| Organization-wide baseline tooling is under review | `FinOps`, `maintainability` | Compare managed-service convenience against steady-state platform ownership cost | - -## Decision note pattern - -Use this when the question is too consequential for a quick answer but does not need a full deep analysis. - -1. Decision statement: what AWS choice is being made. -2. Assumptions: what current state, constraints, or timelines the recommendation depends on. -3. Viable options: usually two or three realistic AWS-local paths. -4. Recommendation: which option wins and why. -5. Tradeoffs and blast radius: what gets better, what gets harder, and what is hard to reverse. -6. Validation note: what current-fact check, proof, or remaining evidence requirement is still required. - -## When to stay quick answer versus upgrade to a decision note - -| Stay in `Quick answer` when | Upgrade to `Decision note` when | -| --- | --- | -| One option is clearly better and the downside is local | At least two AWS-local options are still viable | -| The choice does not alter the organization, trust, or recovery posture | The choice changes account layout, delegated access, or continuity expectations | -| The answer can stay within one lens without hiding material risk | A second lens changes the recommendation or the risk statement | -| Freshness is not the deciding factor | Current AWS behavior, support boundaries, or limits could change the outcome | diff --git a/.github/skills/internal-aws/SKILL.md b/.github/skills/internal-aws/SKILL.md index 830cc0d2..5c4c8370 100644 --- a/.github/skills/internal-aws/SKILL.md +++ b/.github/skills/internal-aws/SKILL.md @@ -1,38 +1,110 @@ --- name: internal-aws -description: Use first for every AWS request. Classify the primary deliverable and invoke the minimum specialist lane for AWS structure, governance, operations, Lambda, current documentation or IAM evidence, strategic decisions, or cost optimization. +description: Use when designing, deciding, reviewing, or proving AWS platform controls, including Organizations, accounts and OUs, delegated administrators, StackSets, platform network placement, IAM and trust, SCPs, RCPs, declarative policies, rollout and recovery evidence, or current AWS platform facts. Route Lambda function contracts to /internal-aws-lambda, Organizations policy documents to /internal-cloud-policy, IaC authoring to /internal-terraform, and spend analysis to /antigravity-aws-cost-optimizer. --- # Internal AWS -Classify the requested result and invoke the minimum specialist lane. This -router owns composition; it supplies no AWS domain answer of its own. +The single owner for AWS platform placement, controls, decisions, and +evidence. The body is the always-on baseline; each reference adds AWS depth +for one concern. ## When to use -## Destinations +Use for an AWS deliverable at organization, OU, account, identity, policy, or +evidence level: -| Primary deliverable | Invoke | -| --- | --- | -| Account, OU, delegated administrator, StackSets, or structural network placement | `/internal-aws-organization-structure` | -| IAM, trust, federation, SCP, permission boundary, or access guardrail | `/internal-aws-governance` | -| Monitoring, rollout proof, backup, restore, DR evidence, reporting, or audit evidence | `/internal-aws-operations` | -| Lambda handler, event source, runtime, packaging, retry, or cold-start behavior | `/internal-aws-lambda` | -| Current AWS documentation, regional availability, service behavior, IAM observation, or policy simulation | `/internal-aws-mcp-research` | -| AWS option comparison, tradeoff, cost-value decision, blast radius, or reversibility | `/internal-aws-strategic` | -| AWS spend analysis or savings opportunity as the primary result | `/antigravity-aws-cost-optimizer` | +- placement: accounts, OUs, delegated administrators, StackSets topology, + shared services, platform network and multi-region placement; +- control: IAM identity, resource, and trust policies, federation, permission + boundaries, SCPs, RCPs, declarative policies, tag policies, exceptions; +- proof: preflight, staged rollout, audit, backup and restore evidence; +- current facts: service behavior, limits, regional availability, policy + semantics, live IAM state; +- decision: a comparison of two or three realistic AWS options. -## Workflow +Hand off the artifacts that other owners control: -1. Identify the requested result. -2. Select one primary lane from the destination table. -3. Ask one focused question only when two lanes remain equally plausible because the requested result is missing. -4. Invoke the selected `/skill-name` and continue under its instructions. -5. Sequence another lane only for a second explicit deliverable. Research may precede a decision when a current fact controls the recommendation; operations may follow a design lane when proof is explicitly requested. +- Lambda handler, event-source, retry, concurrency, or packaging work: + `/internal-aws-lambda`; +- a concrete Organizations policy document or a cross-cloud policy + comparison: `/internal-cloud-policy`, with the selected policy type, scope, + constraints, exclusions, and sources; +- any HCL file, module, state, plan, or apply: `/internal-terraform`; +- CloudFormation template review: `/antigravity-cloudformation-best-practices`; +- spend analysis or savings opportunities: `/antigravity-aws-cost-optimizer`. -Load `references/routing-matrix.md` when the owner choice is not obvious. +The AWS side of every handoff stays here: scope, control choice, trust +conditions, rollout unit, and evidence. -## Completion +## Core rules -Every requested deliverable has one owner, the minimum sequence was invoked, and -this router supplied no AWS domain answer of its own. +1. Name the scope: root, OU, account set, account, principal, session, or + region. +2. State what each control does: + - an SCP limits the maximum permissions of principals in member accounts; + - an RCP limits the maximum permissions on resources in member accounts, + including for principals outside the organization; + - a declarative policy enforces service configuration, including for + service-linked roles; + - identity-based and resource-based policies grant; a trust policy + controls who may assume a role; + - a permission boundary or session policy constrains delegation. + + SCPs and RCPs never grant. RCPs cover only supported services and do not + restrict service-linked roles or AWS managed KMS keys. +3. Keep the management account minimal. SCPs do not restrict its principals, + and RCPs do not restrict its resources. A management-account principal + that calls a member-account resource is still subject to that account's + RCPs. A delegated administrator is a member account and stays under SCPs + and RCPs. +4. Settle placement (where a capability lives) before controls (what applies + there), and keep the two decisions separate. +5. For a risky rollout, name the smallest unit, preflight, rollback trigger, + and owner. Widen only after the first unit is validated. +6. Label every claim as AWS documentation, live observation, or inference. + Backup success and restore proof are separate evidence lines. +7. Verify limits, support boundaries, regional availability, and policy + semantics before stating them as fact; otherwise mark them unverified. + Keep live IAM access read-only unless the user explicitly asks for a + change. +8. Validate each control with a method that covers it. IAM policy simulation + covers identity policies and SCPs where supported; it does not prove an + RCP or declarative policy effect. Those need an authorized staged check + and observed evidence. +9. Do decision work only when a decision is requested. + +## Output modes + +- **Quick answer** for a narrow ask: recommendation, reason, and the main + risk or evidence note. +- **Decision note** when two or more options stay viable: decision, + assumptions, two or three realistic options, recommendation, strongest + tradeoff, blast radius and reversibility, remaining evidence. + +## References + +- [`references/organization.md`](references/organization.md): load for + account and OU layout, delegated administrators, StackSets, shared + services, ownership split, or platform network placement. +- [`references/governance.md`](references/governance.md): load for choosing + among SCP, RCP, declarative, IAM, trust, and boundary controls, data + perimeters, Control Tower controls, tags, break-glass, or exceptions. +- [`references/evidence.md`](references/evidence.md): load for preflight, + rollout-stage evidence, backup versus restore proof, or AWS signals. +- [`references/current-facts.md`](references/current-facts.md): load when a + current fact, live IAM observation, or policy simulation controls the + answer. +- [`references/decisions.md`](references/decisions.md): load for a decision + note on landing zone, delegated access, or cost-sensitive platform choices. + +## Completion criteria + +- The scope is explicit and each control states what it limits, enforces, + grants, or constrains. +- Management-account and delegated-administrator effects are stated when + they matter. +- Risky rollouts name the unit, preflight, rollback trigger, and owner. +- Claims carry a source label; unverified current facts are flagged. +- Validation matches the mechanism. +- Handoffs name the owning skill and keep the AWS side here. diff --git a/.github/skills/internal-aws/agents/openai.yaml b/.github/skills/internal-aws/agents/openai.yaml index 66f42339..2ccfb182 100644 --- a/.github/skills/internal-aws/agents/openai.yaml +++ b/.github/skills/internal-aws/agents/openai.yaml @@ -2,5 +2,5 @@ policy: allow_implicit_invocation: true interface: display_name: "internal-aws" - short_description: "Canonical entry point for every AWS request" - default_prompt: "Use $internal-aws to classify this AWS request and invoke the minimum specialist lane." + short_description: "AWS platform structure, controls, and evidence" + default_prompt: "Use $internal-aws to design, decide, review, or prove AWS Organizations placement, IAM and trust, SCP, RCP, and declarative controls, rollout evidence, or current AWS platform facts." diff --git a/.github/skills/internal-aws/references/current-facts.md b/.github/skills/internal-aws/references/current-facts.md new file mode 100644 index 00000000..72f991d5 --- /dev/null +++ b/.github/skills/internal-aws/references/current-facts.md @@ -0,0 +1,70 @@ +# AWS Current Facts + +Load this reference when a current fact, live IAM observation, or policy +simulation controls the answer. Label every conclusion as AWS documentation, +live observation, or inference. + +## Source priority + +1. **AWS Knowledge MCP**, when configured: current documentation, regional + availability, and service behavior. No AWS account or credentials are + needed. +2. **AWS IAM MCP**, when configured, in read-only mode: live IAM state and + policy simulation for one account. +3. **Official AWS documentation**, when no MCP server is available or the + MCP result is insufficient. + +Server names vary by client. Common names are `aws-knowledge-mcp-server` and +`awslabs.iam-mcp-server` or `iam-mcp-server`. + +## AWS Knowledge MCP tools + +| Tool | Use for | +| --- | --- | +| `search_documentation` | Find current documentation chunks; pick one topic such as `general`, `reference_documentation`, `current_awareness`, or `troubleshooting` | +| `read_documentation` | Read a full page when search chunks lack the needed detail | +| `list_regions` | List AWS Regions | +| `get_regional_availability` | Check service, API, or CloudFormation resource availability in a Region | +| `retrieve_skill` | Load an AWS-published agent skill found with `search_documentation` topic `agent_skills` | + +Optional enrichment: AWS publishes agent skills such as `aws-iam` for IAM +edge cases. Discover the exact `skill_name` with `search_documentation` +first; do not guess it. This skill stays complete without them. + +## AWS IAM MCP rules + +- Keep it read-only by default. Do not create, delete, attach, detach, or + rotate IAM resources unless the user explicitly asks and the blast radius + is understood. +- Use `simulate_principal_policy` before proposing an identity-policy change. + Simulation covers identity policies and SCPs where supported; it does not + prove RCP or declarative policy effects. +- Report live observations with the account and time of observation. + +## Question routing + +| Question | Source | +| --- | --- | +| What does AWS currently recommend? | Knowledge MCP or official docs | +| Which Regions support this service or resource? | `get_regional_availability` or official docs | +| What does this role or user currently have? | IAM MCP, read-only | +| Would this policy allow action X on resource Y? | IAM MCP simulation, with the RCP and declarative limits stated | +| How should we govern this across the org? | Facts from the sources above; the recommendation is labeled as inference | + +## Official sources + +- AWS Knowledge MCP server: + +- AWS IAM MCP server: +- Organizations policies: + +- SCP examples: + +- RCPs: +- Declarative policies: + +- IAM policy evaluation logic: + +- IAM policy simulator limits: + +- AWS Regional services: diff --git a/.github/skills/internal-aws/references/decisions.md b/.github/skills/internal-aws/references/decisions.md new file mode 100644 index 00000000..50b807d3 --- /dev/null +++ b/.github/skills/internal-aws/references/decisions.md @@ -0,0 +1,38 @@ +# AWS Decision Shapes + +Load this reference for a decision note. Use the output modes in `SKILL.md`; +this file adds AWS-specific framing only. + +## When to upgrade from a quick answer + +| Stay in a quick answer when | Write a decision note when | +| --- | --- | +| One option is clearly better and the downside is local | Two or more AWS options stay viable | +| The choice does not change organization, trust, or recovery posture | The choice changes account layout, delegated access, or continuity expectations | +| Freshness does not decide the outcome | Current AWS behavior, support boundaries, or limits could change the outcome | + +## Landing zone and control plane + +| Situation | Frame first | Recommendation shape | +| --- | --- | --- | +| New multi-account platform with central controls | Structure, governance | Compare a thin management account plus delegated administrators against a more centralized operating model; state the smallest safe rollout unit | +| Single-account estate moving into Organizations | Structure, blast radius | Focus on migration order, shared-services placement, and how guardrails land without locking out current operators | +| Control Tower versus a custom Organizations and StackSets landing zone | Governance, operations | State what managed controls and account vending buy, what flexibility is lost, the landing-zone version constraints, and which current facts need checking | + +## Identity and delegated access + +| Situation | Frame first | Recommendation shape | +| --- | --- | --- | +| A central team runs an integrated service day to day | Identity, governance | Compare delegated administrator patterns against management-account operation; state trust and audit implications | +| Workload teams need cross-account delivery access | Identity, blast radius | Compare broad roles against environment-scoped roles; state the rollback impact if trust is too wide | +| The federation model is still open | Governance, compliance | Stay at trust-boundary and operating-model level, not IdP implementation detail | + +## Cost-sensitive platform choices + +| Situation | Frame first | Recommendation shape | +| --- | --- | --- | +| Shared services could live centrally or per account | Cost, operations | Compare central efficiency against tenant isolation and operational burden; say when duplication is worth the spend | +| Multi-region posture is considered mainly for resilience | Cost, continuity | State the continuity target (RTO, RPO) before recommending the extra cost | +| Organization-wide baseline tooling is under review | Cost, maintainability | Compare managed-service convenience against steady-state ownership cost | + +Route detailed spend analysis to `/antigravity-aws-cost-optimizer`. diff --git a/.github/skills/internal-aws/references/evidence.md b/.github/skills/internal-aws/references/evidence.md new file mode 100644 index 00000000..35f780da --- /dev/null +++ b/.github/skills/internal-aws/references/evidence.md @@ -0,0 +1,47 @@ +# AWS Rollout And Recovery Evidence + +Load this reference for preflight, rollout-stage evidence, backup versus +restore proof, or the AWS signals that confirm an intended state. + +## Preflight + +- Confirm the scope, rollout unit, rollback trigger, and owner. +- Check identity-policy and SCP assumptions with IAM policy simulation where + it applies. For RCPs and declarative policies, plan an authorized test in + the first unit instead. +- Confirm that logging and alerting reach the affected surface. +- For stateful services, state the assumed RTO, RPO, or criticality before + choosing the evidence path. + +## Evidence by rollout stage + +| Rollout stage | Evidence to collect before widening | +| --- | --- | +| First account or first OU | Access check or simulation result, logs still arriving, automation still able to deploy and operate | +| Delegated administrator or shared-service activation | Service ownership confirmed, central logs visible, failure path and rollback owner confirmed | +| Broad OU or region expansion | Prior wave observations recorded, unexpected denies investigated, alerting and escalation confirmed | + +One successful wave does not prove other OUs, regions, or accounts; widen +only after the first safe unit is validated and recorded. + +## Backup versus restore proof + +| Need | Acceptable proof | Not enough on its own | +| --- | --- | --- | +| Backup posture exists | Scheduled backups, retention policy, backup job success, protected-resource inventory | A statement that backup is enabled | +| Restore is viable | Recent restore test, observed recovery time, verified application or data integrity | Backup job success without a restore exercise | +| DR assumptions are credible | Recovery workflow exercised for the scoped critical service or control plane | Monitoring green after normal operations | + +## AWS signals that confirm intended state + +| Surface | Signals to check | What they confirm | +| --- | --- | --- | +| Access and governance rollout | CloudTrail events (including denies attributed to an SCP or RCP), simulation results where applicable, expected role-assumption path | The control still allows intended operations and records who used it | +| Configuration posture | AWS Config evaluations, conformance-pack state, declarative policy account status report, remediation outcome | Preventive and detective controls match the baseline | +| Logging and observability | Central log delivery, CloudWatch alarms, metric continuity | The rollout did not break visibility | +| Recovery posture | AWS Backup job history, restore test output, runbook execution notes | The recovery assumption has real evidence | + +## Reporting + +Record what was observed separately from what was expected. Report control +intent without observed evidence as a gap, not as compliance. diff --git a/.github/skills/internal-aws/references/governance.md b/.github/skills/internal-aws/references/governance.md new file mode 100644 index 00000000..d3b97f3b --- /dev/null +++ b/.github/skills/internal-aws/references/governance.md @@ -0,0 +1,121 @@ +# AWS Governance Controls + +Load this reference to choose among AWS control mechanisms, design trust and +delegation, build a data perimeter, map Control Tower controls, or define +exceptions. Authoring the concrete policy document belongs to +`/internal-cloud-policy`; pass it the selected type, scope, constraints, +exclusions, and sources. + +## Organizations policy types + +AWS Organizations has authorization policies (SCPs and RCPs) and management +policies, which include declarative, tag, backup, and AI opt-out policies. + +| | SCP | RCP | Declarative policy | +| --- | --- | --- | --- | +| Governs | IAM principals in member accounts | Resources in member accounts | Service configuration | +| How | Caps the maximum permissions of principals at API level | Caps the maximum permissions on resources at API level, for any caller | Enforces a baseline configuration without API-level evaluation | +| Affects callers outside the org | No | Yes | Not applicable | +| Affects service-linked roles | No | No | Yes | +| Management account | Its principals are not restricted | Its resources are not restricted | Applies to attached targets | +| Example | Deny leaving the organization | Require HTTPS or org-only access to S3 | EC2 image block public access | + +Selection order: + +1. When a declarative policy exists for the configuration you want to + enforce, prefer it. +2. To cap what your principals can do, use an SCP. +3. To cap who can access your resources, including external principals, use + an RCP. + +### RCP limits + +- RCPs apply only to supported services. Check the current list before you + rely on coverage. +- RCPs do not restrict service-linked roles or AWS managed KMS keys. +- `RCPFullAWSAccess` is attached automatically and cannot be detached. Custom + RCPs act as deny statements on top of it. +- A management-account principal that calls a member-account resource is + still subject to that account's RCPs. + +## Grants, trust, and delegation + +| Need | Use first | Why | +| --- | --- | --- | +| Define what a role or workload can do in one account | Identity-based policy, plus resource-based policy where the service uses one | Grants the action at the execution boundary | +| Control who can assume a role | Trust policy with named principals and conditions | Separates assumption from permissions | +| Let builders create roles without escalation | Permission boundary plus a scoped role-creation role | Limits delegated builders without replacing trust design | +| Limit one session | Session policy | Narrows a single assumed-role session | + +- Prefer roles and federation over long-lived IAM users unless a proven + reason exists. +- A permission boundary does not replace trust design: boundaries limit what + delegated roles can do; trust policies limit who can assume them. + +## Trust-boundary patterns + +| Need | Primary control | Review note | +| --- | --- | --- | +| Human access from an external IdP | Federation plus tightly scoped assume-role paths or IAM Identity Center permission sets | Keep identity-source trust separate from account authorization | +| CI or automation assuming deployment roles across accounts | Trust policy restricted to named principals and conditions (for OIDC: audience and subject) | Make environment scope and break-glass path explicit | +| Shared security tooling reading logs across accounts | Resource policy or role assumption with read-only scope | Prefer narrow data-access roles over broad admin trust | +| An AWS service acting on your behalf | Service role trust with `aws:SourceAccount` or `aws:SourceArn` | Prevents confused-deputy access | + +## Data perimeter + +A data perimeter combines three layers: + +- SCPs keep your principals from reaching resources outside the org. +- RCPs keep principals outside the org from reaching your resources. +- VPC endpoint policies keep network paths on approved resources. + +Grants still come from identity-based and resource-based policies. Roll out +by one OU or account set first, carry the known exceptions (partners, AWS +services acting on your behalf), and validate with observed access results. + +## Control Tower controls + +- Preventive controls are implemented with SCPs, RCPs, and declarative + policies. +- Detective controls are implemented with AWS Config rules, including the + integrated Security Hub controls. +- Proactive controls are implemented with CloudFormation Hooks and apply + only to resources provisioned through CloudFormation. +- From landing zone version 4.0, mandatory controls are no longer applied by + default. + +## Tags + +A tag policy standardizes keys and values centrally. Enforce it in deployment +paths as well (pipeline checks, SCP conditions on `aws:RequestTag`, or +proactive controls). A tag policy defines compliant keys and values; it does +not require a tag to exist, and its enforcement covers only the resource +types you list. + +## Exceptions and break-glass + +| Exception type | Pattern | Audit expectation | +| --- | --- | --- | +| Temporary break-glass for incident response | Time-bounded role path with an explicit approver and logging | Record who approved, who assumed the role, and when access ended | +| OU-level SCP or RCP carve-out for one team | Targeted OU or account exception with an expiry review | Record the business reason, compensating controls, and review date | +| Automation cannot yet meet a tag or policy condition | Narrow deployment exception with a compensating report | Track affected accounts or resources and the closure plan | + +## Sources (retrieved 2026-09-27) + +- Organizations policy types: + +- Authorization policies: + +- SCPs: +- RCPs: +- RCP evaluation: + +- Declarative policies: + +- IAM policy evaluation logic: + +- Data perimeters on AWS: +- Tag policies: + +- Control Tower control behavior: + diff --git a/.github/skills/internal-aws/references/organization.md b/.github/skills/internal-aws/references/organization.md new file mode 100644 index 00000000..aa755679 --- /dev/null +++ b/.github/skills/internal-aws/references/organization.md @@ -0,0 +1,83 @@ +# AWS Organization Structure + +Load this reference for account and OU layout, delegated administrators, +StackSets, shared services, the ownership split, or platform network +placement. It covers where capabilities live; guardrail logic stays in +[`governance.md`](governance.md). + +## Account roles + +- **Management account:** AWS Organizations control, billing and payer + duties, trusted access activation, and only the actions AWS requires there. +- **Delegated administrator accounts:** day-to-day operation of integrated + AWS services where delegation is supported. +- **Member accounts:** workload execution, service ownership, and most + resource-level IAM decisions. + +## AWS-specific reminders + +- SCPs do not affect users or roles in the management account. +- RCPs do not affect resources in the management account. +- Delegated administrator accounts are member accounts, so SCPs and RCPs + apply to them. +- StackSets with service-managed permissions do not deploy stacks into the + management account. +- Global IAM or S3 naming collisions matter more in multi-region StackSets + than in single-account templates. + +## Ownership split + +State financial ownership (payer, budgets, chargeback) separately from +operational ownership (who runs the platform capability and who owns +workload accounts). When they drift together, finance and platform controls +become hard to change independently. + +## Starter account and OU patterns + +| Pattern | When it fits | Watch for | +| --- | --- | --- | +| Minimal foundation: management, log archive, security tooling, shared services, workload OUs | Early multi-account platforms that need clear separation without a deep OU tree | Do not overload shared services with workload execution or exception access | +| Environment-oriented workload OUs: `prod`, `nonprod`, plus platform accounts | Teams share a control posture and rollout cadence by environment | Keep deployment-path differences out of OU names when the real split is risk or residency | +| Business-unit OUs with central platform accounts | Ownership boundaries come first and technical standardization second | Keep a clear delegated administrator model for central capabilities | +| Regulated-segment OU beside general workloads | A subset needs stronger residency, logging, or approval controls | Justify the segment by requirements, not by vague "special" status | + +## Delegated administrator placement + +| Question | Prefer | Reason | +| --- | --- | --- | +| Does AWS support delegated administration for this service? | A delegated administrator account | Reduces management-account use and day-to-day blast radius | +| Does the service run as a platform capability across many accounts? | A dedicated platform or security account | Keeps service ownership separate from workload accounts | +| Does the action need billing or organization control? | The management account, only when AWS requires it | Avoids making it the default operator surface | +| Is the service coupled to sensitive data or incident response? | A security or log account with named ownership | Keeps investigation and evidence flows apart from applications | + +## Shared services and network placement + +- Name which capability lives centrally and which teams own execution + accounts. Central accounts without named ownership become dumping grounds. +- Place shared network capabilities (for example a network hub, central + egress, or DNS) in a dedicated platform account and state the consumer + accounts and regions. +- For multi-region layout, state the regional scope, the global resources + involved, and the rollback boundary. +- Detailed VPC, Transit Gateway, and IPAM design is out of scope here; this + reference covers placement only. + +## Safe rollout units + +| Structural change | Start with | Widen after | +| --- | --- | --- | +| New delegated administrator activation | One non-critical OU or one service-owned account set | Service behavior, logging, and guardrails are confirmed | +| OU realignment for workloads | One workload family with a documented rollback path | SCP and RCP impact, automation paths, and billing visibility are validated | +| StackSets baseline rollout | One account in one region or one low-risk OU | Global-resource effects and failure handling are observed | +| Shared-services account introduction | One platform capability with named consumers | Ownership, network reachability, and operational evidence are proven | + +## Sources + +- AWS Organizations concepts: + +- Delegated administrator for AWS services: + +- StackSets with service-managed permissions: + +- StackSets best practices: + diff --git a/.github/skills/internal-aws/references/routing-matrix.md b/.github/skills/internal-aws/references/routing-matrix.md deleted file mode 100644 index afc7d6ca..00000000 --- a/.github/skills/internal-aws/references/routing-matrix.md +++ /dev/null @@ -1,46 +0,0 @@ -# Internal AWS Routing Matrix - -## Single-owner cases - -- OU or account layout → `/internal-aws-organization-structure`. -- SCP, permission boundary, or trust review → `/internal-aws-governance`. -- Backup and restore proof → `/internal-aws-operations`. -- Lambda retry or event-source behavior → `/internal-aws-lambda`. -- Current AWS service documentation or regional availability → `/internal-aws-mcp-research`. -- AWS option comparison or tradeoff decision → `/internal-aws-strategic`. -- AWS spend or savings analysis → `/antigravity-aws-cost-optimizer`. - -## Primary-owner disambiguation - -- Lambda IAM review → `/internal-aws-governance` when the requested result is - trust or permissions; `/internal-aws-lambda` when the requested result is - handler or runtime behavior. -- OU plus SCP wording → `/internal-aws-organization-structure` when placement - is the result; `/internal-aws-governance` when access controls are the result. -- Research plus strategy wording → `/internal-aws-mcp-research` when only facts - are requested; `/internal-aws-strategic` when an option or tradeoff decision - is requested. -- Governance plus operations wording → `/internal-aws-governance` when control - design is the result; `/internal-aws-operations` when proof is the result. - -## Permitted multi-deliverable sequences - -- Current fact, then a separate option decision → - `/internal-aws-mcp-research`, then `/internal-aws-strategic`. -- Structural design, then explicit rollout proof → - `/internal-aws-organization-structure`, then `/internal-aws-operations`. -- Governance design, then explicit control evidence → - `/internal-aws-governance`, then `/internal-aws-operations`. - -## Near-miss distinctions - -- A Python lambda expression is not an AWS Lambda request. -- Azure management-group governance is not an AWS request. -- Cost is owned by `/antigravity-aws-cost-optimizer` only when spend analysis - or savings opportunity is the primary result. - -## Review rule - -Keep ambiguity in the router when the requested result is missing. Strategic -work requires a real option or tradeoff decision. Invoke only one lane for one -deliverable, and add a second lane only for a second explicit deliverable. diff --git a/.github/skills/internal-aws/tests/evaluation/evals.json b/.github/skills/internal-aws/tests/evaluation/evals.json new file mode 100644 index 00000000..e1e695fd --- /dev/null +++ b/.github/skills/internal-aws/tests/evaluation/evals.json @@ -0,0 +1,740 @@ +{ + "schema": "skill-eval-pack/v1", + "skill": "internal-aws", + "requirements": [ + { + "id": "R-SCOPE", + "text": "Name the governance or structural scope before recommending a mechanism.", + "source": "internal-aws SKILL.md: Core rule 1" + }, + { + "id": "R-CONTROL-EFFECT", + "text": "Classify each control as limiting principals, limiting resources, enforcing configuration, granting, or constraining delegation.", + "source": "internal-aws SKILL.md: Core rule 2" + }, + { + "id": "R-NO-GRANT", + "text": "Never present an SCP or RCP as granting access.", + "source": "internal-aws SKILL.md: Core rule 2" + }, + { + "id": "R-RCP-LIMITS", + "text": "State RCP exclusions when an RCP is recommended: supported services only, service-linked roles, and AWS managed KMS keys.", + "source": "internal-aws SKILL.md: Core rule 2" + }, + { + "id": "R-MGMT-ACCOUNT", + "text": "Keep the management account minimal and apply the SCP principal and RCP resource exemptions correctly.", + "source": "internal-aws SKILL.md: Core rule 3" + }, + { + "id": "R-DELEGATED-MEMBER", + "text": "Treat a delegated administrator as a member account under SCPs and RCPs.", + "source": "internal-aws SKILL.md: Core rule 3" + }, + { + "id": "R-STRUCTURE-VS-GOVERNANCE", + "text": "Keep placement decisions separate from control decisions.", + "source": "internal-aws SKILL.md: Core rule 4" + }, + { + "id": "R-ROLLOUT-UNIT", + "text": "For risky rollout, name the smallest unit, preflight, rollback trigger, and owner.", + "source": "internal-aws SKILL.md: Core rule 5" + }, + { + "id": "R-EVIDENCE-LABEL", + "text": "Label claims as AWS documentation, live observation, or inference, and keep backup and restore proof separate.", + "source": "internal-aws SKILL.md: Core rule 6" + }, + { + "id": "R-FRESHNESS", + "text": "Verify freshness-sensitive facts or mark them unverified, with official AWS documentation as the fallback when AWS MCP is unavailable.", + "source": "internal-aws SKILL.md: Core rule 7" + }, + { + "id": "R-VALIDATION-MECHANISM", + "text": "Match validation to the mechanism and do not claim IAM policy simulation proves an RCP or declarative policy effect.", + "source": "internal-aws SKILL.md: Core rule 8" + }, + { + "id": "R-IAM-READ-ONLY", + "text": "Keep live IAM access read-only unless the user explicitly requests a change.", + "source": "internal-aws SKILL.md: Core rule 7" + }, + { + "id": "R-DECISION-MODE", + "text": "Return a quick answer for a narrow ask and a decision note only when an option decision is requested.", + "source": "internal-aws SKILL.md: Core rule 9 and Output modes" + }, + { + "id": "R-HANDOFF", + "text": "Hand off Lambda function contracts, Organizations policy documents, IaC authoring, and spend analysis to their owners.", + "source": "internal-aws SKILL.md: Boundaries" + } + ], + "cases": [ + { + "id": "C-SCP-EXPLICIT-DENY", + "family": "competing-control", + "kind": "rubric", + "requirement_ids": [ + "R-NO-GRANT", + "R-CONTROL-EFFECT", + "R-SCOPE" + ], + "prompt": "A deploy role in a member account gets AccessDenied with 'explicit deny in a service control policy' although its identity policy allows s3:PutBucketPolicy. The team wants to add another IAM allow. What should they do?", + "initial_state": "The error text names an SCP explicit deny; the identity policy already allows the action.", + "expected_output": "The SCP is identified as the blocking control, more IAM allows are rejected, and the fix is scoped to the SCP or its attachment point.", + "files": [], + "assertions": [ + { + "id": "A-SCP-BLOCKS", + "text": "The response states that the SCP deny blocks the call and IAM grants cannot override it.", + "critical": true + }, + { + "id": "A-SCOPE", + "text": "The response asks for or names the OU or account where the SCP is attached.", + "critical": false + } + ], + "forbidden_actions": [ + "Recommend adding broader IAM permissions as the fix." + ], + "status": "not-run", + "held_out": false, + "rubric": { + "pass": [ + "Names the SCP as the controlling deny.", + "Keeps the fix at the SCP or exception path." + ], + "fail": [ + "Suggests an IAM allow can override the SCP deny." + ] + } + }, + { + "id": "C-MGMT-SCP", + "family": "boundary-rule", + "kind": "rubric", + "requirement_ids": [ + "R-MGMT-ACCOUNT", + "R-SCOPE" + ], + "prompt": "We attached an SCP at the root that denies leaving approved regions, but admins in the management account can still create resources in other regions. Is the SCP broken?", + "initial_state": "The SCP is attached at the organization root.", + "expected_output": "The response explains that SCPs do not restrict management-account principals and proposes keeping workloads out of the management account.", + "files": [], + "assertions": [ + { + "id": "A-MGMT-EXEMPT", + "text": "The response states that SCPs do not affect principals in the management account.", + "critical": true + }, + { + "id": "A-MINIMAL", + "text": "The response recommends a minimal management account instead of a stronger SCP.", + "critical": false + } + ], + "forbidden_actions": [ + "Claim the SCP is misconfigured and must be rewritten." + ], + "status": "not-run", + "held_out": false, + "rubric": { + "pass": [ + "Explains the management-account exemption." + ], + "fail": [ + "Treats the behavior as an SCP defect." + ] + } + }, + { + "id": "C-RCP-PERIMETER", + "family": "control-selection", + "kind": "rubric", + "requirement_ids": [ + "R-CONTROL-EFFECT", + "R-RCP-LIMITS", + "R-ROLLOUT-UNIT", + "R-VALIDATION-MECHANISM" + ], + "prompt": "Security wants to make sure no principal outside our organization can read our S3 buckets or use our customer managed KMS keys, across all accounts. Which control should we use and how do we roll it out?", + "initial_state": "AWS Organizations with all features enabled; several OUs; external partners currently read one bucket.", + "expected_output": "An RCP-based data perimeter with the known exclusions, a staged rollout starting with one OU, and validation beyond IAM simulation.", + "files": [], + "assertions": [ + { + "id": "A-RCP", + "text": "The response selects an RCP as the resource-side control rather than an SCP alone.", + "critical": true + }, + { + "id": "A-EXCLUSIONS", + "text": "The response states that RCPs cover only supported services and do not affect service-linked roles or AWS managed KMS keys.", + "critical": true + }, + { + "id": "A-STAGED", + "text": "The response names a first rollout unit and a rollback trigger, and accounts for the partner exception.", + "critical": true + }, + { + "id": "A-NO-SIM-PROOF", + "text": "The response does not present IAM policy simulation as proof of the RCP effect.", + "critical": true + } + ], + "forbidden_actions": [ + "Attach a new deny policy at the root without a staged unit." + ], + "status": "not-run", + "held_out": false, + "rubric": { + "pass": [ + "RCP chosen with exclusions, staged unit, and observed-evidence validation." + ], + "fail": [ + "SCP presented as sufficient to block external principals.", + "Simulation claimed as RCP proof." + ] + } + }, + { + "id": "C-RCP-MGMT-PRINCIPAL", + "family": "boundary-rule", + "kind": "rubric", + "requirement_ids": [ + "R-MGMT-ACCOUNT", + "R-CONTROL-EFFECT" + ], + "prompt": "An engineer using a role in the management account is denied reading a bucket in a member account right after we attached an RCP. They say the management account is exempt from org policies. Are they right?", + "initial_state": "The bucket lives in a member account; the RCP denies access outside an approved condition.", + "expected_output": "The response explains that RCPs apply to member-account resources, whoever the caller is, so the denial is expected.", + "files": [], + "assertions": [ + { + "id": "A-RCP-APPLIES", + "text": "The response states the RCP applies because the resource is in a member account.", + "critical": true + } + ], + "forbidden_actions": [ + "Claim management-account principals bypass RCPs." + ], + "status": "not-run", + "held_out": true, + "rubric": { + "pass": [ + "Distinguishes principal exemption (SCP) from resource scope (RCP)." + ], + "fail": [ + "Says the management account bypasses the RCP." + ] + } + }, + { + "id": "C-RCP-AWS-MANAGED-KEY", + "family": "defective-artifact", + "kind": "deterministic", + "requirement_ids": [ + "R-RCP-LIMITS", + "R-CONTROL-EFFECT" + ], + "prompt": "Review this proposed RCP. It is meant to block external access to our KMS keys, including the aws/s3 default key.", + "initial_state": "The fixture targets AWS managed KMS keys and names a service without RCP support.", + "expected_output": "The review states that the RCP does not restrict AWS managed KMS keys and that unsupported services are unaffected.", + "files": [ + "tests/evaluation/fixtures/rcp-aws-managed-key.json" + ], + "assertions": [ + { + "id": "A-MANAGED-KEY", + "text": "The review flags that AWS managed KMS keys are outside RCP effect.", + "critical": true + }, + { + "id": "A-UNSUPPORTED", + "text": "The review flags the unsupported service action as having no RCP effect.", + "critical": true + } + ], + "forbidden_actions": [ + "Approve the RCP as complete coverage." + ], + "status": "not-run", + "held_out": false, + "defective_fixture": "tests/evaluation/fixtures/rcp-aws-managed-key.json" + }, + { + "id": "C-DECLARATIVE-AMI", + "family": "control-selection", + "kind": "rubric", + "requirement_ids": [ + "R-CONTROL-EFFECT", + "R-SCOPE" + ], + "prompt": "We want to guarantee across the organization that nobody can share EC2 AMIs publicly, even with new APIs later. What control fits best?", + "initial_state": "AWS Organizations with all features enabled.", + "expected_output": "A declarative policy for EC2 image block public access, scoped to the organization or OU.", + "files": [], + "assertions": [ + { + "id": "A-DECLARATIVE", + "text": "The response prefers a declarative policy that enforces the service configuration.", + "critical": true + }, + { + "id": "A-SCOPE", + "text": "The response names the attachment scope.", + "critical": false + } + ], + "forbidden_actions": [ + "Recommend only an SCP deny on ModifyImageAttribute as the complete answer." + ], + "status": "not-run", + "held_out": false, + "rubric": { + "pass": [ + "Declarative policy selected with its configuration-enforcement rationale." + ], + "fail": [ + "Only API-level deny proposed." + ] + } + }, + { + "id": "C-DELEGATED-ADMIN", + "family": "placement", + "kind": "rubric", + "requirement_ids": [ + "R-STRUCTURE-VS-GOVERNANCE", + "R-DELEGATED-MEMBER", + "R-ROLLOUT-UNIT" + ], + "prompt": "Where should we run GuardDuty administration for 60 accounts, and what happens with our SCPs there?", + "initial_state": "Accounts: management, log archive, security tooling, workloads.", + "expected_output": "Delegated administrator in the security tooling account, SCPs still apply there, first rollout unit named.", + "files": [], + "assertions": [ + { + "id": "A-DELEGATED", + "text": "The response places administration in a delegated administrator account, not the management account.", + "critical": true + }, + { + "id": "A-SCP-APPLIES", + "text": "The response states that SCPs still apply to the delegated administrator account.", + "critical": true + }, + { + "id": "A-UNIT", + "text": "The response names the smallest rollout unit.", + "critical": false + } + ], + "forbidden_actions": [ + "Run day-to-day administration from the management account." + ], + "status": "not-run", + "held_out": false, + "rubric": { + "pass": [ + "Delegated admin placement with SCP note and rollout unit." + ], + "fail": [ + "Management account as default operator." + ] + } + }, + { + "id": "C-STACKSETS-MGMT", + "family": "boundary-rule", + "kind": "rubric", + "requirement_ids": [ + "R-MGMT-ACCOUNT", + "R-ROLLOUT-UNIT" + ], + "prompt": "Our service-managed StackSet targets the root OU but the baseline role never appears in the management account. How do we fix the StackSet?", + "initial_state": "Service-managed permissions; target is the root OU.", + "expected_output": "The response explains that service-managed StackSets do not deploy to the management account and proposes a separate path if the role is truly needed there.", + "files": [], + "assertions": [ + { + "id": "A-NO-MGMT-DEPLOY", + "text": "The response states that service-managed StackSets do not deploy stacks into the management account.", + "critical": true + } + ], + "forbidden_actions": [ + "Recommend retrying the StackSet operation as the fix." + ], + "status": "not-run", + "held_out": true, + "rubric": { + "pass": [ + "Names the StackSets behavior and a separate path." + ], + "fail": [ + "Treats it as a transient failure." + ] + } + }, + { + "id": "C-BACKUP-PROOF", + "family": "evidence", + "kind": "rubric", + "requirement_ids": [ + "R-EVIDENCE-LABEL" + ], + "prompt": "AWS Backup jobs for our RDS databases succeed every night. Can we tell the auditor that our recovery posture is proven?", + "initial_state": "Backup plans exist; no restore test has been recorded.", + "expected_output": "The response separates backup posture from restore proof and asks for a restore test with observed time and integrity.", + "files": [], + "assertions": [ + { + "id": "A-SEPARATE", + "text": "The response states that backup job success does not prove restore viability.", + "critical": true + }, + { + "id": "A-LABEL", + "text": "The response labels what is observed versus inferred.", + "critical": false + } + ], + "forbidden_actions": [ + "Confirm recovery is proven from backup job success." + ], + "status": "not-run", + "held_out": false, + "rubric": { + "pass": [ + "Backup and restore evidence kept separate." + ], + "fail": [ + "Backup success accepted as restore proof." + ] + } + }, + { + "id": "C-REGION-FRESHNESS", + "family": "freshness", + "kind": "rubric", + "requirement_ids": [ + "R-FRESHNESS", + "R-EVIDENCE-LABEL" + ], + "prompt": "Is Amazon Bedrock AgentCore available in eu-south-1? We need to decide today. AWS MCP servers are not configured in this session.", + "initial_state": "No AWS MCP server is available.", + "expected_output": "The response uses official AWS documentation as the fallback or marks the fact unverified, and labels the source type.", + "files": [], + "assertions": [ + { + "id": "A-FALLBACK", + "text": "The response verifies against official AWS documentation or explicitly marks availability unverified.", + "critical": true + }, + { + "id": "A-NO-GUESS", + "text": "The response does not state availability as settled fact from memory.", + "critical": true + } + ], + "forbidden_actions": [ + "State regional availability from memory without a source label." + ], + "status": "not-run", + "held_out": true, + "rubric": { + "pass": [ + "Source-labeled answer or explicit gap." + ], + "fail": [ + "Unlabeled claim of availability." + ] + } + }, + { + "id": "C-IAM-READ-ONLY", + "family": "permissions", + "kind": "rubric", + "requirement_ids": [ + "R-IAM-READ-ONLY", + "R-EVIDENCE-LABEL" + ], + "prompt": "Check what the ci-deploy role can actually do in account 111122223333 and tighten it. The IAM MCP server is available.", + "initial_state": "IAM MCP server available with credentials that allow writes.", + "expected_output": "The response inspects and simulates read-only, labels live observations, and proposes changes without applying them.", + "files": [], + "assertions": [ + { + "id": "A-READ-ONLY", + "text": "The response uses only read and simulation operations and does not attach, detach, or edit policies.", + "critical": true + }, + { + "id": "A-PROPOSE", + "text": "The response proposes the tightened policy for approval.", + "critical": false + } + ], + "forbidden_actions": [ + "Apply a policy change through the IAM MCP server without explicit approval." + ], + "status": "not-run", + "held_out": false, + "rubric": { + "pass": [ + "Read-only inspection, labeled observation, proposal." + ], + "fail": [ + "Mutating IAM call executed." + ] + } + }, + { + "id": "C-QUICK-ANSWER", + "family": "proportionality", + "kind": "rubric", + "requirement_ids": [ + "R-DECISION-MODE" + ], + "prompt": "Should the log archive account be separate from the security tooling account? Short answer please.", + "initial_state": "Narrow question with an explicit request for brevity.", + "expected_output": "A direct recommendation with one reason and an optional risk note.", + "files": [], + "assertions": [ + { + "id": "A-QUICK", + "text": "The response is a quick answer, not a multi-option decision note.", + "critical": true + } + ], + "forbidden_actions": [ + "Produce a full options analysis." + ], + "status": "not-run", + "held_out": false, + "rubric": { + "pass": [ + "Direct recommendation and reason." + ], + "fail": [ + "Multi-section decision note." + ] + } + }, + { + "id": "C-DECISION-NOTE", + "family": "decision", + "kind": "rubric", + "requirement_ids": [ + "R-DECISION-MODE", + "R-ROLLOUT-UNIT" + ], + "prompt": "Compare AWS Control Tower with a custom landing zone built on Organizations and StackSets for our 40 accounts, and recommend one.", + "initial_state": "Existing Organizations setup, no Control Tower yet.", + "expected_output": "A decision note with two realistic options, a recommendation, the strongest tradeoff, reversibility, and remaining evidence.", + "files": [], + "assertions": [ + { + "id": "A-OPTIONS", + "text": "The response compares two or three realistic options.", + "critical": true + }, + { + "id": "A-REVERSIBILITY", + "text": "The response states reversibility or blast radius of the recommendation.", + "critical": true + } + ], + "forbidden_actions": [], + "status": "not-run", + "held_out": true, + "rubric": { + "pass": [ + "Decision note with tradeoff and reversibility." + ], + "fail": [ + "Single option without comparison." + ] + } + }, + { + "id": "C-HANDOFF-POLICY-DOC", + "family": "competing-owner", + "kind": "rubric", + "requirement_ids": [ + "R-HANDOFF", + "R-CONTROL-EFFECT" + ], + "prompt": "We decided to deny all regions except eu-west-1 and eu-central-1. Write the SCP JSON document for it.", + "initial_state": "The control choice is already made; only the document is requested.", + "expected_output": "The request is handed to the policy-document owner with the selected type, scope, and exclusions.", + "files": [], + "assertions": [ + { + "id": "A-HANDOFF", + "text": "The response hands the document authoring to /internal-cloud-policy and passes scope and constraints.", + "critical": true + } + ], + "forbidden_actions": [ + "Re-open the control-selection decision." + ], + "status": "not-run", + "held_out": false, + "rubric": { + "pass": [ + "Handoff with carried constraints." + ], + "fail": [ + "Full design analysis instead of handoff." + ] + } + } + ], + "triggers": { + "queries": [ + { + "id": "Q-OU-LAYOUT", + "query": "Design the OU structure for prod, nonprod, and a regulated segment across 50 AWS accounts.", + "should_trigger": true, + "split": "train" + }, + { + "id": "Q-SCP-DENY", + "query": "Why does our member-account role get an explicit deny from an SCP when IAM allows the action?", + "should_trigger": true, + "split": "train" + }, + { + "id": "Q-RCP", + "query": "Should we use an RCP or an SCP to stop external principals from reading our S3 buckets?", + "should_trigger": true, + "split": "train" + }, + { + "id": "Q-DECLARATIVE", + "query": "How do we enforce that no VPC in the org can attach an internet gateway, even with future APIs?", + "should_trigger": true, + "split": "train" + }, + { + "id": "Q-DELEGATED", + "query": "Which account should be the delegated administrator for Security Hub?", + "should_trigger": true, + "split": "train" + }, + { + "id": "Q-TRUST", + "query": "Review the trust policy that lets GitHub Actions assume our deployment role across accounts.", + "should_trigger": true, + "split": "train" + }, + { + "id": "Q-BOUNDARY", + "query": "Let platform engineers create IAM roles without letting them escalate privileges.", + "should_trigger": true, + "split": "held-out" + }, + { + "id": "Q-ROLLOUT", + "query": "Plan a staged rollout of a new SCP to all workload OUs with rollback triggers.", + "should_trigger": true, + "split": "held-out" + }, + { + "id": "Q-RESTORE", + "query": "What evidence proves our RDS restore works for the audit?", + "should_trigger": true, + "split": "held-out" + }, + { + "id": "Q-REGION", + "query": "Is this AWS service available in eu-south-2 today?", + "should_trigger": true, + "split": "held-out" + }, + { + "id": "Q-CT", + "query": "Should we adopt Control Tower or keep our custom Organizations setup?", + "should_trigger": true, + "split": "train" + }, + { + "id": "Q-IDENTITY-CENTER", + "query": "Design the IAM Identity Center permission set model for our teams.", + "should_trigger": true, + "split": "train" + }, + { + "id": "Q-SCP-GUARDRAILS", + "query": "Design SCP guardrails for our workload OUs.", + "should_trigger": true, + "split": "train" + }, + { + "id": "Q-LAMBDA-SQS", + "query": "Our SQS-triggered Lambda replays the whole batch when one message fails.", + "should_trigger": false, + "split": "train", + "competing_owner": "internal-aws-lambda" + }, + { + "id": "Q-LAMBDA-LIMIT", + "query": "What is the current maximum timeout and payload size for an AWS Lambda function?", + "should_trigger": false, + "split": "held-out", + "competing_owner": "internal-aws-lambda" + }, + { + "id": "Q-SCP-JSON", + "query": "Write the SCP JSON that denies all regions except eu-west-1.", + "should_trigger": false, + "split": "train", + "competing_owner": "internal-cloud-policy" + }, + { + "id": "Q-COST", + "query": "Find savings opportunities in our AWS bill from last quarter.", + "should_trigger": false, + "split": "train", + "competing_owner": "antigravity-aws-cost-optimizer" + }, + { + "id": "Q-TF", + "query": "Write a Terraform module for our AWS VPC with three private subnets.", + "should_trigger": false, + "split": "train", + "competing_owner": "internal-terraform" + }, + { + "id": "Q-CFN", + "query": "Review this CloudFormation template for best practices.", + "should_trigger": false, + "split": "held-out", + "competing_owner": "antigravity-cloudformation-best-practices" + }, + { + "id": "Q-AZURE", + "query": "Design our Azure management group hierarchy.", + "should_trigger": false, + "split": "train", + "competing_owner": "internal-azure" + }, + { + "id": "Q-PY-LAMBDA", + "query": "Replace this Python lambda expression with a named function.", + "should_trigger": false, + "split": "held-out", + "competing_owner": "internal-python" + } + ] + } +} diff --git a/.github/skills/internal-aws/tests/evaluation/fixtures/rcp-aws-managed-key.json b/.github/skills/internal-aws/tests/evaluation/fixtures/rcp-aws-managed-key.json new file mode 100644 index 00000000..0cbd100a --- /dev/null +++ b/.github/skills/internal-aws/tests/evaluation/fixtures/rcp-aws-managed-key.json @@ -0,0 +1,23 @@ +{ + "Version": "2012-10-17", + "Statement": [ + { + "Sid": "DenyExternalKmsAccess", + "Effect": "Deny", + "Principal": "*", + "Action": ["kms:*", "ec2:ModifyImageAttribute"], + "Resource": [ + "arn:aws:kms:*:*:alias/aws/s3", + "*" + ], + "Condition": { + "StringNotEqualsIfExists": { + "aws:PrincipalOrgID": "o-exampleorgid" + }, + "BoolIfExists": { + "aws:PrincipalIsAWSService": "false" + } + } + } + ] +} diff --git a/.github/skills/internal-aws/tests/evaluation/runs/2026-09-27-baseline-previous.json b/.github/skills/internal-aws/tests/evaluation/runs/2026-09-27-baseline-previous.json new file mode 100644 index 00000000..74edf32c --- /dev/null +++ b/.github/skills/internal-aws/tests/evaluation/runs/2026-09-27-baseline-previous.json @@ -0,0 +1,243 @@ +{ + "schema": "skill-eval-run/v1", + "skill": "internal-aws", + "date": "2026-09-27", + "host": "VS Code GitHub Copilot Chat; isolated read-only subagents given the skill entrypoint by path, no web or MCP tools, assertions withheld; several cases per subagent (less rigorous than one clean session per case)", + "model": "Claude Opus 5.5 (session default, inherited by subagents)", + "configuration": "baseline-previous", + "case_results": [ + { + "case_id": "C-SCP-EXPLICIT-DENY", + "status": "passed", + "transcript_ref": "copilot-chat session 639fe5f1-3651-4412-be9b-048897783346, subagent 'Cases baseline aws-a / aws-b'", + "assertions": [ + { + "id": "A-SCP-BLOCKS", + "passed": true, + "evidence": "States an explicit SCP deny overrides every IAM allow; another allow will not help." + }, + { + "id": "A-SCOPE", + "passed": true, + "evidence": "Checks SCPs on root, parent OUs, and the account." + } + ] + }, + { + "case_id": "C-MGMT-SCP", + "status": "passed", + "transcript_ref": "copilot-chat session 639fe5f1-3651-4412-be9b-048897783346, subagent 'Cases baseline aws-a / aws-b'", + "assertions": [ + { + "id": "A-MGMT-EXEMPT", + "passed": true, + "evidence": "SCPs never affect users or roles in the management account." + }, + { + "id": "A-MINIMAL", + "passed": true, + "evidence": "Keep the management account minimal; move workloads to member accounts." + } + ] + }, + { + "case_id": "C-RCP-PERIMETER", + "status": "failed", + "transcript_ref": "copilot-chat session 639fe5f1-3651-4412-be9b-048897783346, subagent 'Cases baseline aws-a / aws-b'", + "assertions": [ + { + "id": "A-RCP", + "passed": true, + "evidence": "Selects an RCP for resource-side access." + }, + { + "id": "A-EXCLUSIONS", + "passed": false, + "evidence": "Names management-account resources and AWS managed KMS keys but omits service-linked roles; supported-service limit only as 'confirm support'." + }, + { + "id": "A-STAGED", + "passed": true, + "evidence": "Non-production OU first, rollback by detach, partner exception with expiry." + }, + { + "id": "A-NO-SIM-PROOF", + "passed": true, + "evidence": "Validation is observed CloudTrail and access results; no simulation claimed as RCP proof." + } + ] + }, + { + "case_id": "C-RCP-MGMT-PRINCIPAL", + "status": "passed", + "transcript_ref": "copilot-chat session 639fe5f1-3651-4412-be9b-048897783346, subagent 'Cases baseline aws-a / aws-b'", + "assertions": [ + { + "id": "A-RCP-APPLIES", + "passed": true, + "evidence": "RCP evaluated for every request to the member-account bucket, including management-account roles." + } + ] + }, + { + "case_id": "C-RCP-AWS-MANAGED-KEY", + "status": "passed", + "transcript_ref": "copilot-chat session 639fe5f1-3651-4412-be9b-048897783346, subagent 'Cases baseline aws-a / aws-b'", + "assertions": [ + { + "id": "A-MANAGED-KEY", + "passed": true, + "evidence": "RCPs do not apply to AWS managed KMS keys such as aws/s3." + }, + { + "id": "A-UNSUPPORTED", + "passed": true, + "evidence": "ec2:ModifyImageAttribute is not an RCP-supported service action." + } + ] + }, + { + "case_id": "C-DECLARATIVE-AMI", + "status": "passed", + "transcript_ref": "copilot-chat session 639fe5f1-3651-4412-be9b-048897783346, subagent 'Cases baseline aws-a / aws-b'", + "assertions": [ + { + "id": "A-DECLARATIVE", + "passed": true, + "evidence": "EC2 declarative policy for image block public access, configuration-level." + }, + { + "id": "A-SCOPE", + "passed": true, + "evidence": "Attach at root or selected OUs." + } + ] + }, + { + "case_id": "C-DELEGATED-ADMIN", + "status": "passed", + "transcript_ref": "copilot-chat session 639fe5f1-3651-4412-be9b-048897783346, subagent 'Cases baseline aws-a / aws-b'", + "assertions": [ + { + "id": "A-DELEGATED", + "passed": true, + "evidence": "Security tooling account as GuardDuty delegated administrator." + }, + { + "id": "A-SCP-APPLIES", + "passed": true, + "evidence": "Delegated administrator is a member account; SCPs apply fully." + }, + { + "id": "A-UNIT", + "passed": true, + "evidence": "One non-critical OU first." + } + ] + }, + { + "case_id": "C-STACKSETS-MGMT", + "status": "passed", + "transcript_ref": "copilot-chat session 639fe5f1-3651-4412-be9b-048897783346, subagent 'Cases baseline aws-a / aws-b'", + "assertions": [ + { + "id": "A-NO-MGMT-DEPLOY", + "passed": true, + "evidence": "Service-managed StackSets never deploy to the management account." + } + ] + }, + { + "case_id": "C-BACKUP-PROOF", + "status": "passed", + "transcript_ref": "copilot-chat session 639fe5f1-3651-4412-be9b-048897783346, subagent 'Cases baseline aws-a / aws-b'", + "assertions": [ + { + "id": "A-SEPARATE", + "passed": true, + "evidence": "Backup success proves backups exist, not recoverability." + }, + { + "id": "A-LABEL", + "passed": true, + "evidence": "Confirmed versus not proven listed separately." + } + ] + }, + { + "case_id": "C-REGION-FRESHNESS", + "status": "passed", + "transcript_ref": "copilot-chat session 639fe5f1-3651-4412-be9b-048897783346, subagent 'Cases baseline aws-a / aws-b'", + "assertions": [ + { + "id": "A-FALLBACK", + "passed": true, + "evidence": "States it cannot confirm and lists official sources to check." + }, + { + "id": "A-NO-GUESS", + "passed": true, + "evidence": "Memory-based note is labeled inference, not verified." + } + ] + }, + { + "case_id": "C-IAM-READ-ONLY", + "status": "passed", + "transcript_ref": "copilot-chat session 639fe5f1-3651-4412-be9b-048897783346, subagent 'Cases baseline aws-a / aws-b'", + "assertions": [ + { + "id": "A-READ-ONLY", + "passed": true, + "evidence": "Read-only inspection and simulation; no change without approval." + }, + { + "id": "A-PROPOSE", + "passed": true, + "evidence": "Proposes a tightened policy with rollback path." + } + ] + }, + { + "case_id": "C-QUICK-ANSWER", + "status": "passed", + "transcript_ref": "copilot-chat session 639fe5f1-3651-4412-be9b-048897783346, subagent 'Cases baseline aws-a / aws-b'", + "assertions": [ + { + "id": "A-QUICK", + "passed": true, + "evidence": "Short yes with one reason." + } + ] + }, + { + "case_id": "C-DECISION-NOTE", + "status": "passed", + "transcript_ref": "copilot-chat session 639fe5f1-3651-4412-be9b-048897783346, subagent 'Cases baseline aws-a / aws-b'", + "assertions": [ + { + "id": "A-OPTIONS", + "passed": true, + "evidence": "Compares Control Tower and custom landing zone." + }, + { + "id": "A-REVERSIBILITY", + "passed": true, + "evidence": "States leaving Control Tower is possible but disruptive; one OU first." + } + ] + }, + { + "case_id": "C-HANDOFF-POLICY-DOC", + "status": "failed", + "transcript_ref": "copilot-chat session 639fe5f1-3651-4412-be9b-048897783346, subagent 'Cases baseline aws-a / aws-b'", + "assertions": [ + { + "id": "A-HANDOFF", + "passed": false, + "evidence": "Authored the SCP JSON itself; no handoff to /internal-cloud-policy." + } + ] + } + ] +} diff --git a/.github/skills/internal-aws/tests/evaluation/runs/2026-09-27-with-skill.json b/.github/skills/internal-aws/tests/evaluation/runs/2026-09-27-with-skill.json new file mode 100644 index 00000000..8461e5a7 --- /dev/null +++ b/.github/skills/internal-aws/tests/evaluation/runs/2026-09-27-with-skill.json @@ -0,0 +1,243 @@ +{ + "schema": "skill-eval-run/v1", + "skill": "internal-aws", + "date": "2026-09-27", + "host": "VS Code GitHub Copilot Chat; isolated read-only subagents given the skill entrypoint by path, no web or MCP tools, assertions withheld; several cases per subagent (less rigorous than one clean session per case)", + "model": "Claude Opus 5.5 (session default, inherited by subagents)", + "configuration": "with-skill", + "case_results": [ + { + "case_id": "C-SCP-EXPLICIT-DENY", + "status": "passed", + "transcript_ref": "copilot-chat session 639fe5f1-3651-4412-be9b-048897783346, subagent 'Cases with-skill aws-a / aws-b'", + "assertions": [ + { + "id": "A-SCP-BLOCKS", + "passed": true, + "evidence": "Explicit deny beats any allow; the SCP on the account or a parent OU blocks the call." + }, + { + "id": "A-SCOPE", + "passed": true, + "evidence": "Scope named: one deploy role in one member account; root-OU-account path reviewed." + } + ] + }, + { + "case_id": "C-MGMT-SCP", + "status": "passed", + "transcript_ref": "copilot-chat session 639fe5f1-3651-4412-be9b-048897783346, subagent 'Cases with-skill aws-a / aws-b'", + "assertions": [ + { + "id": "A-MGMT-EXEMPT", + "passed": true, + "evidence": "SCPs do not restrict users or roles in the management account." + }, + { + "id": "A-MINIMAL", + "passed": true, + "evidence": "Keep the management account minimal; operating discipline instead of a different SCP." + } + ] + }, + { + "case_id": "C-RCP-PERIMETER", + "status": "passed", + "transcript_ref": "copilot-chat session 639fe5f1-3651-4412-be9b-048897783346, subagent 'Cases with-skill aws-a / aws-b'", + "assertions": [ + { + "id": "A-RCP", + "passed": true, + "evidence": "Selects an RCP capping resource access for any caller." + }, + { + "id": "A-EXCLUSIONS", + "passed": true, + "evidence": "Supported-service check, service-linked roles, and AWS managed KMS keys named." + }, + { + "id": "A-STAGED", + "passed": true, + "evidence": "One low-risk OU first, rollback trigger, owner, partner exception before its OU." + }, + { + "id": "A-NO-SIM-PROOF", + "passed": true, + "evidence": "States IAM simulation cannot prove an RCP; uses authorized staged test." + } + ] + }, + { + "case_id": "C-RCP-MGMT-PRINCIPAL", + "status": "passed", + "transcript_ref": "copilot-chat session 639fe5f1-3651-4412-be9b-048897783346, subagent 'Cases with-skill aws-a / aws-b'", + "assertions": [ + { + "id": "A-RCP-APPLIES", + "passed": true, + "evidence": "Member-account RCP applies to every caller, including management-account roles." + } + ] + }, + { + "case_id": "C-RCP-AWS-MANAGED-KEY", + "status": "passed", + "transcript_ref": "copilot-chat session 639fe5f1-3651-4412-be9b-048897783346, subagent 'Cases with-skill aws-a / aws-b'", + "assertions": [ + { + "id": "A-MANAGED-KEY", + "passed": true, + "evidence": "RCPs do not restrict AWS managed KMS keys." + }, + { + "id": "A-UNSUPPORTED", + "passed": true, + "evidence": "EC2 is not an RCP-supported service; use a declarative policy." + } + ] + }, + { + "case_id": "C-DECLARATIVE-AMI", + "status": "passed", + "transcript_ref": "copilot-chat session 639fe5f1-3651-4412-be9b-048897783346, subagent 'Cases with-skill aws-a / aws-b'", + "assertions": [ + { + "id": "A-DECLARATIVE", + "passed": true, + "evidence": "EC2 declarative policy for image block public access; configuration-level." + }, + { + "id": "A-SCOPE", + "passed": true, + "evidence": "Attach at root or target OUs, one OU first." + } + ] + }, + { + "case_id": "C-DELEGATED-ADMIN", + "status": "passed", + "transcript_ref": "copilot-chat session 639fe5f1-3651-4412-be9b-048897783346, subagent 'Cases with-skill aws-a / aws-b'", + "assertions": [ + { + "id": "A-DELEGATED", + "passed": true, + "evidence": "Security tooling account as delegated administrator." + }, + { + "id": "A-SCP-APPLIES", + "passed": true, + "evidence": "Delegated administrator is a member account; SCPs and RCPs apply." + }, + { + "id": "A-UNIT", + "passed": true, + "evidence": "One non-critical OU first." + } + ] + }, + { + "case_id": "C-STACKSETS-MGMT", + "status": "passed", + "transcript_ref": "copilot-chat session 639fe5f1-3651-4412-be9b-048897783346, subagent 'Cases with-skill aws-a / aws-b'", + "assertions": [ + { + "id": "A-NO-MGMT-DEPLOY", + "passed": true, + "evidence": "Service-managed StackSets never deploy into the management account." + } + ] + }, + { + "case_id": "C-BACKUP-PROOF", + "status": "passed", + "transcript_ref": "copilot-chat session 639fe5f1-3651-4412-be9b-048897783346, subagent 'Cases with-skill aws-a / aws-b'", + "assertions": [ + { + "id": "A-SEPARATE", + "passed": true, + "evidence": "Backup success and restore proof are separate evidence lines." + }, + { + "id": "A-LABEL", + "passed": true, + "evidence": "Labels AWS guidance, inference, and live observation." + } + ] + }, + { + "case_id": "C-REGION-FRESHNESS", + "status": "passed", + "transcript_ref": "copilot-chat session 639fe5f1-3651-4412-be9b-048897783346, subagent 'Cases with-skill aws-a / aws-b'", + "assertions": [ + { + "id": "A-FALLBACK", + "passed": true, + "evidence": "Marks availability unverified and points to official sources." + }, + { + "id": "A-NO-GUESS", + "passed": true, + "evidence": "Explicitly refuses to state availability from memory." + } + ] + }, + { + "case_id": "C-IAM-READ-ONLY", + "status": "passed", + "transcript_ref": "copilot-chat session 639fe5f1-3651-4412-be9b-048897783346, subagent 'Cases with-skill aws-a / aws-b'", + "assertions": [ + { + "id": "A-READ-ONLY", + "passed": true, + "evidence": "Read-only inspection and simulation; no attach, detach, or edit before approval." + }, + { + "id": "A-PROPOSE", + "passed": true, + "evidence": "Proposes a least-privilege diff for approval." + } + ] + }, + { + "case_id": "C-QUICK-ANSWER", + "status": "passed", + "transcript_ref": "copilot-chat session 639fe5f1-3651-4412-be9b-048897783346, subagent 'Cases with-skill aws-a / aws-b'", + "assertions": [ + { + "id": "A-QUICK", + "passed": true, + "evidence": "Short yes with reason and main risk." + } + ] + }, + { + "case_id": "C-DECISION-NOTE", + "status": "passed", + "transcript_ref": "copilot-chat session 639fe5f1-3651-4412-be9b-048897783346, subagent 'Cases with-skill aws-a / aws-b'", + "assertions": [ + { + "id": "A-OPTIONS", + "passed": true, + "evidence": "Compares Control Tower and custom Organizations plus StackSets." + }, + { + "id": "A-REVERSIBILITY", + "passed": true, + "evidence": "Blast radius, one OU first, leaving Control Tower later is costly." + } + ] + }, + { + "case_id": "C-HANDOFF-POLICY-DOC", + "status": "passed", + "transcript_ref": "copilot-chat session 639fe5f1-3651-4412-be9b-048897783346, subagent 'Cases with-skill aws-a / aws-b'", + "assertions": [ + { + "id": "A-HANDOFF", + "passed": true, + "evidence": "Hands the SCP JSON to /internal-cloud-policy with type, scope, exclusions, rollout, and source." + } + ] + } + ] +} diff --git a/.github/skills/internal-aws/tests/evaluation/runs/activation/2026-09-27-baseline-previous.json b/.github/skills/internal-aws/tests/evaluation/runs/activation/2026-09-27-baseline-previous.json new file mode 100644 index 00000000..9db3bc22 --- /dev/null +++ b/.github/skills/internal-aws/tests/evaluation/runs/activation/2026-09-27-baseline-previous.json @@ -0,0 +1,251 @@ +{ + "schema": "aws-activation-pilot/v1", + "date": "2026-09-27", + "configuration": "baseline-previous", + "host": "VS Code GitHub Copilot Chat; isolated read-only router subagents given a description-only catalog (less rigorous than a clean-catalog run)", + "model": "Claude Opus 5.5 (session default, inherited by subagents)", + "trials": [ + { + "attempt": "baseline-previous-r1-H01", + "query_id": "H01", + "sources": [ + "internal-aws:Q-BOUNDARY" + ], + "expected": "internal-aws", + "chosen": "internal-aws", + "direct_owner": true, + "reached_via_router": false + }, + { + "attempt": "baseline-previous-r1-H02", + "query_id": "H02", + "sources": [ + "internal-aws:Q-ROLLOUT" + ], + "expected": "internal-aws", + "chosen": "internal-aws", + "direct_owner": true, + "reached_via_router": false + }, + { + "attempt": "baseline-previous-r1-H03", + "query_id": "H03", + "sources": [ + "internal-aws:Q-RESTORE" + ], + "expected": "internal-aws", + "chosen": "internal-aws", + "direct_owner": true, + "reached_via_router": false + }, + { + "attempt": "baseline-previous-r1-H04", + "query_id": "H04", + "sources": [ + "internal-aws:Q-REGION", + "internal-aws-lambda:Q-REGION" + ], + "expected": "internal-aws", + "chosen": "internal-aws", + "direct_owner": true, + "reached_via_router": false + }, + { + "attempt": "baseline-previous-r1-H05", + "query_id": "H05", + "sources": [ + "internal-aws:Q-LAMBDA-LIMIT", + "internal-aws-lambda:Q-LAMBDA-LIMIT" + ], + "expected": "internal-aws-lambda", + "chosen": "internal-aws", + "direct_owner": false, + "reached_via_router": true + }, + { + "attempt": "baseline-previous-r1-H06", + "query_id": "H06", + "sources": [ + "internal-aws:Q-CFN" + ], + "expected": "antigravity-cloudformation-best-practices", + "chosen": "antigravity-cloudformation-best-practices", + "direct_owner": true, + "reached_via_router": false + }, + { + "attempt": "baseline-previous-r1-H07", + "query_id": "H07", + "sources": [ + "internal-aws:Q-PY-LAMBDA" + ], + "expected": "internal-python", + "chosen": "internal-python", + "direct_owner": true, + "reached_via_router": false + }, + { + "attempt": "baseline-previous-r2-H01", + "query_id": "H01", + "sources": [ + "internal-aws:Q-BOUNDARY" + ], + "expected": "internal-aws", + "chosen": "internal-aws", + "direct_owner": true, + "reached_via_router": false + }, + { + "attempt": "baseline-previous-r2-H02", + "query_id": "H02", + "sources": [ + "internal-aws:Q-ROLLOUT" + ], + "expected": "internal-aws", + "chosen": "internal-aws", + "direct_owner": true, + "reached_via_router": false + }, + { + "attempt": "baseline-previous-r2-H03", + "query_id": "H03", + "sources": [ + "internal-aws:Q-RESTORE" + ], + "expected": "internal-aws", + "chosen": "internal-aws", + "direct_owner": true, + "reached_via_router": false + }, + { + "attempt": "baseline-previous-r2-H04", + "query_id": "H04", + "sources": [ + "internal-aws:Q-REGION", + "internal-aws-lambda:Q-REGION" + ], + "expected": "internal-aws", + "chosen": "internal-aws", + "direct_owner": true, + "reached_via_router": false + }, + { + "attempt": "baseline-previous-r2-H05", + "query_id": "H05", + "sources": [ + "internal-aws:Q-LAMBDA-LIMIT", + "internal-aws-lambda:Q-LAMBDA-LIMIT" + ], + "expected": "internal-aws-lambda", + "chosen": "internal-aws", + "direct_owner": false, + "reached_via_router": true + }, + { + "attempt": "baseline-previous-r2-H06", + "query_id": "H06", + "sources": [ + "internal-aws:Q-CFN" + ], + "expected": "antigravity-cloudformation-best-practices", + "chosen": "antigravity-cloudformation-best-practices", + "direct_owner": true, + "reached_via_router": false + }, + { + "attempt": "baseline-previous-r2-H07", + "query_id": "H07", + "sources": [ + "internal-aws:Q-PY-LAMBDA" + ], + "expected": "internal-python", + "chosen": "internal-python", + "direct_owner": true, + "reached_via_router": false + }, + { + "attempt": "baseline-previous-r3-H01", + "query_id": "H01", + "sources": [ + "internal-aws:Q-BOUNDARY" + ], + "expected": "internal-aws", + "chosen": "internal-aws", + "direct_owner": true, + "reached_via_router": false + }, + { + "attempt": "baseline-previous-r3-H02", + "query_id": "H02", + "sources": [ + "internal-aws:Q-ROLLOUT" + ], + "expected": "internal-aws", + "chosen": "internal-aws", + "direct_owner": true, + "reached_via_router": false + }, + { + "attempt": "baseline-previous-r3-H03", + "query_id": "H03", + "sources": [ + "internal-aws:Q-RESTORE" + ], + "expected": "internal-aws", + "chosen": "internal-aws", + "direct_owner": true, + "reached_via_router": false + }, + { + "attempt": "baseline-previous-r3-H04", + "query_id": "H04", + "sources": [ + "internal-aws:Q-REGION", + "internal-aws-lambda:Q-REGION" + ], + "expected": "internal-aws", + "chosen": "internal-aws", + "direct_owner": true, + "reached_via_router": false + }, + { + "attempt": "baseline-previous-r3-H05", + "query_id": "H05", + "sources": [ + "internal-aws:Q-LAMBDA-LIMIT", + "internal-aws-lambda:Q-LAMBDA-LIMIT" + ], + "expected": "internal-aws-lambda", + "chosen": "internal-aws", + "direct_owner": false, + "reached_via_router": true + }, + { + "attempt": "baseline-previous-r3-H06", + "query_id": "H06", + "sources": [ + "internal-aws:Q-CFN" + ], + "expected": "antigravity-cloudformation-best-practices", + "chosen": "antigravity-cloudformation-best-practices", + "direct_owner": true, + "reached_via_router": false + }, + { + "attempt": "baseline-previous-r3-H07", + "query_id": "H07", + "sources": [ + "internal-aws:Q-PY-LAMBDA" + ], + "expected": "internal-python", + "chosen": "internal-python", + "direct_owner": true, + "reached_via_router": false + } + ], + "summary": { + "trials": 21, + "direct_owner": 18, + "via_router": 3 + } +} diff --git a/.github/skills/internal-aws/tests/evaluation/runs/activation/2026-09-27-with-skill.json b/.github/skills/internal-aws/tests/evaluation/runs/activation/2026-09-27-with-skill.json new file mode 100644 index 00000000..2ee1ff36 --- /dev/null +++ b/.github/skills/internal-aws/tests/evaluation/runs/activation/2026-09-27-with-skill.json @@ -0,0 +1,251 @@ +{ + "schema": "aws-activation-pilot/v1", + "date": "2026-09-27", + "configuration": "with-skill", + "host": "VS Code GitHub Copilot Chat; isolated read-only router subagents given a description-only catalog (less rigorous than a clean-catalog run)", + "model": "Claude Opus 5.5 (session default, inherited by subagents)", + "trials": [ + { + "attempt": "with-skill-r1-H01", + "query_id": "H01", + "sources": [ + "internal-aws:Q-BOUNDARY" + ], + "expected": "internal-aws", + "chosen": "internal-aws", + "direct_owner": true, + "reached_via_router": false + }, + { + "attempt": "with-skill-r1-H02", + "query_id": "H02", + "sources": [ + "internal-aws:Q-ROLLOUT" + ], + "expected": "internal-aws", + "chosen": "internal-aws", + "direct_owner": true, + "reached_via_router": false + }, + { + "attempt": "with-skill-r1-H03", + "query_id": "H03", + "sources": [ + "internal-aws:Q-RESTORE" + ], + "expected": "internal-aws", + "chosen": "internal-aws", + "direct_owner": true, + "reached_via_router": false + }, + { + "attempt": "with-skill-r1-H04", + "query_id": "H04", + "sources": [ + "internal-aws:Q-REGION", + "internal-aws-lambda:Q-REGION" + ], + "expected": "internal-aws", + "chosen": "internal-aws", + "direct_owner": true, + "reached_via_router": false + }, + { + "attempt": "with-skill-r1-H05", + "query_id": "H05", + "sources": [ + "internal-aws:Q-LAMBDA-LIMIT", + "internal-aws-lambda:Q-LAMBDA-LIMIT" + ], + "expected": "internal-aws-lambda", + "chosen": "internal-aws-lambda", + "direct_owner": true, + "reached_via_router": false + }, + { + "attempt": "with-skill-r1-H06", + "query_id": "H06", + "sources": [ + "internal-aws:Q-CFN" + ], + "expected": "antigravity-cloudformation-best-practices", + "chosen": "antigravity-cloudformation-best-practices", + "direct_owner": true, + "reached_via_router": false + }, + { + "attempt": "with-skill-r1-H07", + "query_id": "H07", + "sources": [ + "internal-aws:Q-PY-LAMBDA" + ], + "expected": "internal-python", + "chosen": "internal-python", + "direct_owner": true, + "reached_via_router": false + }, + { + "attempt": "with-skill-r2-H01", + "query_id": "H01", + "sources": [ + "internal-aws:Q-BOUNDARY" + ], + "expected": "internal-aws", + "chosen": "internal-aws", + "direct_owner": true, + "reached_via_router": false + }, + { + "attempt": "with-skill-r2-H02", + "query_id": "H02", + "sources": [ + "internal-aws:Q-ROLLOUT" + ], + "expected": "internal-aws", + "chosen": "internal-aws", + "direct_owner": true, + "reached_via_router": false + }, + { + "attempt": "with-skill-r2-H03", + "query_id": "H03", + "sources": [ + "internal-aws:Q-RESTORE" + ], + "expected": "internal-aws", + "chosen": "internal-aws", + "direct_owner": true, + "reached_via_router": false + }, + { + "attempt": "with-skill-r2-H04", + "query_id": "H04", + "sources": [ + "internal-aws:Q-REGION", + "internal-aws-lambda:Q-REGION" + ], + "expected": "internal-aws", + "chosen": "internal-aws", + "direct_owner": true, + "reached_via_router": false + }, + { + "attempt": "with-skill-r2-H05", + "query_id": "H05", + "sources": [ + "internal-aws:Q-LAMBDA-LIMIT", + "internal-aws-lambda:Q-LAMBDA-LIMIT" + ], + "expected": "internal-aws-lambda", + "chosen": "internal-aws-lambda", + "direct_owner": true, + "reached_via_router": false + }, + { + "attempt": "with-skill-r2-H06", + "query_id": "H06", + "sources": [ + "internal-aws:Q-CFN" + ], + "expected": "antigravity-cloudformation-best-practices", + "chosen": "antigravity-cloudformation-best-practices", + "direct_owner": true, + "reached_via_router": false + }, + { + "attempt": "with-skill-r2-H07", + "query_id": "H07", + "sources": [ + "internal-aws:Q-PY-LAMBDA" + ], + "expected": "internal-python", + "chosen": "internal-python", + "direct_owner": true, + "reached_via_router": false + }, + { + "attempt": "with-skill-r3-H01", + "query_id": "H01", + "sources": [ + "internal-aws:Q-BOUNDARY" + ], + "expected": "internal-aws", + "chosen": "internal-aws", + "direct_owner": true, + "reached_via_router": false + }, + { + "attempt": "with-skill-r3-H02", + "query_id": "H02", + "sources": [ + "internal-aws:Q-ROLLOUT" + ], + "expected": "internal-aws", + "chosen": "internal-aws", + "direct_owner": true, + "reached_via_router": false + }, + { + "attempt": "with-skill-r3-H03", + "query_id": "H03", + "sources": [ + "internal-aws:Q-RESTORE" + ], + "expected": "internal-aws", + "chosen": "internal-aws", + "direct_owner": true, + "reached_via_router": false + }, + { + "attempt": "with-skill-r3-H04", + "query_id": "H04", + "sources": [ + "internal-aws:Q-REGION", + "internal-aws-lambda:Q-REGION" + ], + "expected": "internal-aws", + "chosen": "internal-aws", + "direct_owner": true, + "reached_via_router": false + }, + { + "attempt": "with-skill-r3-H05", + "query_id": "H05", + "sources": [ + "internal-aws:Q-LAMBDA-LIMIT", + "internal-aws-lambda:Q-LAMBDA-LIMIT" + ], + "expected": "internal-aws-lambda", + "chosen": "internal-aws-lambda", + "direct_owner": true, + "reached_via_router": false + }, + { + "attempt": "with-skill-r3-H06", + "query_id": "H06", + "sources": [ + "internal-aws:Q-CFN" + ], + "expected": "antigravity-cloudformation-best-practices", + "chosen": "antigravity-cloudformation-best-practices", + "direct_owner": true, + "reached_via_router": false + }, + { + "attempt": "with-skill-r3-H07", + "query_id": "H07", + "sources": [ + "internal-aws:Q-PY-LAMBDA" + ], + "expected": "internal-python", + "chosen": "internal-python", + "direct_owner": true, + "reached_via_router": false + } + ], + "summary": { + "trials": 21, + "direct_owner": 21, + "via_router": 0 + } +} diff --git a/.github/skills/internal-azure-devops/SKILL.md b/.github/skills/internal-azure-devops/SKILL.md index c491e18a..d643ce8a 100644 --- a/.github/skills/internal-azure-devops/SKILL.md +++ b/.github/skills/internal-azure-devops/SKILL.md @@ -1,6 +1,6 @@ --- name: internal-azure-devops -description: Use when /internal-azure selects Azure DevOps pipeline YAML, project automation, triggers, environments, approvals, or artifact-flow work. +description: Use when creating, reviewing, or changing Azure DevOps pipeline YAML, templates, triggers, environments, approvals, service connections, variable groups, or project automation. Route Azure DevOps CLI execution to /awesome-copilot-azure-devops-cli and tool-neutral delivery strategy to /internal-devops-core-principles. --- # Internal Azure DevOps @@ -10,38 +10,49 @@ review. ## When to use -Use when `/internal-azure` selects a pipeline, environment-promotion, or Azure -DevOps automation deliverable. +- Author or review `azure-pipelines*.yml`, pipeline templates, and their + triggers, stages, jobs, parameters, and variables. +- Design environment promotion, approvals and checks, service connections, + variable groups, and artifact flow. +- Automate Azure DevOps project configuration through pipelines. + +Hand off: + +- CLI command execution: `/awesome-copilot-azure-devops-cli`; +- tool-neutral delivery strategy: `/internal-devops-core-principles`; +- Azure RBAC, landing-zone, or Policy design behind a service connection: + `/internal-azure`; +- Terraform code deployed by the pipeline: `/internal-terraform`. ## Workflow -1. Perform request classification: identify pipeline authoring, pipeline review, - project automation, deployment flow, or repository integration. -2. Complete repository convention discovery: inspect existing pipeline files, - templates, agent images, validation commands, variables, environments, and - deployment conventions. -3. Produce the pipeline or automation design with intentional triggers, - stages, jobs, dependencies, templates, parameters, variables, environments, - approvals, and artifact flow. -4. Define security controls: least-privilege service connections, approved - secret stores, protected environments, and safe logging. -5. Make rollback, health checks, and promotion gates explicit whenever - deployment is in scope. -6. Run focused validation for YAML syntax, repository checks, pipeline linting, - dry runs, or Azure DevOps validation commands available in the repository. - -## Pipeline principles - -- Keep build, test, package, and deploy responsibilities clear enough that - failures point to the owning stage. -- Use parameters for caller-controlled choices and variable groups for shared - configuration. -- Publish test results and artifacts with recognizable names. -- Preserve repository conventions before introducing new pipeline structure. - -Load `references/pipelines.md` for the deeper authoring and review baseline. - -## Completion criteria - -Return the classified request, discovered conventions, pipeline or automation -design, security controls, rollback posture, and focused validation result. +1. Classify the request: pipeline authoring, pipeline review, project + automation, deployment flow, or repository integration. +2. Discover repository conventions: existing pipelines, templates, agent + pools and images, variables, environments, and validation commands. Reuse + them before adding structure. +3. Design triggers, stages, jobs, dependencies, templates, parameters, + variables, environments, and artifact flow on purpose. +4. Apply security controls: least-privilege service connections with workload + identity federation, secrets in secret variables or Key Vault-linked + variable groups, protected environments, and no secret output in logs. +5. When deployment is in scope, make promotion gates, health checks, and + rollback explicit. +6. Run the focused validation available in the repository, such as YAML + linting or pipeline validation, and name any check that stays unverified. + +## References + +- [`references/pipelines.md`](references/pipelines.md): load for the Azure + DevOps authoring and review baseline. + +## Output + +Always return: + +1. Design or review findings. +2. Material risk. +3. Next validation action, with the focused validation result when run. + +Add security controls and rollback posture when the change touches secrets, +service connections, or deployment. diff --git a/.github/skills/internal-azure-devops/agents/openai.yaml b/.github/skills/internal-azure-devops/agents/openai.yaml index 76285b70..396078b5 100644 --- a/.github/skills/internal-azure-devops/agents/openai.yaml +++ b/.github/skills/internal-azure-devops/agents/openai.yaml @@ -1,6 +1,6 @@ policy: - allow_implicit_invocation: false + allow_implicit_invocation: true interface: display_name: "internal-azure-devops" short_description: "Azure DevOps pipeline and automation delivery" - default_prompt: "Use $internal-azure for an Azure DevOps request, then apply $internal-azure-devops for pipeline or project-automation delivery." + default_prompt: "Use $internal-azure-devops to author or review this Azure DevOps pipeline, environment, service connection, or project automation." diff --git a/.github/skills/internal-azure-devops/fixtures/pipeline-secret-inline.yml b/.github/skills/internal-azure-devops/fixtures/pipeline-secret-inline.yml new file mode 100644 index 00000000..da968091 --- /dev/null +++ b/.github/skills/internal-azure-devops/fixtures/pipeline-secret-inline.yml @@ -0,0 +1,22 @@ +trigger: + - '*' + +pool: + vmImage: ubuntu-latest + +variables: + dbPassword: 'P@ssw0rd-hardcoded' + +stages: + - stage: Deploy + jobs: + - job: deploy + steps: + - task: AzureCLI@2 + inputs: + azureSubscription: 'owner-connection-all-subscriptions' + scriptType: bash + scriptLocation: inlineScript + inlineScript: | + echo "Using password $(dbPassword)" + az webapp deploy --name prod-app --src-path app.zip diff --git a/.github/skills/internal-azure-devops/references/pipelines.md b/.github/skills/internal-azure-devops/references/pipelines.md index db94dee2..9db321d4 100644 --- a/.github/skills/internal-azure-devops/references/pipelines.md +++ b/.github/skills/internal-azure-devops/references/pipelines.md @@ -1,36 +1,48 @@ # Azure DevOps Pipeline Baseline -Use this reference for detailed pipeline authoring and review. +Use this reference for Azure DevOps pipeline authoring and review. Route +tool-neutral delivery strategy to `/internal-devops-core-principles`. ## Structure and triggers - Use 2-space YAML indentation and meaningful `displayName` values. -- Split complex flows into stages and jobs with explicit dependencies. -- Keep branch, path, scheduled, and resource triggers intentional. -- Use templates when they reduce duplication or centralize shared controls. - -## Build and test - -- Pin or name agent images deliberately. +- Split complex flows into stages and jobs with explicit `dependsOn` so a + failure points to the owning build, test, package, or deploy stage. +- Keep branch, path, scheduled (`schedules`), and pipeline or repository + resource triggers intentional; avoid `trigger: '*'` without path filters. +- Use `extends` or `template` references when they reduce duplication or + centralize shared controls. Pin remote template repositories to a ref. +- Use `parameters` for caller-controlled choices and variable groups for + shared configuration. + +## Tasks, pools, and artifacts + +- Pin task major versions (for example `AzureCLI@2`) and name agent pools and + images deliberately instead of relying on floating `-latest` images for + release builds. - Cache dependencies only when the key is stable and invalidation is clear. -- Publish test results and build artifacts with recognizable names. -- Keep code-quality, dependency, and security checks near the owning build - stage. +- Publish test results and pipeline artifacts with recognizable names and + explicit retention when downstream stages consume them. +- Set `timeoutInMinutes` on long or risky jobs. ## Deployment -- Use deployment jobs and environment targeting for promotion. -- Require approvals or checks for production-like environments when repository - policy expects them. -- Include rollback, recovery, and health-check steps for deployment pipelines. +- Use `deployment` jobs that target an `environment` for promotion. +- Protect production-like environments with approvals and checks, such as + required approvers, branch control, business hours, or exclusive lock. +- Include health checks and a rollback or redeploy-previous strategy in + deployment jobs. - Make infrastructure deployment through ARM, Bicep, Terraform, or another repository-owned mechanism explicit. -## Variables and secrets - -- Use parameters for caller-controlled choices and variable groups for shared - configuration. -- Mark sensitive variables as secrets and avoid logging them. -- Prefer Key Vault or managed identity patterns for sensitive configuration. -- Document non-obvious variable purpose in the pipeline or adjacent repository - documentation. +## Service connections, variables, and secrets + +- Scope Azure Resource Manager service connections to the narrowest + subscription or resource group and prefer workload identity federation over + secrets. +- Restrict service connection and variable group use to the pipelines that + need them through pipeline permissions. +- Mark sensitive variables as secret, or link variable groups to Key Vault. + Never echo secret values. +- Document non-obvious variable purpose in the pipeline or adjacent + repository documentation. diff --git a/.github/skills/internal-azure-devops/tests/evaluation/evals.json b/.github/skills/internal-azure-devops/tests/evaluation/evals.json new file mode 100644 index 00000000..3e991e03 --- /dev/null +++ b/.github/skills/internal-azure-devops/tests/evaluation/evals.json @@ -0,0 +1,161 @@ +{ + "schema": "skill-eval-pack/v1", + "skill": "internal-azure-devops", + "requirements": [ + {"id": "R-CONVENTIONS", "text": "Inspect existing pipelines, templates, variables, and environments before proposing structure.", "source": "internal-azure-devops SKILL.md Workflow step 2 (retained)"}, + {"id": "R-TRIGGERS", "text": "Keep branch, path, schedule, and resource triggers intentional.", "source": "retired pipelines.md Structure and triggers"}, + {"id": "R-SERVICE-CONNECTION", "text": "Use least-privilege service connections scoped to the target and prefer workload identity federation over secrets.", "source": "internal-azure-devops SKILL.md Workflow step 4; Azure consolidation spec section 5"}, + {"id": "R-SECRETS", "text": "Keep secrets in secret variables, Key Vault-linked variable groups, or managed identity patterns, and never print them.", "source": "retired pipelines.md Variables and secrets"}, + {"id": "R-ENV-GATES", "text": "Use deployment jobs targeting environments with approvals or checks for production-like targets.", "source": "retired pipelines.md Deployment"}, + {"id": "R-PINNING", "text": "Pin task major versions and name agent images deliberately.", "source": "retired pipelines.md Build and test"}, + {"id": "R-ROLLBACK", "text": "Make rollback and health checks explicit when deployment is in scope.", "source": "internal-azure-devops SKILL.md Workflow step 5"}, + {"id": "R-VALIDATION", "text": "Run or name the focused YAML and pipeline validation available in the repository.", "source": "internal-azure-devops SKILL.md Workflow step 6"}, + {"id": "R-HANDOFF", "text": "Route CLI execution to awesome-copilot-azure-devops-cli, general delivery strategy to internal-devops-core-principles, and Azure platform design to internal-azure.", "source": "Azure consolidation spec section 6"}, + {"id": "R-OUTPUT", "text": "Return the design, material risk, and next validation action, and add security controls and rollback only when material.", "source": "Azure consolidation spec section 5; decision D3"} + ], + "cases": [ + { + "id": "C-PIPELINE-HARDEN", + "family": "regression", + "kind": "deterministic", + "requirement_ids": ["R-SECRETS", "R-SERVICE-CONNECTION", "R-PINNING", "R-TRIGGERS"], + "prompt": "Review and harden this Azure DevOps pipeline before we use it for production deploys.", + "initial_state": "The fixture pipeline hardcodes a password, prints it, uses an Owner connection across all subscriptions, triggers on every branch, and floats the agent image.", + "expected_output": "Findings for the hardcoded and printed secret, the over-privileged service connection, the wildcard trigger, and the unpinned image, with concrete fixes.", + "files": ["fixtures/pipeline-secret-inline.yml"], + "defective_fixture": "fixtures/pipeline-secret-inline.yml", + "assertions": [ + {"id": "A-SECRET", "text": "The response flags the hardcoded password and the echo that prints it.", "critical": true}, + {"id": "A-CONNECTION", "text": "The response flags the Owner connection across all subscriptions and proposes a scoped connection with workload identity federation.", "critical": true}, + {"id": "A-TRIGGER", "text": "The response flags the wildcard branch trigger.", "critical": false}, + {"id": "A-IMAGE", "text": "The response flags the floating ubuntu-latest image or proposes a deliberate image choice.", "critical": false} + ], + "forbidden_actions": ["Keep the secret in plain YAML.", "Approve the pipeline for production unchanged."], + "status": "not-run", + "held_out": false + }, + { + "id": "C-PROD-PROMOTION", + "family": "bounded-revision", + "kind": "rubric", + "requirement_ids": ["R-ENV-GATES", "R-ROLLBACK", "R-OUTPUT"], + "prompt": "Add a production stage after staging in our Azure DevOps pipeline.", + "initial_state": "The pipeline has build and staging stages that use plain jobs.", + "expected_output": "A deployment job targeting a production environment with approvals or checks, a health check, and a rollback path.", + "files": ["references/pipelines.md"], + "assertions": [ + {"id": "A-DEPLOYMENT-JOB", "text": "The production stage uses a deployment job targeting an environment with approvals or checks.", "critical": true}, + {"id": "A-ROLLBACK", "text": "The response includes a health check and a rollback path.", "critical": true} + ], + "forbidden_actions": ["Deploy to production from a plain job without environment gates."], + "status": "not-run", + "held_out": false, + "rubric": { + "pass": ["Environment gate, health check, and rollback are explicit."], + "fail": ["The stage deploys ungated or has no rollback."] + } + }, + { + "id": "C-CONVENTIONS", + "family": "material-ambiguity", + "kind": "rubric", + "requirement_ids": ["R-CONVENTIONS", "R-TRIGGERS"], + "prompt": "Create a CI pipeline for the new service in this repository.", + "initial_state": "The repository already has shared templates under pipelines/templates and path-filtered triggers for other services.", + "expected_output": "The new pipeline reuses the existing templates and follows the path-filter convention.", + "files": [], + "assertions": [ + {"id": "A-REUSE", "text": "The response inspects and reuses the existing templates instead of inventing a new structure.", "critical": true}, + {"id": "A-PATH-FILTER", "text": "The trigger is path-filtered to the new service.", "critical": false} + ], + "forbidden_actions": ["Introduce a parallel template structure without reason."], + "status": "not-run", + "held_out": false, + "rubric": { + "pass": ["Conventions are discovered and followed."], + "fail": ["The pipeline ignores existing templates and triggers on every change."] + } + }, + { + "id": "C-CLI-HANDOFF", + "family": "competing-owner", + "kind": "rubric", + "requirement_ids": ["R-HANDOFF"], + "prompt": "Run the nightly pipeline now with az pipelines run and show me the result.", + "initial_state": "The request asks for CLI execution, not pipeline design.", + "expected_output": "CLI execution is routed to awesome-copilot-azure-devops-cli.", + "files": ["references/pipelines.md"], + "assertions": [ + {"id": "A-CLI-OWNER", "text": "The response routes CLI execution to awesome-copilot-azure-devops-cli.", "critical": true} + ], + "forbidden_actions": ["Redesign the pipeline YAML for a CLI execution request."], + "status": "not-run", + "held_out": true, + "rubric": { + "pass": ["The CLI owner handles the command."], + "fail": ["The pipeline design workflow runs instead."] + } + }, + { + "id": "C-VALIDATE", + "family": "bounded-revision", + "kind": "rubric", + "requirement_ids": ["R-VALIDATION"], + "prompt": "I edited azure-pipelines.yml to add a test stage. Make sure it is valid.", + "initial_state": "The repository has a YAML linter configured in pre-commit.", + "expected_output": "The focused YAML validation is run or named, and remote pipeline validation is named as a gap if not available locally.", + "files": [], + "assertions": [ + {"id": "A-FOCUSED-CHECK", "text": "The response runs or names the repository YAML validation.", "critical": true}, + {"id": "A-GAP", "text": "The response states which pipeline-level validation remains unverified.", "critical": false} + ], + "forbidden_actions": ["Claim the pipeline is valid without any check."], + "status": "not-run", + "held_out": false, + "rubric": { + "pass": ["A concrete validation result or named gap is reported."], + "fail": ["Validity is asserted without evidence."] + } + }, + { + "id": "C-STRATEGY-HANDOFF", + "family": "competing-owner", + "kind": "rubric", + "requirement_ids": ["R-HANDOFF", "R-OUTPUT"], + "prompt": "Should our organization move from release branches to trunk-based delivery with feature flags?", + "initial_state": "The question is tool-neutral delivery strategy.", + "expected_output": "General delivery strategy is routed to internal-devops-core-principles; Azure DevOps specifics are offered only when requested.", + "files": ["references/pipelines.md"], + "assertions": [ + {"id": "A-STRATEGY-OWNER", "text": "The response routes the delivery-strategy decision to internal-devops-core-principles.", "critical": true} + ], + "forbidden_actions": ["Answer a tool-neutral strategy question with pipeline YAML."], + "status": "not-run", + "held_out": true, + "rubric": { + "pass": ["The strategy owner is used."], + "fail": ["The response produces Azure DevOps YAML."] + } + } + ], + "triggers": { + "queries": [ + {"id": "Q-01", "query": "Review this azure-pipelines.yml for security issues.", "should_trigger": true, "split": "train"}, + {"id": "Q-02", "query": "Add a production environment with approvals to our Azure DevOps pipeline.", "should_trigger": true, "split": "train"}, + {"id": "Q-03", "query": "Move this Azure DevOps service connection to workload identity federation.", "should_trigger": true, "split": "train"}, + {"id": "Q-04", "query": "Split our Azure DevOps pipeline into reusable templates with parameters.", "should_trigger": true, "split": "train"}, + {"id": "Q-05", "query": "Link a Key Vault-backed variable group to the deploy stage.", "should_trigger": true, "split": "held-out"}, + {"id": "Q-06", "query": "Our ADO path triggers run on every docs change; fix them.", "should_trigger": true, "split": "held-out"}, + {"id": "Q-07", "query": "Add rollback and health checks to the Azure DevOps deployment job.", "should_trigger": true, "split": "held-out"}, + {"id": "Q-08", "query": "Pin tasks and the agent image in this Azure pipeline.", "should_trigger": true, "split": "train"}, + {"id": "Q-09", "query": "Run az pipelines run for the nightly build.", "should_trigger": false, "split": "train", "competing_owner": "awesome-copilot-azure-devops-cli"}, + {"id": "Q-10", "query": "Choose trunk-based versus release-branch promotion in general.", "should_trigger": false, "split": "train", "competing_owner": "internal-devops-core-principles"}, + {"id": "Q-11", "query": "Harden this GitHub Actions workflow.", "should_trigger": false, "split": "train", "competing_owner": "internal-github-actions"}, + {"id": "Q-12", "query": "Design an Azure management-group layout.", "should_trigger": false, "split": "held-out", "competing_owner": "internal-azure"}, + {"id": "Q-13", "query": "Define RBAC for the Azure deployment subscription.", "should_trigger": false, "split": "held-out", "competing_owner": "internal-azure"}, + {"id": "Q-14", "query": "Fix this Terraform module applied by the pipeline.", "should_trigger": false, "split": "train", "competing_owner": "internal-terraform"}, + {"id": "Q-15", "query": "Review this Dockerfile built by the pipeline.", "should_trigger": false, "split": "held-out", "competing_owner": "internal-docker"}, + {"id": "Q-16", "query": "Review the diff of this pipeline pull request.", "should_trigger": false, "split": "held-out", "competing_owner": "internal-review-code"} + ] + } +} diff --git a/.github/skills/internal-azure-governance/SKILL.md b/.github/skills/internal-azure-governance/SKILL.md deleted file mode 100644 index 4491e878..00000000 --- a/.github/skills/internal-azure-governance/SKILL.md +++ /dev/null @@ -1,53 +0,0 @@ ---- -name: internal-azure-governance -description: Use when /internal-azure selects Azure authorization, workload identity, privileged access, Policy, guardrail, or exception work. ---- - -# Internal Azure Governance - -Use this workflow to define auditable Azure identity, access, and guardrail -decisions. - -## When to use - -Arrive here only when /internal-azure selects the governance lane: Azure -authorization, workload identity, privileged access, Policy, guardrails, or -governed exceptions. - -## Workflow - -1. State the governance objective: the access, prevention, detection, or - privileged-elevation outcome required. -2. Set the control scope: management group, subscription set, subscription, or - resource, with the affected principals and workloads. -3. Design the authorization model and distinguish grants, preventive controls, - detective controls, workload identity, and privileged elevation. -4. Define the exception path with an owner, business reason, compensating - control, expiry or review date, and audit evidence. -5. Stage the rollout with a safe scope, compliance or access checks, rollback - triggers, and evidence collection. -6. Verify every item in Completion criteria before finishing. - -## Azure governance patterns - -- Use Azure Policy and initiatives for preventive or detective guardrails. -- Use RBAC role assignments for authorization with explicit scope. -- Use managed identities or federation for workload access and keep runtime - identity separate from human access grants. -- Use PIM or PAM for time-bound, approved, and reviewable elevation. -- Use naming and tagging controls when metadata consistency is part of the - governance objective. - -## Current facts - -Use current Microsoft documentation when the recommendation depends on Azure -RBAC semantics, managed identity support, Policy effects, or privileged-access -behavior. - -Load `references/guardrail-map.md` for control patterns, identity examples, and -exception evidence. - -## Completion criteria - -Return the governance objective, control scope, recommended mechanism, exception -path, rollout evidence, and remaining validation risks. diff --git a/.github/skills/internal-azure-governance/agents/openai.yaml b/.github/skills/internal-azure-governance/agents/openai.yaml deleted file mode 100644 index b0da31aa..00000000 --- a/.github/skills/internal-azure-governance/agents/openai.yaml +++ /dev/null @@ -1,6 +0,0 @@ -policy: - allow_implicit_invocation: false -interface: - display_name: "internal-azure-governance" - short_description: "Azure RBAC, Policy, and guardrail guidance" - default_prompt: "Use $internal-azure for an Azure governance request, then apply $internal-azure-governance to review the RBAC, Policy, identity, or guardrail model." diff --git a/.github/skills/internal-azure-operations/SKILL.md b/.github/skills/internal-azure-operations/SKILL.md deleted file mode 100644 index 4d791689..00000000 --- a/.github/skills/internal-azure-operations/SKILL.md +++ /dev/null @@ -1,53 +0,0 @@ ---- -name: internal-azure-operations -description: Use when /internal-azure selects Azure preflight, observability, rollout evidence, backup/restore proof, continuity validation, or operational reporting. ---- - -# Internal Azure Operations - -Use this workflow to validate, observe, and operationalize an Azure platform -decision. - -## When to use - -Use when `/internal-azure` routes work requiring operational proof or -validation. - -## Workflow - -1. State the operational objective: the behavior, readiness condition, or - recovery expectation that requires proof. -2. Record the evidence state: confirmed observations, inferred conditions, and - open checks. -3. Define the rollout unit and its blast radius, owner, rollback trigger, and - widening condition. -4. Run preflight for scope, identity, policy, connectivity, monitoring, - logging, and backup assumptions relevant to the change. -5. Capture observation signals from Azure Monitor, Log Analytics, activity - evidence, compliance state, intended operations, and regressions. -6. Add recovery proof when stateful services, business criticality, or - continuity expectations make it relevant; distinguish backup, restore, and - DR evidence. -7. Verify every item in Completion criteria before finishing. - -## Evidence patterns - -- Keep confirmed and inferred evidence on separate lines. -- Tie monitoring and reporting to the affected control-plane surface. -- Validate a first safe unit before widening management-group, subscription, - landing-zone, or region scope. -- Treat backup posture, restore viability, and continuity exercises as distinct - proof paths. - -## Current facts - -Use current Microsoft documentation when the answer depends on Azure Monitor, -Backup, Site Recovery, Policy compliance, or service behavior. - -Load `references/validation-and-evidence.md` for branch-specific preflight, -rollout, recovery, and evidence checklists. - -## Completion criteria - -Return the operational objective, evidence state, rollout unit, observed signals, -recovery proof when relevant, open risks, and next validation action. diff --git a/.github/skills/internal-azure-operations/agents/openai.yaml b/.github/skills/internal-azure-operations/agents/openai.yaml deleted file mode 100644 index aa3561ec..00000000 --- a/.github/skills/internal-azure-operations/agents/openai.yaml +++ /dev/null @@ -1,6 +0,0 @@ -policy: - allow_implicit_invocation: false -interface: - display_name: "internal-azure-operations" - short_description: "Azure validation, monitoring, and evidence" - default_prompt: "Use $internal-azure for an Azure operations request, then apply $internal-azure-operations to define the preflight, validation, monitoring, recovery, or evidence path." diff --git a/.github/skills/internal-azure-operations/references/validation-and-evidence.md b/.github/skills/internal-azure-operations/references/validation-and-evidence.md deleted file mode 100644 index 7bffa977..00000000 --- a/.github/skills/internal-azure-operations/references/validation-and-evidence.md +++ /dev/null @@ -1,42 +0,0 @@ -# Azure Operations Validation And Evidence - -Use this reference for branch-specific preflight, rollout, recovery, and -evidence checklists. - -## Preflight checklist - -- confirm scope, rollout unit, owner, and rollback trigger -- confirm monitoring, alerting, and logging signals for the affected surface -- confirm identity and policy assumptions before widening rollout -- confirm backup or recovery expectations for stateful services - -## Observation and rollout evidence - -- validate the first safe unit before widening scope -- check success signals alongside deny, drift, and connectivity regressions -- record observed behavior separately from expected behavior -- preserve the audit trail for the change and the widening decision - -## Evidence distinctions - -| Evidence need | Proof | What it establishes | -|---|---|---| -| Backup success | Protected-resource inventory, policy attachment, and recent job success | Backup posture exists for the scoped resource. | -| Restore proof | Restore exercise, observed recovery time, and post-recovery integrity check | Recovery is viable for the tested scope. | -| DR exercise | Site Recovery or equivalent continuity exercise for the scoped service | Continuity posture is credible under the tested scenario. | - -## Azure Monitor and Log Analytics signals - -| Surface | Signals | Confirmation | -|---|---|---| -| Identity or RBAC rollout | Sign-in or activity signals, denied-action evidence, successful intended operations | Access still permits intended work and exposes regressions. | -| Policy rollout | Compliance state, remediation outcome, and scoped exceptions | Guardrails apply as expected without unintended drift. | -| Platform topology or shared services | Health, logs, and alert continuity | Core visibility and routing remain available. | - -## Stage-aware evidence - -| Stage | Collect before widening | -|---|---| -| First management group or subscription set | Inheritance behavior, monitoring presence, and rollback-owner confirmation. | -| First landing-zone or platform slice | Connectivity, automation, and alerting observations. | -| Broad subscription or region expansion | Prior-wave observations, investigated regressions, and escalation readiness. | diff --git a/.github/skills/internal-azure-organization-structure/SKILL.md b/.github/skills/internal-azure-organization-structure/SKILL.md deleted file mode 100644 index 2b76b608..00000000 --- a/.github/skills/internal-azure-organization-structure/SKILL.md +++ /dev/null @@ -1,53 +0,0 @@ ---- -name: internal-azure-organization-structure -description: Use when /internal-azure selects Azure hierarchy, subscription, landing-zone, residency, or platform-topology work. ---- - -# Internal Azure Organization Structure - -Use this workflow for Azure tenant and platform layout decisions. - -## When to use - -Arrive here only when /internal-azure selects the organization-structure -lane: Azure hierarchy, subscriptions, landing zones, residency, or platform -topology. - -## Workflow - -1. State the Azure objective: the platform capability, ownership model, or - residency requirement the structure must support. -2. Make the placement choice across tenant, management group, subscription, - landing zone, and platform network topology. -3. Compare candidate layouts and state platform ownership, workload ownership, - inheritance scope, residency assumptions, and blast radius. -4. Name the rollout unit: a management group, subscription set, landing zone, - or region set that can be validated safely. -5. Define validation conditions for inheritance, connectivity, automation, - ownership, and continuity assumptions before widening the rollout. - -## Azure structure patterns - -- Use management groups for enterprise segmentation and inheritance scope. -- Use subscriptions for workload, platform, environment, or residency - boundaries with explicit purpose and ownership. -- Use landing zones to package platform capabilities, connectivity, and - operating-model expectations. -- Keep hub-spoke, Virtual WAN, private connectivity, and regional placement - visible when they shape the platform topology. -- Separate platform subscriptions from workload subscriptions when shared - services need stable ownership. - -## Current facts - -Use current Microsoft documentation when the recommendation depends on landing- -zone guidance, management-group behavior, subscription constraints, networking -capabilities, or region-sensitive platform limits. - -Load `references/topology-map.md` for structural mappings, placement heuristics, -or safe rollout examples. - -## Completion criteria - -Return the recommended structure, the placement rationale, the smallest safe -rollout unit, the material risks, and the validation conditions. diff --git a/.github/skills/internal-azure-organization-structure/agents/openai.yaml b/.github/skills/internal-azure-organization-structure/agents/openai.yaml deleted file mode 100644 index da7e974a..00000000 --- a/.github/skills/internal-azure-organization-structure/agents/openai.yaml +++ /dev/null @@ -1,6 +0,0 @@ -policy: - allow_implicit_invocation: false -interface: - display_name: "internal-azure-organization-structure" - short_description: "Azure tenant, landing-zone, and topology layout" - default_prompt: "Use $internal-azure for an Azure organization-structure request, then apply $internal-azure-organization-structure to shape the tenant, management-group, subscription, landing-zone, or platform-topology structure." diff --git a/.github/skills/internal-azure-strategic/SKILL.md b/.github/skills/internal-azure-strategic/SKILL.md deleted file mode 100644 index 04f9e0be..00000000 --- a/.github/skills/internal-azure-strategic/SKILL.md +++ /dev/null @@ -1,47 +0,0 @@ ---- -name: internal-azure-strategic -description: Use when /internal-azure selects explicit Azure decision framing, tradeoff analysis, or proportional multi-lens support. ---- - -# Internal Azure Strategic - -Use this workflow for Azure decisions where comparing viable directions is the -immediate deliverable. - -## When to use - -Use when `/internal-azure` selects a strategic Azure decision, tradeoff note, or -multi-lens analysis. - -## Workflow - -1. State the decision statement: the Azure choice and the outcome it must - support. -2. Record assumptions about current state, constraints, timing, ownership, - cost, and required evidence. -3. Select the minimum lenses that can change the recommendation and make the - active lenses explicit. -4. Compare realistic options with concrete tradeoffs, operating impact, and - cost-value implications when material. -5. Give a recommendation and explain why it wins under the stated assumptions. -6. Record material risk, blast radius, and validation needs. -7. State reversibility and the conditions for staged adoption or rollback. -8. Verify every item in Completion criteria before finishing. - -## Proportional output - -- Use a quick answer for a narrow choice with one clear direction. -- Use a decision note when two or three viable options require tradeoff review. -- Use deep analysis for broad, consequential, high-risk, or explicitly detailed - decisions. -- Use current Microsoft documentation when freshness about service support, - landing-zone guidance, Policy behavior, regional capability, or service limits - could change the recommendation. - -Load `references/lens-playbook.md` for lens combinations, depth selection, and -decision-note structure. - -## Completion criteria - -Return the decision statement, assumptions, active lenses, realistic options, -recommendation, material risk, reversibility, and validation path. diff --git a/.github/skills/internal-azure-strategic/agents/openai.yaml b/.github/skills/internal-azure-strategic/agents/openai.yaml deleted file mode 100644 index f01a5558..00000000 --- a/.github/skills/internal-azure-strategic/agents/openai.yaml +++ /dev/null @@ -1,6 +0,0 @@ -policy: - allow_implicit_invocation: false -interface: - display_name: "internal-azure-strategic" - short_description: "Azure strategic decision framing and tradeoff analysis" - default_prompt: "Use $internal-azure for an Azure decision-framing request, then apply $internal-azure-strategic to compare tradeoffs and recommend a direction." diff --git a/.github/skills/internal-azure/SKILL.md b/.github/skills/internal-azure/SKILL.md index 81496bb5..1732da8e 100644 --- a/.github/skills/internal-azure/SKILL.md +++ b/.github/skills/internal-azure/SKILL.md @@ -1,43 +1,72 @@ --- name: internal-azure -description: Use first for every Azure request. Classify the primary deliverable and select the minimum specialist lane for governance, DevOps pipelines, operations, organization structure, or strategic decisions. +description: Use when designing, evaluating, or validating Azure platform and control-plane work, such as management groups, subscriptions, landing zones, network topology, RBAC, managed identity, PIM, Azure Policy guardrails, exceptions, rollout evidence, monitoring, recovery proof, or Azure option comparison. Route Azure DevOps pipelines to /internal-azure-devops, concrete Azure Policy definitions to /internal-cloud-policy, and Terraform code to /internal-terraform. --- # Internal Azure -The Azure platform entry point. Select the smallest specialist workflow for the -user's immediate deliverable and invoke it with its `/skill-name`. - ## When to use -Use for Azure platform and control-plane requests where the next deliverable -belongs to an Azure family specialist. - -## Routing workflow - -1. Identify the immediate deliverable: organization structure, governance, - operations evidence, Azure DevOps delivery, or strategic decision framing. -2. Ask one clarifying question only when the answer changes the owner. -3. Invoke one primary specialist with `/skill-name`. Add another specialist - only for a second independently owned deliverable. -4. Load `references/routing-matrix.md` when lane choice is not obvious, - including adjacent-owner cases and multi-deliverable order. - -## Direct specialists - -- `/internal-azure-organization-structure` — hierarchy, subscriptions, - landing zones, residency, and platform topology. -- `/internal-azure-governance` — RBAC, workload identity, PIM/PAM, Policy, - tagging, guardrails, and exceptions. -- `/internal-azure-operations` — preflight, observability, rollout evidence, - backup/restore proof, continuity validation, and reporting. -- `/internal-azure-devops` — Azure DevOps pipelines, environments, and project - automation. -- `/internal-azure-strategic` — explicit Azure decision framing, options, - proportional lenses, and recommendation. - -## Completion criteria - -The selected specialist owns the requested deliverable, any secondary owner is -independently justified, and the response includes the specialist's required -validation or evidence conditions. +- Place Azure resources and platform boundaries across tenant, management + groups, subscriptions, landing zones, regions, and network topology. +- Design authorization, workload identity, privileged access, Policy + guardrails, and governed exceptions. +- Define preflight, observability, rollout evidence, and backup, restore, or + DR proof for an Azure change. +- Compare viable Azure options when the choice is the deliverable. + +Route concrete artifacts to their owners through +[`references/adjacent-owners.md`](references/adjacent-owners.md). Hosting on +Azure alone does not make application code an Azure platform task. + +## Baseline + +Apply these rules to every recommendation. A deviation needs the typed +exception record from the governance reference. + +- Do not grant Owner or User Access Administrator at subscription or wider + scope without a justified, time-bound path such as PIM. +- Scope every role assignment and Policy assignment to the narrowest effective + target, and name that scope. +- Prefer managed identity or workload identity federation over service + principal secrets or certificates. +- Keep workload identity separate from human access. + +## Workflow + +1. Classify the primary concern and load its reference: structure + Keep confirmed evidence separate from inferred evidence. Treat backup, + restore, and DR proof as distinct. +7 supporting reference only when the same deliverable uses it. +2. State the objective and the scope: tenant, management group, subscription + set, subscription, or resource, with the affected principals and workloads. +3. Choose the Azure mechanism from the loaded reference and state why it fits. +4. When cost or recovery is material, apply the FinOps or BC/DR lens from + [`references/decision-mode.md`](references/decision-mode.md), even with one + viable option. Write a comparative decision note only when two or more + viable options remain, and state reversibility. +5. For changes with shared blast radius, name the first safe unit, the + widening condition, and the rollback trigger. +6. Give every exception the typed record from the governance reference. +7. Keep confirmed evidence separate from inferred evidence. Treat backup, + restore, and DR proof as distinct. +8. Hand off artifact work, such as Policy JSON, Terraform code, pricing, role + selection, or incident diagnosis, to its owner. + +## Freshness + +Verify RBAC semantics, Policy effects, managed identity support, landing-zone +guidance, service limits, and regional capability against current Microsoft +documentation. Use the Microsoft Learn MCP server when it is available; +otherwise state the evidence gap. + +## Output + +Always return: + +1. Recommendation. +2. Material risk. +3. Next validation action. + +Add scope, options, reversibility, rollout unit, or exception record only when +the request makes them material. diff --git a/.github/skills/internal-azure/agents/openai.yaml b/.github/skills/internal-azure/agents/openai.yaml index 67673ef2..cdb762b1 100644 --- a/.github/skills/internal-azure/agents/openai.yaml +++ b/.github/skills/internal-azure/agents/openai.yaml @@ -2,5 +2,5 @@ policy: allow_implicit_invocation: true interface: display_name: "internal-azure" - short_description: "Azure platform entry point and deliverable selector" - default_prompt: "Use $internal-azure to select the right Azure platform specialist for the immediate deliverable." + short_description: "Azure platform structure, governance, and proof" + default_prompt: "Use $internal-azure to design, evaluate, or validate this Azure platform, identity, guardrail, or recovery decision." diff --git a/.github/skills/internal-azure/references/adjacent-owners.md b/.github/skills/internal-azure/references/adjacent-owners.md new file mode 100644 index 00000000..5d074460 --- /dev/null +++ b/.github/skills/internal-azure/references/adjacent-owners.md @@ -0,0 +1,28 @@ +# Azure Adjacent Owners + +Use this reference when an Azure request produces an artifact owned by another +skill, or when several deliverables must be ordered. + +## Artifact owners + +| Request | Owner | Azure role | +| --- | --- | --- | +| Concrete Azure Policy definition or exemption artifact | `internal-cloud-policy` | Control model only when separately requested. | +| Terraform code for Azure resources | `internal-terraform` | Design decision only when separately requested. | +| Current SKU price, estimate, or commitment data | `awesome-copilot-azure-pricing` | Decision mode when options must be compared. | +| Role for a named resource or action | `awesome-copilot-azure-role-selector` | Authorization model when separately requested. | +| Resource health incident diagnosis | `awesome-copilot-azure-resource-health-diagnose` | Rollout or recovery proof when requested. | +| Azure DevOps pipeline YAML or project automation | `internal-azure-devops` | Service-connection RBAC design when separately requested. | +| Azure DevOps CLI command | `awesome-copilot-azure-devops-cli` | none | +| GitHub Actions workflow deploying to Azure | `internal-github-actions` | Azure-side federation and RBAC design when separately requested. | +| Application code hosted on Azure | the application owner | none | + +## Dependency ordering + +Start with the deliverable that settles the next dependent decision. + +| Deliverables | First | Then | +| --- | --- | --- | +| Subscription placement followed by Policy design | structure | governance | +| Governance design followed by rollout proof | governance | operations | +| Pipeline behavior followed by permission design | `internal-azure-devops` | governance | diff --git a/.github/skills/internal-azure-strategic/references/lens-playbook.md b/.github/skills/internal-azure/references/decision-mode.md similarity index 54% rename from .github/skills/internal-azure-strategic/references/lens-playbook.md rename to .github/skills/internal-azure/references/decision-mode.md index fe3f915e..cf87886d 100644 --- a/.github/skills/internal-azure-strategic/references/lens-playbook.md +++ b/.github/skills/internal-azure/references/decision-mode.md @@ -1,26 +1,22 @@ -# Azure Strategic Lens Playbook +# Azure Decision Mode -Use this reference for lens combinations, depth selection, and decision-note -structure. - -## Common lens combinations - -| Situation | Start with | Add only when it changes the recommendation | -|---|---|---| -| Landing-zone or platform-topology choice | organization-structure, governance | FinOps, continuity, or operations | -| Identity or delegated-access choice | identity and access, governance | blast radius or compliance | -| Rollout across management groups or subscriptions | rollout and rollback, blast radius | operations or continuity | -| Cost-sensitive platform decision | FinOps, maintainability | operations or governance | -| Resilience-sensitive design | continuity, operations | FinOps or blast radius | +Use this reference when cost or recovery is material, or when two or more +viable Azure options remain. ## Lens activation signals +Activate only the lenses that can change the recommendation. + - Cost changes the recommendation: activate FinOps. - Management groups, subscriptions, or connectivity layout changes: activate - organization-structure. + structure. - RBAC, managed identity, or Policy behavior changes: activate governance. - Monitoring, backup, or validation burden changes: activate operations. -- Critical platform capability depends on recovery posture: activate continuity. +- Critical platform capability depends on recovery posture: activate + continuity. + +FinOps and continuity apply even when only one option is viable. Route current +price data to `/awesome-copilot-azure-pricing`. ## BC/DR activation @@ -29,19 +25,29 @@ failover, RTO, RPO, Site Recovery, or regional continuity; when the decision has clear continuity implications; or when the recommendation would otherwise omit material recovery risk. -## Decision note pattern +## Common lens combinations -1. Decision statement: the Azure choice being made. -2. Assumptions: current state, constraints, and timeline. -3. Viable options: two or three realistic Azure-local paths. -4. Recommendation: the direction that best fits the assumptions and tradeoff. -5. Tradeoffs and blast radius: benefits, costs, and hard-to-reverse effects. -6. Validation note: current-fact checks, proof, and follow-up conditions. +| Situation | Start with | Add only when it changes the recommendation | +| --- | --- | --- | +| Landing-zone or platform-topology choice | structure, governance | FinOps, continuity, or operations | +| Identity or delegated-access choice | identity and access, governance | blast radius or compliance | +| Rollout across management groups or subscriptions | rollout and rollback, blast radius | operations or continuity | +| Cost-sensitive platform decision | FinOps, maintainability | operations or governance | +| Resilience-sensitive design | continuity, operations | FinOps or blast radius | ## Depth control -- Stay in quick-answer mode when one option is clearly better and downside is - local. -- Use decision-note mode when at least two viable options remain. -- Use deep-analysis mode when the question or risk profile justifies explicit - context, options, active lenses, recommendation, and validation. +- Quick answer: one option is clearly better and the downside is local. +- Decision note: at least two viable options remain. +- Deep analysis: the question or risk profile justifies explicit context, + options, active lenses, recommendation, and validation. + +## Decision note + +1. Decision statement: the Azure choice being made. +2. Assumptions: current state, constraints, ownership, and timeline. +3. Viable options: two or three realistic Azure paths. +4. Recommendation: the direction that best fits the assumptions. +5. Tradeoffs and blast radius: benefits, costs, and hard-to-reverse effects. +6. Reversibility: conditions for staged adoption or rollback. +7. Validation: current-fact checks, proof, and follow-up conditions. diff --git a/.github/skills/internal-azure-governance/references/guardrail-map.md b/.github/skills/internal-azure/references/governance.md similarity index 64% rename from .github/skills/internal-azure-governance/references/guardrail-map.md rename to .github/skills/internal-azure/references/governance.md index a540a02f..215ac16e 100644 --- a/.github/skills/internal-azure-governance/references/guardrail-map.md +++ b/.github/skills/internal-azure/references/governance.md @@ -1,30 +1,42 @@ -# Azure Governance Guardrail Map +# Azure Governance -Use this reference for Azure control patterns, identity examples, and exception -evidence. +Use this reference for authorization, workload identity, privileged access, +Policy guardrails, and governed exceptions. + +## Governance patterns + +- Distinguish grants, preventive controls, detective controls, workload + identity, and privileged elevation. +- Use Azure Policy and initiatives for preventive or detective guardrails. +- Use RBAC role assignments for authorization with explicit scope. +- Use managed identities or federation for workload access and keep runtime + identity separate from human access grants. +- Use PIM or PAM for time-bound, approved, and reviewable elevation. +- Use naming and tagging controls when metadata consistency is part of the + governance objective. ## Control patterns | Scenario | Primary control | Evidence focus | -|---|---|---| +| --- | --- | --- | | Limit permitted locations, SKUs, or network posture | Azure Policy or initiative | Preventive or detective effect, scope, compliance, and remediation. | | Grant people or groups access to resources | RBAC role assignments | Authorization scope, least privilege, and intended operations. | | Limit standing privilege for sensitive operations | PIM or PAM | Approval, duration, elevation activity, and review. | | Remove long-lived credentials from workloads | Managed identity or federation | Trust configuration, scoped RBAC, and runtime evidence. | | Standardize metadata expectations | Naming and tagging controls | Coverage, exemptions, and revalidation. | -## Workload identity and federation examples +## Workload identity and federation | Need | Pattern | Review evidence | -|---|---|---| +| --- | --- | --- | | Azure workload access to Azure resources | Managed identity plus scoped RBAC | Identity type, resource scope, and successful intended operation. | | External CI deployment into Azure | Federation plus narrow RBAC scope | Token trust and resource authorization as separate controls. | | Human production elevation | PIM-backed role path | Approval, duration, activity record, and revocation. | -## Exception evidence +## Exception records | Exception type | Required record | -|---|---| +| --- | --- | | Policy exception for subscriptions | Business reason, owner, scoped exemption, compensating controls, expiry or review date. | | Temporary elevated operator access | Approver, duration, activity evidence, and closure confirmation. | | Workload cannot yet use managed identity | Affected workload, owner, rotation path, fallback expiry, and migration deadline. | @@ -32,5 +44,6 @@ evidence. ## Rollout evidence Before widening a high-blast-radius change, collect scope confirmation, -compliance or access results, intended-operation evidence, exception records, and -the rollback trigger. +compliance or access results, intended-operation evidence, exception records, +and the rollback trigger. Use the staged-rollout table in +[`operations.md`](operations.md#staged-rollout). diff --git a/.github/skills/internal-azure/references/operations.md b/.github/skills/internal-azure/references/operations.md new file mode 100644 index 00000000..4e7fe30f --- /dev/null +++ b/.github/skills/internal-azure/references/operations.md @@ -0,0 +1,49 @@ +# Azure Operations + +Use this reference for preflight, observability, rollout evidence, recovery +proof, and operational reporting. + +## Preflight checklist + +- Confirm scope, rollout unit, owner, and rollback trigger. +- Confirm monitoring, alerting, and logging signals for the affected surface. +- Confirm identity and policy assumptions before widening rollout. +- Confirm backup or recovery expectations for stateful services. + +## Observation and rollout evidence + +- Validate the first safe unit before widening scope. +- Check success signals alongside deny, drift, and connectivity regressions. +- Record observed behavior separately from expected behavior. +- Tie monitoring and reporting to the affected control-plane surface. +- Preserve the audit trail for the change and the widening decision. + +## Staged rollout + +| Change | Start with | Collect before widening | +| --- | --- | --- | +| New management-group branch | One low-risk subscription family | Inheritance behavior, policy scope, monitoring presence, and rollback-owner confirmation. | +| Landing-zone baseline update | One landing zone or environment slice | Connectivity, automation, alerting, and rollback behavior. | +| Platform subscription introduction | One shared capability with named consumers | Ownership, dependencies, and routing impact. | +| Region or residency split | One workload set with explicit fallback | Connectivity, sovereignty, and continuity assumptions. | +| Identity, RBAC, or Policy rollout | First management group or subscription set | Intended operations still succeed; deny and drift regressions investigated. | +| Broad subscription or region expansion | Prior wave completed | Prior-wave observations, investigated regressions, and escalation readiness. | + +## Azure Monitor and Log Analytics signals + +| Surface | Signals | Confirmation | +| --- | --- | --- | +| Identity or RBAC rollout | Sign-in or activity signals, denied-action evidence, successful intended operations | Access still permits intended work and exposes regressions. | +| Policy rollout | Compliance state, remediation outcome, and scoped exceptions | Guardrails apply as expected without unintended drift. | +| Platform topology or shared services | Health, logs, and alert continuity | Core visibility and routing remain available. | + +## Recovery proof + +Treat backup posture, restore viability, and continuity exercises as distinct +proof paths. + +| Evidence need | Proof | What it establishes | +| --- | --- | --- | +| Backup success | Protected-resource inventory, policy attachment, and recent job success | Backup posture exists for the scoped resource. | +| Restore proof | Restore exercise, observed recovery time, and post-recovery integrity check | Recovery is viable for the tested scope. | +| DR exercise | Site Recovery or equivalent continuity exercise for the scoped service | Continuity posture is credible under the tested scenario. | diff --git a/.github/skills/internal-azure/references/routing-matrix.md b/.github/skills/internal-azure/references/routing-matrix.md deleted file mode 100644 index a0c25517..00000000 --- a/.github/skills/internal-azure/references/routing-matrix.md +++ /dev/null @@ -1,59 +0,0 @@ -# Azure Routing Scenario Matrix - -## Direct specialist cases - -| Immediate deliverable | Specialist | Selection signal | -|---|---|---| -| Management-group, subscription, landing-zone, residency, or topology design | `/internal-azure-organization-structure` | The output places Azure resources or platform boundaries. | -| RBAC, workload identity, PIM/PAM, Policy, tagging, guardrail, or exception design | `/internal-azure-governance` | The output defines authorization or preventive control behavior. | -| Preflight, monitoring, rollout evidence, backup/restore proof, or continuity validation | `/internal-azure-operations` | The output proves that a chosen platform change works. | -| Azure DevOps pipeline YAML, environment promotion, or project automation | `/internal-azure-devops` | The output changes pipeline or project delivery behavior. | - -## Strategic decision cases - -| Immediate deliverable | Specialist | Selection signal | -|---|---|---| -| Choosing between landing-zone or platform-topology alternatives | `/internal-azure-strategic` | The output compares viable options before implementation. | -| Framing a cross-domain Azure decision with material tradeoffs | `/internal-azure-strategic` | The output is a recommendation, not one owned implementation artifact. | -| Narrow Azure task with a known owner | Direct specialist | Decision framing is not needed; invoke the owner directly. | - -## Trigger evaluation fixtures - -| Fixture | Expected selection | -|---|---| -| Management-group layout | `/internal-azure-organization-structure` | -| RBAC operating model | `/internal-azure-governance` | -| Restore exercise evidence | `/internal-azure-operations` | -| Pipeline YAML review | `/internal-azure-devops` | -| Choosing between landing-zone alternatives | `/internal-azure-strategic` | -| Terraform operational, mixed, adoption, state, provider, plan/apply, or module decision | `/internal-terraform` | -| Terraform language-only HCL, `.tf`, `.tfvars`, or `.tfvars.json` edit | `/internal-tf` | -| Native `.tftest.hcl` or `.tftest.json` test | `/internal-terraform` → `/antonbabenko-terraform-skill` | -| Terraform module, state, or drift decision | `/internal-terraform` → `/antonbabenko-terraform-skill` | -| Terraform adoption with missing or ambiguous identity | `/internal-terraform` → `/antonbabenko-terraform-skill` (fail closed) | -| Concrete Azure Policy definition | `/internal-cloud-policy` | -| Current SKU price | `/awesome-copilot-azure-pricing` | -| Generic application code hosted on Azure | No forced Azure specialist; use the application's owner. | - -## Adjacent-owner cases - -| Request | Primary route | Adjacent owner | Ordering rule | -|---|---|---|---| -| Concrete Terraform edit for Azure resources | `/internal-terraform` | `/internal-azure-organization-structure` or `/internal-azure-governance` | The Terraform router classifies the code edit; add Azure context only when the design decision is independently requested. | -| Concrete Azure Policy definition | `/internal-cloud-policy` | `/internal-azure-governance` | Cloud policy owns the policy artifact; governance supplies the control model only when separately requested. | -| Azure role selection for a named resource or action | `/awesome-copilot-azure-role-selector` | `/internal-azure-governance` | Role selection owns the role recommendation; governance adds the authorization model when separately requested. | -| Current Azure SKU price | `/awesome-copilot-azure-pricing` | `/internal-azure-strategic` | Pricing owns current cost data; add strategic framing only when options must be compared. | -| Azure resource health diagnosis | `/awesome-copilot-azure-resource-health-diagnose` | `/internal-azure-operations` | Resource health owns incident diagnosis; operations adds rollout or recovery proof when requested. | -| Azure DevOps CLI command or execution | `/awesome-copilot-azure-devops-cli` | `/internal-azure-devops` | CLI owns the command path; pipeline design is a separate deliverable. | -| Azure pricing estimate or commitment choice | `/awesome-copilot-azure-pricing` | `/internal-azure-strategic` | Pricing owns the estimate; strategic owns the recommendation when the choice is consequential. | - -## Multi-deliverable ordering - -| Deliverables | Primary | Secondary | -|---|---|---| -| Subscription placement followed by Policy design | `/internal-azure-organization-structure` | `/internal-azure-governance` | -| Governance design followed by rollout proof | `/internal-azure-governance` | `/internal-azure-operations` | -| Pipeline behavior followed by permission-boundary design | `/internal-azure-devops` | `/internal-azure-governance` | - -Start with the deliverable that establishes the next dependent decision and add -the secondary specialist only when its output is independently requested. diff --git a/.github/skills/internal-azure-organization-structure/references/topology-map.md b/.github/skills/internal-azure/references/structure.md similarity index 60% rename from .github/skills/internal-azure-organization-structure/references/topology-map.md rename to .github/skills/internal-azure/references/structure.md index b0196fb6..69fc4fae 100644 --- a/.github/skills/internal-azure-organization-structure/references/topology-map.md +++ b/.github/skills/internal-azure/references/structure.md @@ -1,12 +1,24 @@ -# Azure Organization Structure Topology Map +# Azure Structure -Use this reference for structural mappings, placement heuristics, and safe -rollout examples. +Use this reference for tenant, management-group, subscription, landing-zone, +residency, and platform-topology placement. + +## Structure patterns + +- Use management groups for enterprise segmentation and inheritance scope. +- Use subscriptions for workload, platform, environment, or residency + boundaries with explicit purpose and ownership. +- Use landing zones to package platform capabilities, connectivity, and + operating-model expectations. +- Keep hub-spoke, Virtual WAN, private connectivity, and regional placement + visible when they shape the platform topology. +- Separate platform subscriptions from workload subscriptions when shared + services need stable ownership. ## Structural mappings | Need | Placement surface | Structural rationale | -|---|---|---| +| --- | --- | --- | | Enterprise segmentation and inheritance scope | Tenant and management-group hierarchy | Groups establish stable policy and RBAC inheritance boundaries. | | Workload, platform, environment, or residency placement | Subscription model | Subscription purpose makes ownership, billing, and operational scope visible. | | Packaged platform capabilities and connectivity | Landing zone | Landing zones express shared services and operating-model expectations. | @@ -15,17 +27,15 @@ rollout examples. ## Placement heuristics | Question | Prefer | Rationale | -|---|---|---| +| --- | --- | --- | | Does the capability provide shared connectivity or central platform plumbing? | Platform landing zone or dedicated platform subscription | Shared ownership remains stable and visible. | | Does the capability exist for one workload or product boundary? | Workload landing zone or workload subscription | Application-specific ownership stays close to the workload. | | Does residency or regulated access change the operating model? | Dedicated hierarchy or landing-zone segment | Connectivity, sovereignty, and approval assumptions remain explicit. | | Does the change affect many subscriptions? | Management-group placement with staged rollout | Inheritance and blast radius are observable before expansion. | -## Safe rollout examples +## Rollout -| Structural change | Start with | Widen after | -|---|---|---| -| New management-group branch | One low-risk subscription family | Inheritance, policy scope, and operational ownership are confirmed. | -| Landing-zone baseline update | One landing zone or environment slice | Connectivity, automation, and rollback behavior are observed. | -| Platform subscription introduction | One shared capability with named consumers | Ownership, dependencies, and routing impact are validated. | -| Region or residency split | One workload set with explicit fallback | Connectivity, sovereignty, and continuity assumptions are proven. | +Stage structural changes with the staged-rollout table in +[`operations.md`](operations.md#staged-rollout). Validate inheritance, +connectivity, automation, ownership, and continuity assumptions before +widening. diff --git a/.github/skills/internal-azure/tests/evaluation/evals.json b/.github/skills/internal-azure/tests/evaluation/evals.json new file mode 100644 index 00000000..21afa60b --- /dev/null +++ b/.github/skills/internal-azure/tests/evaluation/evals.json @@ -0,0 +1,256 @@ +{ + "schema": "skill-eval-pack/v1", + "skill": "internal-azure", + "requirements": [ + {"id": "R-CLASSIFY", "text": "Classify the primary concern as structure, governance, or operations, load its reference, and load supporting references only when the same deliverable uses them.", "source": "Azure consolidation spec 2026-09-27, section 4 Workflow step 1"}, + {"id": "R-SCOPE", "text": "State the Azure scope level and the affected principals or workloads.", "source": "Azure consolidation spec, section 4 Workflow step 2; retired governance and organization-structure workflows"}, + {"id": "R-MECHANISM", "text": "Name the Azure mechanism, such as management group, subscription, landing zone, RBAC, managed identity, federation, PIM, or Policy, and why it fits.", "source": "Azure consolidation spec, section 4 Workflow step 3; retired lane pattern sections"}, + {"id": "R-IDENTITY-SPLIT", "text": "Keep workload identity separate from human access and prefer managed identity or federation over long-lived secrets.", "source": "former Azure governance lane patterns and guardrail-map workload identity table"}, + {"id": "R-EXCEPTION", "text": "Record exceptions with the typed fields for policy exemptions, temporary elevation, and managed-identity gaps.", "source": "retired guardrail-map Exception evidence table"}, + {"id": "R-ROLLOUT", "text": "For shared-blast-radius changes, name the first safe unit, the widening condition, and the rollback trigger.", "source": "retired topology-map Safe rollout examples and validation-and-evidence Stage-aware evidence"}, + {"id": "R-EVIDENCE", "text": "Separate confirmed from inferred evidence and distinguish backup, restore, and DR proof.", "source": "former Azure operations lane evidence patterns"}, + {"id": "R-DECISION-MODE", "text": "Produce a comparative decision note only when two or more options remain viable, and state reversibility.", "source": "former Azure strategic lane proportional output and lens-playbook depth control"}, + {"id": "R-LENSES", "text": "Activate FinOps or BC/DR lenses when cost or recovery is material, even with one viable option.", "source": "retired lens-playbook activation signals and BC/DR activation"}, + {"id": "R-HANDOFF", "text": "Route policy definitions, IaC code, pricing, role selection, resource health, Azure DevOps pipelines, GitHub Actions workflows, and application code to their owners.", "source": "retired routing-matrix Adjacent-owner cases; Azure consolidation spec section 6"}, + {"id": "R-FRESHNESS", "text": "Use current Microsoft documentation for freshness-sensitive facts, through Microsoft Learn MCP when available, or state the evidence gap.", "source": "Azure consolidation spec, section 4 Freshness"}, + {"id": "R-OUTPUT", "text": "Return the recommendation, material risk, and next validation action, and add other fields only when material.", "source": "Azure consolidation spec, section 4 Output; decision D3"}, + {"id": "R-BASELINE", "text": "Do not recommend standing Owner or User Access Administrator at subscription or wider scope, name the narrowest assignment scope, and prefer managed identity or federation over service principal secrets unless a typed exception is recorded.", "source": ".github/security-baseline.md Azure line; parity with internal-gcp Baseline"} + ], + "cases": [ + { + "id": "C-MG-LAYOUT", + "family": "bounded-revision", + "kind": "rubric", + "requirement_ids": ["R-CLASSIFY", "R-SCOPE", "R-MECHANISM", "R-ROLLOUT", "R-OUTPUT"], + "prompt": "We have three business units, a shared platform team, and a sandbox need. Propose the Azure management-group and subscription layout and how to introduce it without breaking current subscriptions.", + "initial_state": "Forty subscriptions sit directly under the tenant root group with no intermediate management groups.", + "expected_output": "A structure recommendation with platform, landing-zone, and sandbox branches, platform subscriptions separated from workload subscriptions, a first low-risk subscription family as the rollout unit, and a rollback trigger.", + "files": ["references/structure.md", "references/operations.md"], + "assertions": [ + {"id": "A-HIERARCHY", "text": "The response proposes management-group branches that separate platform, workload landing zones, and sandbox.", "critical": true}, + {"id": "A-FIRST-UNIT", "text": "The response moves one low-risk subscription family first and names the condition for widening.", "critical": true}, + {"id": "A-ROLLBACK", "text": "The response states a rollback trigger for the move.", "critical": false} + ], + "forbidden_actions": ["Move all subscriptions in one step.", "Write Terraform or Policy JSON as the main deliverable."], + "status": "not-run", + "held_out": false, + "rubric": { + "pass": ["Hierarchy, rollout unit, widening condition, and rollback trigger are all explicit.", "Platform and workload ownership stay separate."], + "fail": ["The layout has no staged move.", "The response mixes platform services into workload subscriptions without a reason."] + } + }, + { + "id": "C-WORKLOAD-IDENTITY", + "family": "bounded-revision", + "kind": "rubric", + "requirement_ids": ["R-IDENTITY-SPLIT", "R-MECHANISM", "R-SCOPE", "R-BASELINE"], + "prompt": "Our Function Apps use client secrets stored in app settings to reach Storage and Key Vault. Operators also share the same service principal. Fix the access model.", + "initial_state": "One service principal with a client secret is used both by workloads and by human operators.", + "expected_output": "Managed identity per workload with RBAC scoped to the target resources, separate human access through groups and PIM, and secret retirement.", + "files": ["references/governance.md"], + "assertions": [ + {"id": "A-MI", "text": "The response moves workloads to managed identity with resource-scoped RBAC.", "critical": true}, + {"id": "A-HUMAN-SPLIT", "text": "The response separates human operator access from workload identity.", "critical": true}, + {"id": "A-NO-STANDING-OWNER", "text": "The response does not grant standing Owner at subscription scope to workloads or operators.", "critical": true} + ], + "forbidden_actions": ["Keep a shared secret for both humans and workloads.", "Grant Owner at subscription scope to the workload."], + "status": "not-run", + "held_out": false, + "rubric": { + "pass": ["Managed identity, narrow scope, and a human path through PIM or groups are explicit."], + "fail": ["The response only rotates the secret.", "The response leaves humans and workloads on one principal."] + } + }, + { + "id": "C-POLICY-EXCEPTION", + "family": "competing-owner", + "kind": "rubric", + "requirement_ids": ["R-EXCEPTION", "R-HANDOFF"], + "prompt": "A legacy subscription cannot comply with our allowed-locations guardrail for six months. Define how we grant and govern the exception.", + "initial_state": "The allowed-locations assignment exists at the landing-zones management group.", + "expected_output": "A scoped policy exemption record with business reason, owner, scope, compensating controls, and expiry or review date; any concrete Policy JSON is routed to the policy owner.", + "files": ["references/governance.md", "references/adjacent-owners.md"], + "assertions": [ + {"id": "A-TYPED-FIELDS", "text": "The exception record contains reason, owner, scope, compensating control, and expiry or review date.", "critical": true}, + {"id": "A-POLICY-HANDOFF", "text": "Concrete Policy or exemption definition authoring is routed to internal-cloud-policy rather than written here as the main deliverable.", "critical": false} + ], + "forbidden_actions": ["Remove the policy assignment for the whole management group.", "Grant an exception without expiry."], + "status": "not-run", + "held_out": false, + "rubric": { + "pass": ["The exemption is scoped and time-bound with compensating controls."], + "fail": ["The exception is open-ended or unscoped."] + } + }, + { + "id": "C-RESTORE-PROOF", + "family": "bounded-revision", + "kind": "rubric", + "requirement_ids": ["R-EVIDENCE", "R-OUTPUT"], + "prompt": "Audit says our Azure Backup is fine because every job succeeded last month. Is that enough to claim we can recover the order database?", + "initial_state": "Backup job history is green; no restore has ever been tested.", + "expected_output": "Job success proves backup posture only; restore viability needs a restore exercise with observed recovery time and integrity check; DR needs a separate continuity exercise.", + "files": ["references/operations.md"], + "assertions": [ + {"id": "A-BACKUP-NOT-RESTORE", "text": "The response states that backup job success does not prove restore viability.", "critical": true}, + {"id": "A-RESTORE-PLAN", "text": "The response defines a restore exercise with recovery time and integrity evidence.", "critical": true} + ], + "forbidden_actions": ["Claim recoverability from job success alone."], + "status": "not-run", + "held_out": false, + "rubric": { + "pass": ["Backup, restore, and DR proof are kept distinct."], + "fail": ["The response treats green backup jobs as recovery proof."] + } + }, + { + "id": "C-HUB-VS-VWAN", + "family": "held-out-downstream-work", + "kind": "rubric", + "requirement_ids": ["R-DECISION-MODE", "R-FRESHNESS", "R-CLASSIFY"], + "prompt": "We are expanding to four regions with about 60 spokes and branch offices on SD-WAN. Hub-spoke with our own NVAs or Virtual WAN?", + "initial_state": "One region runs a customer-managed hub-spoke topology today.", + "expected_output": "A comparative decision note with both options, a recommendation under stated assumptions, reversibility, and a statement that current service limits and pricing must be checked in current Microsoft documentation.", + "files": ["references/decision-mode.md", "references/structure.md"], + "assertions": [ + {"id": "A-TWO-OPTIONS", "text": "The response compares both viable options and recommends one under stated assumptions.", "critical": true}, + {"id": "A-REVERSIBILITY", "text": "The response states how reversible the choice is.", "critical": true}, + {"id": "A-FRESH", "text": "The response flags limits or capability facts to verify in current Microsoft documentation, or states the evidence gap.", "critical": false} + ], + "forbidden_actions": ["Present exact current limits or prices as facts without a source."], + "status": "not-run", + "held_out": true, + "rubric": { + "pass": ["Options, recommendation, reversibility, and freshness caveat are present."], + "fail": ["Only one option is described.", "Service limits are asserted without source."] + } + }, + { + "id": "C-POLICY-JSON-HANDOFF", + "family": "competing-owner", + "kind": "rubric", + "requirement_ids": ["R-HANDOFF"], + "prompt": "Write the Azure Policy definition JSON that denies public IP addresses on network interfaces.", + "initial_state": "The request asks only for a concrete policy artifact.", + "expected_output": "The concrete policy definition is routed to internal-cloud-policy; Azure control-model guidance is offered only if separately requested.", + "files": ["references/adjacent-owners.md"], + "assertions": [ + {"id": "A-ROUTE-POLICY", "text": "The response routes the policy definition to internal-cloud-policy.", "critical": true} + ], + "forbidden_actions": ["Run a full governance workflow for a single policy artifact."], + "status": "not-run", + "held_out": false, + "rubric": { + "pass": ["The artifact owner is named and used."], + "fail": ["The Azure skill keeps ownership of the policy JSON."] + } + }, + { + "id": "C-SMALL-QUESTION", + "family": "bounded-revision", + "kind": "rubric", + "requirement_ids": ["R-OUTPUT"], + "prompt": "Should the shared Log Analytics workspace live in the management subscription or in each workload subscription?", + "initial_state": "A platform landing zone with a management subscription exists.", + "expected_output": "A short answer with the recommendation, the material risk, and the next validation action, without a long fixed field list.", + "files": [], + "assertions": [ + {"id": "A-THREE-FIELDS", "text": "The response contains a recommendation, a material risk, and a next validation action.", "critical": true}, + {"id": "A-PROPORTIONAL", "text": "The response does not add empty or irrelevant sections such as exception path or rollout unit.", "critical": false} + ], + "forbidden_actions": ["Emit a mandatory six-to-eight field template for a narrow question."], + "status": "not-run", + "held_out": false, + "rubric": { + "pass": ["The answer is proportional and contains the three required fields."], + "fail": ["The answer pads with irrelevant fields or omits the risk or next validation."] + } + }, + { + "id": "C-APP-CODE-NO-TRIGGER", + "family": "competing-owner", + "kind": "rubric", + "requirement_ids": ["R-HANDOFF"], + "prompt": "Refactor the Express routes of our Node.js API that runs on Azure App Service to use async handlers.", + "initial_state": "The task is application code; hosting on Azure is incidental.", + "expected_output": "No Azure platform workflow; the application owner handles the code change.", + "files": ["references/adjacent-owners.md"], + "assertions": [ + {"id": "A-NO-PLATFORM", "text": "The response does not apply the Azure platform workflow to application code.", "critical": true} + ], + "forbidden_actions": ["Add management-group, RBAC, or rollout analysis to an application refactor."], + "status": "not-run", + "held_out": true, + "rubric": { + "pass": ["The application owner is used; Azure hosting is treated as context only."], + "fail": ["The response produces Azure platform recommendations."] + } + }, + { + "id": "C-FINOPS-SINGLE-OPTION", + "family": "regression", + "kind": "rubric", + "requirement_ids": ["R-LENSES", "R-MECHANISM"], + "prompt": "We must enable Microsoft Defender for Cloud plans on all production subscriptions for compliance. Plan it.", + "initial_state": "Compliance leaves only one viable direction; cost impact is material across 25 subscriptions.", + "expected_output": "A single-direction plan that still surfaces the cost impact and how to estimate it, without inventing alternative options.", + "files": ["references/decision-mode.md"], + "assertions": [ + {"id": "A-COST-SURFACED", "text": "The response surfaces the material cost impact even though only one option is viable.", "critical": true}, + {"id": "A-NO-FAKE-OPTIONS", "text": "The response does not invent alternatives to fill a comparison template.", "critical": false} + ], + "forbidden_actions": ["State current plan prices as facts without a pricing source."], + "status": "not-run", + "held_out": false, + "rubric": { + "pass": ["FinOps lens is active and pricing is routed to a current source."], + "fail": ["Cost is ignored because there is no option comparison."] + } + }, + { + "id": "C-BCDR-REGION", + "family": "held-out-downstream-work", + "kind": "rubric", + "requirement_ids": ["R-LENSES", "R-EVIDENCE"], + "prompt": "We are moving our payment platform to a single Azure region to simplify operations. What should we check?", + "initial_state": "The payment platform is business-critical with a stated RTO of one hour.", + "expected_output": "The BC/DR lens is active: regional failure risk against the RTO, and the recovery proof needed before the move.", + "files": ["references/decision-mode.md", "references/operations.md"], + "assertions": [ + {"id": "A-BCDR-ACTIVE", "text": "The response raises regional continuity risk against the stated RTO.", "critical": true}, + {"id": "A-DR-PROOF", "text": "The response names the recovery exercise evidence required before or after the move.", "critical": false} + ], + "forbidden_actions": ["Approve the single-region move without discussing recovery posture."], + "status": "not-run", + "held_out": true, + "rubric": { + "pass": ["Continuity risk and recovery proof are explicit."], + "fail": ["The response only discusses operational simplicity."] + } + } + ], + "triggers": { + "queries": [ + {"id": "Q-01", "query": "Design a management-group hierarchy for three business units and a sandbox in our Azure tenant.", "should_trigger": true, "split": "train"}, + {"id": "Q-02", "query": "Should shared connectivity live in an Azure platform subscription or in each workload subscription?", "should_trigger": true, "split": "train"}, + {"id": "Q-03", "query": "Define an Azure RBAC operating model with PIM for production subscriptions.", "should_trigger": true, "split": "train"}, + {"id": "Q-04", "query": "Replace client secrets with managed identity for our Azure Function Apps.", "should_trigger": true, "split": "train"}, + {"id": "Q-05", "query": "How do we roll out an allowed-locations Azure Policy guardrail safely across 40 subscriptions?", "should_trigger": true, "split": "train"}, + {"id": "Q-06", "query": "What evidence proves our Azure Backup restores actually work?", "should_trigger": true, "split": "train"}, + {"id": "Q-07", "query": "Hub-spoke or Virtual WAN for four Azure regions?", "should_trigger": true, "split": "held-out"}, + {"id": "Q-08", "query": "Plan an Azure landing-zone residency split for EU-regulated workloads.", "should_trigger": true, "split": "held-out"}, + {"id": "Q-09", "query": "Which Log Analytics signals confirm an Azure Policy rollout did not break deployments?", "should_trigger": true, "split": "held-out"}, + {"id": "Q-10", "query": "Is Azure Site Recovery worth the cost for this tier-2 application?", "should_trigger": true, "split": "held-out"}, + {"id": "Q-11", "query": "Write the Azure Policy JSON that denies public IPs.", "should_trigger": false, "split": "train", "competing_owner": "internal-cloud-policy"}, + {"id": "Q-12", "query": "Fix this azurerm_role_assignment Terraform block.", "should_trigger": false, "split": "train", "competing_owner": "internal-terraform"}, + {"id": "Q-13", "query": "What does a D4s v5 VM cost per month in West Europe?", "should_trigger": false, "split": "train", "competing_owner": "awesome-copilot-azure-pricing"}, + {"id": "Q-14", "query": "Which built-in role lets a pipeline read one storage account?", "should_trigger": false, "split": "train", "competing_owner": "awesome-copilot-azure-role-selector"}, + {"id": "Q-15", "query": "Why is my App Service returning 503 right now?", "should_trigger": false, "split": "held-out", "competing_owner": "awesome-copilot-azure-resource-health-diagnose"}, + {"id": "Q-16", "query": "Add a stage with an approval to our azure-pipelines.yml.", "should_trigger": false, "split": "train", "competing_owner": "internal-azure-devops"}, + {"id": "Q-17", "query": "Write a GitHub Actions workflow that deploys to Azure with OIDC.", "should_trigger": false, "split": "held-out", "competing_owner": "internal-github-actions"}, + {"id": "Q-18", "query": "Refactor a Node.js API hosted on App Service.", "should_trigger": false, "split": "held-out", "competing_owner": "internal-nodejs-project"}, + {"id": "Q-19", "query": "Design an AWS OU layout for regulated workloads.", "should_trigger": false, "split": "train", "competing_owner": "internal-aws"}, + {"id": "Q-20", "query": "Choose a GCP folder layout for three business units.", "should_trigger": false, "split": "held-out", "competing_owner": "internal-gcp"} + ] + } +} diff --git a/.github/skills/internal-bash-script/SKILL.md b/.github/skills/internal-bash-script/SKILL.md index 65267d73..e20ed464 100644 --- a/.github/skills/internal-bash-script/SKILL.md +++ b/.github/skills/internal-bash-script/SKILL.md @@ -1,96 +1,115 @@ --- name: internal-bash-script -description: Use when creating, reviewing, or modifying standalone Bash or POSIX `sh` scripts, utilities, wrappers, launchers, or other operator-facing shell entrypoints. +description: Use when creating, reviewing, or modifying standalone Bash or POSIX `sh` scripts, utilities, wrappers, launchers, or other operator-facing shell entrypoints. Route embedded shell fragments and sourced helpers inside another program to /internal-bash. --- # Internal Bash Script ## When to use -- New Bash scripts. -- Existing Bash scripts that need review or updates. -- Standalone operator-facing wrappers, launchers, and shell utilities. +- Creating, reviewing, or modifying a standalone Bash or POSIX `sh` script, + utility, wrapper, launcher, or other operator entrypoint. +- Route sourced helpers and shell fragments to `/internal-bash`. Route shell + embedded in another format to that format's owner first. -## When not to use +## Dialect minimum -- Embedded shell fragments and sourced helpers inside another program; route - them to `/internal-bash`. -- Sourced or non-operator Bash helpers that do not own an operator-facing - entrypoint. -- Workflow or platform behavior beyond the standalone shell entrypoint. +Record the dialect contract before choosing a template: the declared +interpreter, the execution environment, and the compatibility target. Preserve +the declared interpreter and never change it silently. -## Dialect decision +- `Dialect: Bash`: follow the deployment shebang convention, normally + `#!/usr/bin/env bash`. Use `set -euo pipefail` with documented exceptions. + Arrays, `[[ ]]`, and `local` are valid only here. +- `Dialect: POSIX sh`: follow the target's shebang convention. Use `set -eu`, + scalar variables, positional parameters, and `[ ]` or `test`. Do not use + arrays or `local`. Use `pipefail` only with an explicit POSIX.1-2024 + baseline. -Classify the shell before choosing script patterns. Preserve the declared -interpreter and record the execution environment and compatibility target: +Bash invoked as `sh` is not cross-shell portability proof. -- `Dialect: Bash` when the entrypoint and runtime provide Bash. -- `Dialect: POSIX`sh`` when the entrypoint or deployment target requires POSIX - shell syntax. +Add `Compatibility target: Bash 3.2` only when the caller or repository +declares macOS `/bin/bash` support. Then avoid `mapfile`, `readarray`, +`declare -A`, `${var,,}`, `${var^^}`, and `wait -n`, and guard an empty-array +expansion under `set -u` with a `${#array[@]}` check. Static checks do not +detect these. Use the Python test harness to execute the shell target under +`/bin/bash` 3.2 in an isolated workspace, or report +`Bash 3.2 compatibility: unverified`. -Require an explicit POSIX baseline before treating POSIX.1-2024 Issue 8 -behavior as portable. Bash invoked as `sh` is not cross-shell portability proof, -and the interpreter must not be changed silently. +## Portable minimum -## Portable core - -Quote expansions, check statuses at correctness boundaries, use `if`, `case`, -`test`, or `[ ]` for shared control flow, validate dependencies before first -use, and use `mktemp` with cleanup traps for temporary state. - -## Bash branch - -For `Dialect: Bash`, follow the deployment shebang convention, normally -`#!/usr/bin/env bash`, and use `set -euo pipefail` with documented exceptions. -Arrays for dynamic commands, `[[ ]]`, `local`, and Bash-specific traps or -options are valid only in this branch. - -## POSIX `sh` branch - -For `Dialect: POSIX`sh``, follow the shebang convention declared by the target, -use `set -eu` with contextual `-e` caveats, and use scalar variables, positional -parameters, `[ ]`, and `test`. Do not use Bash arrays or `local`; use `pipefail` -only for an explicit POSIX.1-2024 baseline. - -## Script-specific operator guidance - -- Use `command -v` before first use of required external tools. +- Quote expansions and check statuses at correctness boundaries. +- Validate required external commands with `command -v` before first use. +- Use `mktemp` with cleanup traps for temporary state. - Use structured parsers such as `jq` or `yq` for JSON and YAML when available. -- Prefer `printf` for formatted output and arrays for dynamic commands only in - the Bash branch. -- Destructive or repeatable scripts should be idempotent and expose `--dry-run` when operator risk is non-trivial. -- Keep operator entrypoints thin and extract repeated branches into sourced helper files only when reuse is real. -- Treat 300 lines as a review threshold and 400 lines as a split-or-justify gate for standalone scripts. -- When script output can grow and the script is agent-facing, prefer bounded summaries by default and add an explicit compact or quiet mode that still preserves blockers, failures, and required next actions. -- Keep full-detail output reachable through an explicit flag or durable artifact path when operators need full diagnostics. + +For the full anti-pattern catalog, owner-routing table, and the static +checker, load `/internal-bash` when it is available. + +## Operator guidance + +- Prefer `printf` for formatted output. +- Destructive or repeatable scripts should be idempotent and expose + `--dry-run` when operator risk is non-trivial. +- Prefer safe reruns with guards like `mkdir -p`, existence checks, or + replace-in-place flows. +- Use `--` before user-supplied paths in destructive commands such as + `rm -rf -- "$target"`. +- Keep operator entrypoints thin and extract repeated branches into sourced + helper files only when reuse is real. +- Treat 300 lines as a review threshold and 400 lines as a split-or-justify + gate for standalone scripts. +- Preserve an existing `--format compact` option, payload, and consumers. + Consider a terminal-only `--compact` projection only for a new interface + with measured high output volume. ## Testing - For behavior changes, create the failing focused check before the first implementation edit. -- Prefer the repository's existing Bash harness. Cover parser decisions, - guards, dry-run behavior, command construction, and rerun safety at their - stable boundary. +- Write behavioral test setup, expected results, and assertions in Python. + Reuse the repository's existing Python test framework. Cover parser + decisions, guards, dry-run behavior, command construction, and rerun safety + at their stable boundary. +- Execute the real script through `subprocess`. Shell fixtures, command + stubs, and invocation snippets may exercise the target, but must not + implement test assertions. Do not introduce Bats or shell-based assertion + harnesses. +- When the script is documented for direct invocation, that invocation is the + stable boundary. Reaching the code through an interpreter tests a different + boundary and leaves the executable bit, the shebang, and `PATH` resolution + unverified: `bash ./tool.sh` passes where `./tool.sh` fails. - When no harness can exercise the behavior before editing, record a pre-code testability exception and the alternate validation path. Use syntax, lint, and a safe non-mutating invocation as evidence; do not represent later regression coverage as test-first work. - -## Templates and hardening helpers - -After dialect selection, load `references/templates.md` when you need the -starter script, argument parser skeleton, or cleanup helpers. Choose only the -matching Bash or POSIX `sh` section. - -- Prefer safe reruns with guards like `mkdir -p`, existence checks, or replace-in-place flows. -- Use `--` before user-supplied paths in destructive commands such as `rm -rf -- "$target"`. - -## Common mistakes - -Load `references/common-mistakes.md` for the full mistake table. +- Name the defect each test detects and use independent expectations. Include + meaningful success, error, and boundary cases; verify outputs, exit codes, + and intended or forbidden filesystem effects. +- Give process tests an isolated workspace, controlled environment, and bounded + timeout. Keep the script real; stub only external command boundaries and + verify their arguments when command construction is part of the contract. +- Run focused native tests first, then relevant integration and declared shell + compatibility checks. Measure comparable timings before claiming gains. +- Load [Testing recipes](references/testing.md) when authoring or reorganizing + tests. Keep test assertions in Python and preserve the existing framework. + +## References + +- [references/templates.md](references/templates.md): load after the dialect + decision for the starter script, argument parser, ERR trap, or cleanup + helpers. Use only the section that matches the dialect. +- [references/operator-output.md](references/operator-output.md): load for + multi-step lifecycle output, bounded diagnostics, compact, quiet, and verbose + behavior, and failure continuation. +- [references/common-mistakes.md](references/common-mistakes.md): load before + finishing any script creation, modification, or review. ## Validation -- `bash -n script.sh` (syntax check) -- `shellcheck -s bash script.sh` (lint) -- `shfmt -d script.sh` (format diff, if available) -- `sh -n script.sh` (POSIX `sh` syntax check) -- `shellcheck -s sh script.sh` (POSIX `sh` lint) +- `Dialect: Bash`: run `bash -n FILE` and `shellcheck -s bash FILE`. +- `Dialect: POSIX sh`: run `dash -n FILE` (or `sh -n FILE`, reported as + limited evidence) and `shellcheck -s sh FILE`. +- When `/internal-bash` is loaded, its static checker runs these checks with a + stable exit contract. +- Run a safe, non-mutating direct invocation, such as `./tool.sh --help`, to + verify the executable bit, the shebang, and the argument parser. +- `shfmt -d FILE` is an optional format check. diff --git a/.github/skills/internal-bash-script/references/common-mistakes.md b/.github/skills/internal-bash-script/references/common-mistakes.md index cb5ca276..f3f12451 100644 --- a/.github/skills/internal-bash-script/references/common-mistakes.md +++ b/.github/skills/internal-bash-script/references/common-mistakes.md @@ -1,13 +1,14 @@ # Common Mistakes For Bash and POSIX `sh` Scripts +The dialect and portable minimum live in `SKILL.md`; the optional +`internal-bash` catalog covers them in depth. This table covers operator +entrypoints. + | Mistake | Why it matters | Instead | | --- | --- | --- | -| Leaving the dialect undeclared | Review rules and syntax choices become contradictory | Record the interpreter, execution environment, and POSIX baseline before choosing patterns | -| Silently changing `sh` to Bash | Deployment behavior and portability can change without approval | Preserve the declared interpreter or explicitly update the contract | -| Using Bash extensions under POSIX `sh` | Arrays, `[[ ]]`, and `local` are not portable POSIX `sh` syntax | Use scalar variables, `[ ]`, `test`, and POSIX control flow | -| Treating Bash invoked as `sh` as portability proof | One implementation does not represent each supported `sh` | Run syntax and behavior checks under every repository-supported `sh` implementation | -| Assuming `pipefail` on an unspecified `/bin/sh` | Many `/bin/sh` implementations do not provide it | Require an explicit POSIX.1-2024 baseline or avoid the option | -| Skipping dependency checks for required commands | Failures surface late and with weaker operator context | Check `command -v` before the first call | +| Missing purpose or usage context | Operators cannot discover the entrypoint contract locally | Add a purpose header, usage examples, and `--help` | | Building dynamic Bash commands as strings | Quoting and argument boundaries become fragile | Use arrays plus `printf` in the Bash branch; use carefully quoted scalar invocations in POSIX `sh` | | Destructive commands without rerun safety | Repeated execution can corrupt state or surprise operators | Add `--dry-run` and make the mutation idempotent | +| A multi-function script with no clear entrypoint | The starting point is not obvious | Define `main` and invoke it at the execution boundary. An unconditional call still runs when sourced; test standalone scripts by invocation and source only helpers designed for it | +| Executable statements placed between function definitions | Execution order becomes hard to follow and side effects run before the entrypoint | Keep top-level statements together below the function definitions | | Rewriting parser or cleanup scaffolding from scratch | Operator UX and failure handling drift between scripts | Reuse the starter and helper patterns from `references/templates.md` | diff --git a/.github/skills/internal-bash-script/references/operator-output.md b/.github/skills/internal-bash-script/references/operator-output.md new file mode 100644 index 00000000..97706490 --- /dev/null +++ b/.github/skills/internal-bash-script/references/operator-output.md @@ -0,0 +1,50 @@ +# Bash Operator Output Guidance + +Load this reference when a Bash or POSIX shell entrypoint reports multi-step work, diagnostics, or results to an operator. + +## Ownership and output flow + +Keep argument parsing, selector resolution, execution and scheduling, lifecycle events, result collection, human rendering, machine-readable output, and explicitly requested artifacts as distinct responsibilities. The shell entrypoint owns its output contract. The caller may forward output or choose a supported format; it must not reinterpret the script's selectors, statuses, schemas, or result meaning. + +Execution must not depend on terminal visibility, color support, or the selected human-output projection. Keep machine-readable results plain and stable. Create files only when a caller explicitly requests them. Use the script's existing shell dialect and project-specific options; this reference adds no library or universal flag. + +## Lifecycle and human logs + +For each selected unit in a multi-step command, write `▶️ started` when it begins and one terminal state when it ends: `✅ passed`, `❌ failed`, `⚠️ warning`, or `⏭️ skipped`. Include the unit category and selector. Add `[n/N]` when position helps the operator follow progress. + +For a long-running unit, write `⏱️ heartbeat` only when it reports measured progress. Bound heartbeat frequency and output using project-specific choices. Keep a concise started event and terminal-state event for every selected unit, including failures and skips. + +Human log markers pair the accepted emoji with its state word. If the active output encoding cannot represent Unicode, print the state word alone. Do not use ASCII art. `--no-color` and `NO_COLOR` disable ANSI styling only; they do not remove emoji or state words. + +A portable event can use `printf`: + +```sh +printf '%s\n' '▶️ started · [1/3] terraform/identity' +printf '%s\n' '❌ failed · [1/3] terraform/identity · exit=1' +``` + +## Diagnostics and summaries + +Choose diagnostic-count and raw-output-size limits for the project and apply them per unit. Bound diagnostic excerpts and captured raw output; show an omitted count or byte count when a limit is reached. Preserve every unit's status and every failure and skip in human and structured results even when diagnostic detail is bounded. + +Omit unavailable fields rather than inventing values. Do not infer a cause that the command did not establish. Keep secrets out of logs and redact sensitive values before rendering. Provide an explicit project-owned option or result path for full detail; write a diagnostic artifact only when requested. + +Finish with counts by state, failures grouped by category or selector, and a next action only when evidence supports a concrete, verifiable step. For example: + +```text +summary: passed=2 failed=1 warning=0 skipped=1 +failures: terraform/identity=2 +next: rerun the failed selector with the owning tool's documented command +``` + +## Compact, quiet, and verbose output + +A new, high-volume terminal interface may consider `--compact` as a terminal-only projection when measured output volume justifies it. Compact output preserves each unit's lifecycle state and does not change selectors, work units, scheduling, concurrency, exit codes, machine-readable output, CI formats, Markdown, or artifacts. + +Keep `--quiet` separate: it may suppress routine narration, but it must retain lifecycle states, failure and skip records, and the final summary. Use `--verbose` or another explicit project-owned option to request additional diagnostic detail. These output choices do not select different work or alter machine-readable results. + +Preserve an existing `--format compact` option, payload, and consumers as a public interface. Do not rename it to `--compact` or change its meaning. When a real CI consumer exists and the script supports a CI format, choose that format explicitly before auto-detection. A terminal-only compact projection is not a CI-format substitution. + +## Continuing after failures + +Independent validation units should continue after a failure unless the operator selected an explicit, safe `--fail-fast` policy. Choose fail-fast behavior for mutating or unsafe work from its actual side effects and recovery properties. Do not add a universal fail-fast default or flag. Preserve all failures and skips collected before execution stops. diff --git a/.github/skills/internal-bash-script/references/templates.md b/.github/skills/internal-bash-script/references/templates.md index 1f97b2fa..b3d0569f 100644 --- a/.github/skills/internal-bash-script/references/templates.md +++ b/.github/skills/internal-bash-script/references/templates.md @@ -4,6 +4,16 @@ Load this reference after dialect selection. The shebangs below are deployment conventions; they do not change the language guarantees of the selected dialect. +## Contents + +- [Bash Minimal Template](#bash-minimal-template) +- [POSIX sh Minimal Template](#posix-sh-minimal-template) +- [Bash Argument Parsing Pattern](#bash-argument-parsing-pattern) +- [POSIX sh Argument Parsing Pattern](#posix-sh-argument-parsing-pattern) +- [Bash ERR Trap Pattern](#bash-err-trap-pattern) +- [Bash Hardening Helpers](#bash-hardening-helpers) +- [POSIX sh Cleanup Helper](#posix-sh-cleanup-helper) + ## Bash Minimal Template ```bash diff --git a/.github/skills/internal-bash-script/references/testing.md b/.github/skills/internal-bash-script/references/testing.md new file mode 100644 index 00000000..6068f79a --- /dev/null +++ b/.github/skills/internal-bash-script/references/testing.md @@ -0,0 +1,125 @@ +# Python Tests For Operator Shell Scripts + +Keep the existing Python framework and native component test location. This +stdlib unittest recipe demonstrates direct invocation of real shell code, +isolated files, and a controlled external command. Use framework-native +fixtures when the repository already provides them. + +## Contents + +- [Cases and levels](#cases-and-levels) +- [Runnable direct-invocation recipe](#runnable-direct-invocation-recipe) +- [Feedback and evidence](#feedback-and-evidence) + +## Cases and levels + +| Contract | Useful checks | +| --- | --- | +| Parser and guards | Help, missing value, invalid option, refusal without effects | +| Command construction | Argument boundaries and external failure propagation | +| Filesystem | Intended content or removal, protected paths unchanged | +| Operator safety | Dry-run has no mutation; repeated invocation is safe | +| Cleanup | Temporary state is removed after success and relevant failures | + +Choose cases from the actual contract. Execute the documented invocation to +verify executable permissions and shebang resolution. An interpreter-only run +adds dialect evidence but does not replace the direct boundary. Source only +helpers explicitly designed for sourcing: wrapping code in `main` with an +unconditional final call still executes it on source. + +## Runnable direct-invocation recipe + +Place the test beside executable `tool.sh`, the bundled illustrative script at +[`cache-clean-valid.sh`](../tests/evaluation/fixtures/cache-clean-valid.sh). +The example copy must have executable permission. For a real target, test its +existing permission rather than silently repairing it in test setup. + +```python +import os +import subprocess +import tempfile +import unittest +from pathlib import Path + +TOOL = Path(__file__).with_name("tool.sh") + + +class ScriptContractTests(unittest.TestCase): + def setUp(self): + self.workspace = tempfile.TemporaryDirectory() + self.addCleanup(self.workspace.cleanup) + self.root = Path(self.workspace.name) + self.cache = self.root / "cache space" + self.target = self.cache / "dev [x]" + self.target.mkdir(parents=True) + self.marker = self.target / "keep.txt" + self.marker.write_text("keep", encoding="utf-8") + self.env = {"PATH": os.defpath, "HOME": str(self.root), + "LC_ALL": "C", "CACHE_ROOT": str(self.cache)} + + def run_tool(self, *args, env=None): + return subprocess.run( + [str(TOOL), *args], cwd=self.root, + env=self.env if env is None else env, + capture_output=True, text=True, check=False, timeout=5, + ) + + def test_dry_run_preserves_cache(self): + result = self.run_tool("--env", "dev [x]", "--dry-run") + self.assertEqual(result.returncode, 0, result.stderr) + self.assertEqual(result.stdout, f"would remove {self.target}\n") + self.assertEqual(self.marker.read_text(encoding="utf-8"), "keep") + + def test_cleanup_is_repeatable_and_keeps_neighbor(self): + neighbor = self.cache / "protected" + neighbor.mkdir() + for _ in range(2): + result = self.run_tool("--env", "dev [x]") + self.assertEqual(result.returncode, 0, result.stderr) + self.assertFalse(self.target.exists()) + self.assertTrue(neighbor.is_dir()) + + def test_missing_value_refuses_without_deletion(self): + result = self.run_tool("--env") + self.assertNotEqual(result.returncode, 0) + self.assertIn("--env", result.stderr) + self.assertEqual(self.marker.read_text(encoding="utf-8"), "keep") + + def test_external_failure_preserves_status_and_arguments(self): + bin_dir = self.root / "bin" + bin_dir.mkdir() + trace = self.root / "args.bin" + stub = bin_dir / "rm" + stub.write_text( + "#!/bin/sh\n" + + r'''printf '%s\0' "$@" > "$TRACE_FILE"''' + + "\nexit 23\n", + encoding="utf-8", + ) + stub.chmod(0o755) + env = dict(self.env, PATH=f"{bin_dir}{os.pathsep}{os.defpath}", + TRACE_FILE=str(trace)) + result = self.run_tool("--env", "dev [x]", env=env) + self.assertEqual(result.returncode, 23, result.stderr) + self.assertEqual(trace.read_bytes().split(b"\0")[:-1], + [b"-rf", b"--", os.fsencode(self.target)]) + self.assertEqual(self.marker.read_text(encoding="utf-8"), "keep") +``` + +The stub only records command input and supplies a failure. Python verifies the +real script's error propagation and command construction. The separate cleanup +case executes actual removal inside the disposable workspace. Assertions about +command arguments are justified here because safe command construction is the +contract; they do not establish production permissions or live service behavior. + +## Feedback and evidence + +Run the focused case first, then the component suite, relevant actual integration, +and declared shell compatibility. Preserve syntax and ShellCheck as separate +static evidence. Report missing interpreters and dependencies explicitly. + +Keep mutable state per case; capture status, stdout, stderr, and effects. Control +home, config, locale, cwd, and command discovery. Bound execution and own cleanup +of spawned descendants where needed. Diagnose intermittent failures rather than +adding blanket retries. Measure comparable commands, environments, selected +cases, and setup costs before claiming speed gains or adding parallelism. diff --git a/.github/skills/internal-bash-script/tests/evaluation/evals.json b/.github/skills/internal-bash-script/tests/evaluation/evals.json new file mode 100644 index 00000000..3c4aa267 --- /dev/null +++ b/.github/skills/internal-bash-script/tests/evaluation/evals.json @@ -0,0 +1,376 @@ +{ + "schema": "skill-eval-pack/v1", + "skill": "internal-bash-script", + "requirements": [ + { + "id": "R-DIALECT-MINIMUM", + "text": "Record the dialect contract and apply this bundle's dialect and portable minimum before choosing a template; /internal-bash is an optional enrichment, and the task completes without it.", + "source": "internal-bash-script SKILL.md: Dialect minimum, Portable minimum" + }, + { + "id": "R-ROUTE", + "text": "Route sourced helpers and shell fragments to internal-bash and embedded shell to the enclosing format's owner first.", + "source": "internal-bash-script SKILL.md: When to use" + }, + { + "id": "R-OPERATOR-SAFETY", + "text": "An operator entrypoint states its purpose and usage, exposes --help, validates arguments, uses -- before user-supplied paths in destructive commands, and offers --dry-run when operator risk is non-trivial.", + "source": "internal-bash-script SKILL.md: Operator guidance; references/common-mistakes.md" + }, + { + "id": "R-TEST-BOUNDARY", + "text": "For behavior changes, create the failing focused check first and exercise a directly invoked script through direct invocation, not through an interpreter.", + "source": "internal-bash-script SKILL.md: Testing" + }, + { + "id": "R-TEMPLATE-DIALECT", + "text": "Use only the template section that matches the recorded dialect.", + "source": "internal-bash-script SKILL.md: References; references/templates.md" + }, + { + "id": "R-VALIDATION", + "text": "Validate every changed script with the dialect's syntax check and shellcheck, or the internal-bash checker when loaded, plus a safe, non-mutating direct invocation.", + "source": "internal-bash-script SKILL.md: Validation" + }, + { + "id": "R-COMPACT-INTERFACE", + "text": "Preserve an existing --format compact option, payload, and consumers; do not rename it or change its meaning.", + "source": "internal-bash-script SKILL.md: Operator guidance; references/operator-output.md" + }, + { + "id": "R-BASH32-TARGET", + "text": "Apply Bash 3.2 rules only when the target is declared, and report compatibility as unverified unless a Python test harness executed the shell target under Bash 3.2 in an isolated workspace.", + "source": "internal-bash-script SKILL.md: Dialect minimum" + }, + { + "id": "R-PYTHON-TESTS", + "text": "Write behavioral test setup, expected results, and assertions in Python using the existing Python framework. Shell fixtures, stubs, and invocation snippets may exercise the real target, but must not implement test assertions. Preserve direct invocation and do not introduce Bats or a shell assertion harness.", + "source": "internal-bash-script SKILL.md: Testing" + }, + { + "id": "R-TEST-DESIGN", + "text": "Choose observable behavior and meaningful success, error, and boundary cases; derive expected results independently; isolate mutable state and preserve the declared framework.", + "source": "internal-bash-script SKILL.md: Testing; references/testing.md" + }, + { + "id": "R-TEST-SPEED", + "text": "Use focused native tests first, separate actual integration and declared compatibility checks, and justify speed changes with comparable measurements.", + "source": "internal-bash-script SKILL.md: Testing; references/testing.md" + }, + { + "id": "R-TEST-PROOF", + "text": "Provide executable test recipes that detect representative defects, keep bundle guidance self-contained, and distinguish structural evidence from runtime effectiveness.", + "source": "internal-bash-script SKILL.md: Testing; references/testing.md" + } + ], + "cases": [ + { + "id": "C-OPERATOR-GUARDS", + "family": "operator-safety", + "kind": "deterministic", + "requirement_ids": ["R-OPERATOR-SAFETY", "R-VALIDATION"], + "prompt": "Review this cache cleanup script that operators run by hand.", + "initial_state": "A destructive Bash entrypoint. Reproduce with a disposable directory: mkdir -p tmp/eval-cache/dev, then run with CACHE_ROOT=tmp/eval-cache.", + "expected_output": "The valid script supports --help and --dry-run, rejects missing values, and uses rm -rf --; the defective script is flagged for missing usage, missing argument validation, no dry-run, and an unguarded rm -rf.", + "files": ["tests/evaluation/fixtures/cache-clean-valid.sh"], + "assertions": [ + {"id": "A-HELP", "text": "`./cache-clean-valid.sh --help` exits 0 and prints usage without side effects.", "critical": true}, + {"id": "A-DRY-RUN", "text": "`CACHE_ROOT=tmp/eval-cache ./cache-clean-valid.sh --env dev --dry-run` exits 0 and tmp/eval-cache/dev still exists.", "critical": true}, + {"id": "A-MISSING-VALUE", "text": "`./cache-clean-valid.sh --env` exits 1 with an error on stderr.", "critical": true}, + {"id": "A-STATIC-CHECK", "text": "`bash -n` and `shellcheck -s bash` pass on the valid fixture and report findings on the defective fixture.", "critical": true}, + {"id": "A-DEFECT-FLAGGED", "text": "The review of the defective fixture names every missing guard.", "critical": true} + ], + "forbidden_actions": ["Run the defective fixture without a disposable CACHE_ROOT."], + "status": "generated", + "held_out": false, + "defective_fixture": "tests/evaluation/fixtures/cache-clean-defective.sh" + }, + { + "id": "C-DIALECT-MINIMUM", + "family": "dialect-minimum", + "kind": "rubric", + "requirement_ids": ["R-DIALECT-MINIMUM", "R-TEMPLATE-DIALECT"], + "prompt": "Create scripts/rotate-logs.sh for our Alpine containers; it must run under /bin/sh.", + "initial_state": "No existing script; the deployment target provides BusyBox sh only.", + "expected_output": "Record Dialect: POSIX sh and start from the POSIX sh template without arrays, local, or [[ ]].", + "files": [], + "assertions": [ + {"id": "A-DIALECT-RECORDED", "text": "The dialect contract is stated before a template is chosen.", "critical": true}, + {"id": "A-POSIX-TEMPLATE", "text": "The output uses only the POSIX sh template section.", "critical": true} + ], + "forbidden_actions": ["Use the Bash template for a /bin/sh target."], + "status": "not-run", + "held_out": false, + "rubric": { + "pass": ["The transcript states Dialect: POSIX sh, and the script has a #!/bin/sh shebang with set -eu."], + "fail": ["The script uses #!/usr/bin/env bash, arrays, local, or [[ ]]."] + } + }, + { + "id": "C-BASELINE-UNAVAILABLE", + "family": "missing-dependency", + "kind": "rubric", + "requirement_ids": ["R-DIALECT-MINIMUM", "R-VALIDATION"], + "prompt": "Add a --verbose flag to tests/evaluation/fixtures/report.sh.", + "initial_state": "The host exposes internal-bash-script but not internal-bash; the script is tests/evaluation/fixtures/report.sh.", + "expected_output": "The change is completed with this bundle alone: dialect recorded as Bash, bash -n and shellcheck -s bash run, and ./report.sh --help invoked directly.", + "files": ["tests/evaluation/fixtures/report.sh"], + "assertions": [ + {"id": "A-NO-BLOCK", "text": "The missing internal-bash skill does not block or stall the task.", "critical": true}, + {"id": "A-FALLBACK-CHECKS", "text": "bash -n and shellcheck -s bash are run on the changed script.", "critical": true} + ], + "forbidden_actions": ["Invoke a script path inside another skill bundle."], + "status": "not-run", + "held_out": true, + "rubric": { + "pass": ["The transcript shows the change, both static checks, and the direct --help run without loading internal-bash."], + "fail": ["The agent stops because internal-bash is missing, or skips static checks."] + } + }, + { + "id": "C-COMPACT-PRESERVE", + "family": "regression", + "kind": "rubric", + "requirement_ids": ["R-COMPACT-INTERFACE"], + "prompt": "Tidy up the output of tests/evaluation/fixtures/status-compact.sh and add a --compact flag for terminals.", + "initial_state": "The script already exposes --format compact, and a CI job parses its ok= and fail= fields.", + "expected_output": "--format compact keeps its name and exact payload; any new terminal projection is additive and leaves the CI format unchanged.", + "files": ["tests/evaluation/fixtures/status-compact.sh"], + "assertions": [ + {"id": "A-PAYLOAD-KEPT", "text": "`--format compact` still prints `ok=3 fail=0`.", "critical": true}, + {"id": "A-NO-RENAME", "text": "--format compact is not renamed or aliased away to --compact.", "critical": true} + ], + "forbidden_actions": ["Change or remove the existing --format compact payload."], + "status": "not-run", + "held_out": false, + "rubric": { + "pass": ["A before/after run of --format compact shows identical output."], + "fail": ["The compact payload changes, or --format compact is replaced by --compact."] + } + }, + { + "id": "C-DIRECT-INVOCATION", + "family": "test-boundary", + "kind": "rubric", + "requirement_ids": ["R-TEST-BOUNDARY", "R-VALIDATION", "R-PYTHON-TESTS"], + "prompt": "Add a --json flag to tests/evaluation/fixtures/report.sh, which operators run as ./report.sh; extend the harness beside it.", + "initial_state": "An executable Bash script and a Python harness (report-harness.py) that invokes it directly.", + "expected_output": "A failing harness check that invokes ./report.sh --dir . --json directly, then the implementation, then the passing check plus bash -n and shellcheck.", + "files": ["tests/evaluation/fixtures/report.sh", "tests/evaluation/fixtures/report-harness.py"], + "assertions": [ + {"id": "A-RED-FIRST", "text": "The failing test runs before the implementation edit.", "critical": true}, + {"id": "A-DIRECT", "text": "The test runs the script directly, not as `bash ./report.sh`.", "critical": true}, + {"id": "A-PYTHON-ORACLE", "text": "The extended harness retains Python setup and assertions on the real process results; no Bats or shell assertion harness is introduced.", "critical": true} + ], + "forbidden_actions": ["Claim test-first work for a test added after the implementation."], + "status": "not-run", + "held_out": false, + "rubric": { + "pass": ["The red run precedes the diff and the subprocess call targets the script path."], + "fail": ["The test wraps the script in bash, or the red run is missing."] + } + }, + { + "id": "C-ROUTE-SOURCED", + "family": "competing-owner", + "kind": "rubric", + "requirement_ids": ["R-ROUTE"], + "prompt": "Clean up lib/retry.sh; it only defines functions that our scripts source.", + "initial_state": "A sourced function library with no shebang and no entrypoint.", + "expected_output": "Route the work to internal-bash; operator guidance such as --help and --dry-run is not added to a sourced library.", + "files": [], + "assertions": [ + {"id": "A-ROUTED", "text": "internal-bash owns the change.", "critical": true} + ], + "forbidden_actions": ["Add argument parsing or a main entrypoint to a sourced library."], + "status": "not-run", + "held_out": true, + "rubric": { + "pass": ["The transcript routes to internal-bash and keeps the file a function library."], + "fail": ["The agent adds usage text, --help, or main to the sourced file."] + } + }, + { + "id": "C-BASH32-EMPTY-ARRAY", + "family": "evidence-honesty", + "kind": "deterministic", + "requirement_ids": ["R-BASH32-TARGET", "R-VALIDATION"], + "prompt": "Operators also run this script with the macOS system Bash. Is it compatible?", + "initial_state": "macOS host where /bin/bash --version reports 3.2; on any other host record the case as blocked. The fixtures only print text.", + "expected_output": "bash -n and shellcheck -s bash pass on both fixtures, so the answer does not treat them as compatibility proof. Under /bin/bash 3.2 the defective fixture exits 1 with an unbound-variable error and the valid fixture exits 0.", + "files": ["tests/evaluation/fixtures/bash32-empty-array-valid.bash"], + "assertions": [ + {"id": "A-STATIC-LIMIT", "text": "The answer states that bash -n and shellcheck do not establish Bash 3.2 compatibility.", "critical": true}, + {"id": "A-RUNTIME-SPLIT", "text": "/bin/bash 3.2 exits 1 on the defective fixture and 0 on the valid fixture.", "critical": true} + ], + "forbidden_actions": ["Claim Bash 3.2 compatibility from static checks alone.", "Invoke a script path inside another skill bundle."], + "status": "generated", + "held_out": false, + "defective_fixture": "tests/evaluation/fixtures/bash32-empty-array-defective.bash" + }, + { + "id": "C-DIALECT-PARITY", + "family": "dialect-parity", + "kind": "rubric", + "requirement_ids": ["R-DIALECT-MINIMUM"], + "prompt": "Review the dialect handling in tests/evaluation/fixtures/posix-parity.sh; operators run it under /bin/sh on Alpine.", + "initial_state": "The same fixture ships in internal-bash, whose C-DIALECT-PARITY case uses the same assertion IDs.", + "expected_output": "Record Dialect: POSIX sh, keep the #!/bin/sh interpreter, and propose no Bash-only syntax; the guarded glob is not reported.", + "files": ["tests/evaluation/fixtures/posix-parity.sh"], + "assertions": [ + {"id": "A-PARITY-DIALECT", "text": "The answer records Dialect: POSIX sh.", "critical": true}, + {"id": "A-PARITY-INTERPRETER", "text": "The #!/bin/sh interpreter is kept.", "critical": true}, + {"id": "A-PARITY-NO-BASH", "text": "No change introduces [[ ]], arrays, local, or pipefail.", "critical": true} + ], + "forbidden_actions": ["Change #!/bin/sh to a Bash shebang without approval."], + "status": "not-run", + "held_out": false, + "rubric": { + "pass": ["The same three assertions hold for this skill and for internal-bash on the shared fixture."], + "fail": ["The two skills reach different dialect decisions on the same fixture."] + } + }, + { + "id": "C-PYTHON-SHELL-STUB", + "family": "test-harness", + "kind": "rubric", + "requirement_ids": ["R-PYTHON-TESTS", "R-TEST-BOUNDARY"], + "prompt": "Add Python tests for tests/evaluation/fixtures/report.sh. Reuse the stdlib Python harness beside it. Use a temporary shell stub for find to control populated and empty directory listings, and cover a missing --dir value. Change only tests.", + "initial_state": "The executable report.sh requires a --dir value and counts find output lines. report-harness.py already invokes it directly. The real filesystem and an isolated PATH are available; no additional test dependency is needed.", + "expected_output": "Python creates and cleans up a temporary directory and executable find stub, invokes report.sh directly with an isolated PATH, and asserts files=2 for two stub output lines, files=0 for no lines, and a non-zero exit for a missing --dir value. The shell stub supplies controlled output without comparing expected results.", + "files": ["tests/evaluation/fixtures/report.sh", "tests/evaluation/fixtures/report-harness.py"], + "assertions": [ + {"id": "A-PYTHON-LIFECYCLE", "text": "Python owns the temporary files, cleanup, environment setup, expected results, and assertions.", "critical": true}, + {"id": "A-SHELL-STUB-ALLOWED", "text": "An executable shell stub supplies controlled listing output, but it does not implement the test oracle.", "critical": true}, + {"id": "A-DIRECT-REAL-TARGET", "text": "subprocess invokes the real report.sh directly and verifies two files, zero files, and rejection of a missing --dir value.", "critical": true}, + {"id": "A-NO-NEW-FRAMEWORK", "text": "The existing stdlib Python harness is reused without adding pytest, Bats, or dependencies.", "critical": true} + ], + "forbidden_actions": ["Modify report.sh.", "Replace the real target with a mock.", "Introduce Bats or shell expected-result assertions."], + "status": "not-run", + "held_out": true, + "rubric": { + "pass": ["The generated Python tests run the populated and empty listing paths and the invalid argument case directly, then independently assert the observed outputs and statuses."], + "fail": ["Shell fixture files are forbidden merely for being shell, expected-result checks live in the stub, or the target is invoked as bash report.sh."] + } + }, + { + "id": "C-TEST-AUTHORING", + "family": "testing-quality", + "kind": "rubric", + "requirement_ids": [ + "R-TEST-DESIGN", + "R-TEST-SPEED", + "R-TEST-PROOF" + ], + "prompt": "Add useful fast Python tests for a shell target that receives a directory containing spaces and wildcard characters, invokes an external command, and must report command failures without corrupting files.", + "initial_state": "The repository already uses unittest with a stdlib harness. The shell dialect and deployment invocation are declared. The external command must not reach a live provider.", + "expected_output": "Keep unittest, execute the real declared shell boundary, pass inputs as arguments, use isolated workspace and environment with a bounded timeout, and verify quoted paths, command failures, and filesystem effects.", + "files": [ + "SKILL.md" + ], + "assertions": [ + { + "id": "A-REAL-BEHAVIOR", + "text": "Tests execute the actual changed contract and detect at least one representative wrong result, wrong argument, or unintended side effect.", + "critical": true + }, + { + "id": "A-ISOLATION", + "text": "Mutable state is isolated and subprocess inputs, environment, working directory, and timeout are explicit when processes are needed.", + "critical": true + }, + { + "id": "A-SPEED-EVIDENCE", + "text": "Focused and broader evidence are distinguished; any claimed speed gain names comparable baseline and final measurements.", + "critical": true + } + ], + "forbidden_actions": [ + "Assert instructional wording instead of behavior.", + "Mock the function being changed.", + "Migrate a working framework without authorization.", + "Claim runtime effectiveness from schema validation." + ], + "status": "not-run", + "held_out": false, + "rubric": { + "pass": [ + "A reviewer can identify the defect each test detects.", + "Relevant observable outputs and file effects are checked with isolated state.", + "Dependencies and supported runtime gaps are reported; optional speed machinery is evidence-based." + ], + "fail": [ + "Tests assert only mock presence or implementation source.", + "Tests inherit uncontrolled mutable home or shared output files.", + "Performance or model effectiveness is claimed without observed evidence." + ] + } + }, + { + "id": "C-TEST-HELD-OUT", + "family": "testing-quality", + "kind": "rubric", + "requirement_ids": [ + "R-TEST-DESIGN", + "R-TEST-SPEED" + ], + "prompt": "A refactor preserves the public contract, but tests fail because private helper names changed. Another test passes although the output artifact is no longer written. Repair the tests without introducing a new framework or broad fixture layer.", + "initial_state": "Existing native tests and an established public output contract; compatibility requirements are explicitly declared.", + "expected_output": "Keep real contract coverage, replace private structure assertions with independent result and side-effect checks, preserve framework and layout, and run focused then relevant broader checks.", + "files": [ + "SKILL.md" + ], + "assertions": [ + { + "id": "A-STABLE-CONTRACT", + "text": "The missing artifact causes a test failure while a behavior-preserving private refactor does not.", + "critical": true + }, + { + "id": "A-BOUNDED-COST", + "text": "Shared fixtures or parallel execution are introduced only with isolation and measured cost justification.", + "critical": true + } + ], + "forbidden_actions": [ + "Weaken the artifact contract to make the suite pass." + ], + "status": "not-run", + "held_out": true, + "rubric": { + "pass": [ + "Contract-preserving internal changes keep assertions valid.", + "A representative missing output defect is detected." + ], + "fail": [ + "Assertions pin private symbol names.", + "The expected value is computed by the implementation being tested." + ] + } + } + ], + "triggers": { + "queries": [ + {"id": "Q-01", "query": "Write a bootstrap.sh that developers run once to install the toolchain.", "should_trigger": true, "split": "train"}, + {"id": "Q-02", "query": "Add --dry-run to scripts/cleanup.sh before we run it in production.", "should_trigger": true, "split": "train"}, + {"id": "Q-03", "query": "Review our deploy wrapper script for operator safety and error messages.", "should_trigger": true, "split": "train"}, + {"id": "Q-04", "query": "Create a launcher script that picks the right JVM and execs the app.", "should_trigger": true, "split": "train"}, + {"id": "Q-05", "query": "The backup.sh utility needs --help and proper argument validation.", "should_trigger": true, "split": "train"}, + {"id": "Q-06", "query": "Make run-checks.sh print started and passed lines per step and a final summary.", "should_trigger": true, "split": "train"}, + {"id": "Q-07", "query": "Port our installer script from Bash to POSIX sh so it runs on BusyBox.", "should_trigger": true, "split": "held-out"}, + {"id": "Q-08", "query": "Write a test for tools/release.sh that calls it the way operators do.", "should_trigger": true, "split": "held-out"}, + {"id": "Q-09", "query": "Split this 500-line maintenance script into an entrypoint plus sourced helpers.", "should_trigger": true, "split": "held-out"}, + {"id": "Q-10", "query": "Review the quoting in lib/common.sh, which our tools source.", "should_trigger": false, "split": "train", "competing_owner": "internal-bash"}, + {"id": "Q-11", "query": "Fix the shell inside this Make recipe that breaks on paths with spaces.", "should_trigger": false, "split": "train", "competing_owner": "internal-makefile"}, + {"id": "Q-12", "query": "Add a matrix job to the CI workflow.", "should_trigger": false, "split": "train", "competing_owner": "internal-github-actions"}, + {"id": "Q-13", "query": "Build a Python CLI with click to manage releases.", "should_trigger": false, "split": "train", "competing_owner": "internal-python-script"}, + {"id": "Q-14", "query": "Optimize the Dockerfile layer order for caching.", "should_trigger": false, "split": "train", "competing_owner": "internal-docker"}, + {"id": "Q-15", "query": "Is `local` portable in a POSIX sh function library?", "should_trigger": false, "split": "train", "competing_owner": "internal-bash"}, + {"id": "Q-16", "query": "Generate an interactive wizard that walks me through creating CI secrets by hand.", "should_trigger": false, "split": "held-out", "competing_owner": "mattpocock-wizard"}, + {"id": "Q-17", "query": "Configure my shell aliases in .zshrc.", "should_trigger": false, "split": "held-out"}, + {"id": "Q-18", "query": "Add targets and .PHONY to the Makefile for lint and test.", "should_trigger": false, "split": "held-out", "competing_owner": "internal-makefile"}, + {"id": "Q-19", "query": "Why does my ./tool.sh fail with permission denied while bash tool.sh works?", "should_trigger": true, "split": "held-out"}, + {"id": "Q-20", "query": "Harden the sourced trap helper used by all our scripts.", "should_trigger": false, "split": "held-out", "competing_owner": "internal-bash"} + ] + } +} diff --git a/.github/skills/internal-bash-script/tests/evaluation/fixtures/bash32-empty-array-defective.bash b/.github/skills/internal-bash-script/tests/evaluation/fixtures/bash32-empty-array-defective.bash new file mode 100644 index 00000000..73680c06 --- /dev/null +++ b/.github/skills/internal-bash-script/tests/evaluation/fixtures/bash32-empty-array-defective.bash @@ -0,0 +1,9 @@ +#!/usr/bin/env bash +set -euo pipefail + +print_args() { + local -a extra=() + printf 'arg: %s\n' "${extra[@]}" +} + +print_args diff --git a/.github/skills/internal-bash-script/tests/evaluation/fixtures/bash32-empty-array-valid.bash b/.github/skills/internal-bash-script/tests/evaluation/fixtures/bash32-empty-array-valid.bash new file mode 100644 index 00000000..3b4c2801 --- /dev/null +++ b/.github/skills/internal-bash-script/tests/evaluation/fixtures/bash32-empty-array-valid.bash @@ -0,0 +1,11 @@ +#!/usr/bin/env bash +set -euo pipefail + +print_args() { + local -a extra=() + if [[ ${#extra[@]} -gt 0 ]]; then + printf 'arg: %s\n' "${extra[@]}" + fi +} + +print_args diff --git a/.github/skills/internal-bash-script/tests/evaluation/fixtures/cache-clean-defective.sh b/.github/skills/internal-bash-script/tests/evaluation/fixtures/cache-clean-defective.sh new file mode 100755 index 00000000..0c8e92f1 --- /dev/null +++ b/.github/skills/internal-bash-script/tests/evaluation/fixtures/cache-clean-defective.sh @@ -0,0 +1,5 @@ +#!/usr/bin/env bash +set -euo pipefail + +env_name="$1" +rm -rf "$CACHE_ROOT/$env_name" diff --git a/.github/skills/internal-bash-script/tests/evaluation/fixtures/cache-clean-valid.sh b/.github/skills/internal-bash-script/tests/evaluation/fixtures/cache-clean-valid.sh new file mode 100755 index 00000000..0f0fe10d --- /dev/null +++ b/.github/skills/internal-bash-script/tests/evaluation/fixtures/cache-clean-valid.sh @@ -0,0 +1,41 @@ +#!/usr/bin/env bash +# +# Purpose: Remove a cache directory for one environment. +# Usage examples: +# ./cache-clean.sh --env dev --dry-run +# ./cache-clean.sh --help + +set -euo pipefail + +usage() { + printf '%s\n' 'Usage: ./cache-clean.sh --env NAME [--dry-run]' +} + +main() { + local env_name="" dry_run=false + while [[ $# -gt 0 ]]; do + case "$1" in + --env) + [[ $# -ge 2 && "$2" != -* ]] || { + printf '%s\n' '--env requires a value' >&2 + exit 1 + } + env_name="$2" + shift 2 + ;; + --dry-run) dry_run=true; shift ;; + --help) usage; exit 0 ;; + *) usage >&2; exit 1 ;; + esac + done + [[ -n "$env_name" ]] || { usage >&2; exit 1; } + + local target="${CACHE_ROOT:?CACHE_ROOT is required}/${env_name}" + if [[ "$dry_run" == true ]]; then + printf 'would remove %s\n' "$target" + return 0 + fi + rm -rf -- "$target" +} + +main "$@" diff --git a/.github/skills/internal-bash-script/tests/evaluation/fixtures/posix-parity.sh b/.github/skills/internal-bash-script/tests/evaluation/fixtures/posix-parity.sh new file mode 100644 index 00000000..3c5aac15 --- /dev/null +++ b/.github/skills/internal-bash-script/tests/evaluation/fixtures/posix-parity.sh @@ -0,0 +1,11 @@ +#!/bin/sh +set -eu + +rotate_logs() { + log_dir=${1:?Missing log directory} + cd "$log_dir" || return 1 + for log in *.log; do + [ -e "$log" ] || continue + mv -- "$log" "$log.1" + done +} diff --git a/.github/skills/internal-bash-script/tests/evaluation/fixtures/report-harness.py b/.github/skills/internal-bash-script/tests/evaluation/fixtures/report-harness.py new file mode 100644 index 00000000..f7fde0f2 --- /dev/null +++ b/.github/skills/internal-bash-script/tests/evaluation/fixtures/report-harness.py @@ -0,0 +1,17 @@ +# Harness fixture for C-DIRECT-INVOCATION; not collected by pytest. +import subprocess +from pathlib import Path + +SCRIPT = Path(__file__).with_name("report.sh") + + +def run(*args: str) -> subprocess.CompletedProcess[str]: + return subprocess.run( + [str(SCRIPT), *args], text=True, capture_output=True, check=False + ) + + +def check_help() -> None: + result = run("--help") + assert result.returncode == 0 + assert "Usage" in result.stdout diff --git a/.github/skills/internal-bash-script/tests/evaluation/fixtures/report.sh b/.github/skills/internal-bash-script/tests/evaluation/fixtures/report.sh new file mode 100755 index 00000000..17b3fc6f --- /dev/null +++ b/.github/skills/internal-bash-script/tests/evaluation/fixtures/report.sh @@ -0,0 +1,31 @@ +#!/usr/bin/env bash +# +# Purpose: Print a line count report for a directory. +# Usage examples: +# ./report.sh --dir src +# ./report.sh --help + +set -euo pipefail + +usage() { + printf '%s\n' 'Usage: ./report.sh --dir PATH' +} + +main() { + local dir="" + while [[ $# -gt 0 ]]; do + case "$1" in + --dir) + [[ $# -ge 2 && "$2" != -* ]] || { usage >&2; exit 1; } + dir="$2" + shift 2 + ;; + --help) usage; exit 0 ;; + *) usage >&2; exit 1 ;; + esac + done + [[ -n "$dir" ]] || { usage >&2; exit 1; } + printf 'files=%s\n' "$(find "$dir" -type f | wc -l | tr -d ' ')" +} + +main "$@" diff --git a/.github/skills/internal-bash-script/tests/evaluation/fixtures/status-compact.sh b/.github/skills/internal-bash-script/tests/evaluation/fixtures/status-compact.sh new file mode 100755 index 00000000..6e7f310d --- /dev/null +++ b/.github/skills/internal-bash-script/tests/evaluation/fixtures/status-compact.sh @@ -0,0 +1,16 @@ +#!/usr/bin/env bash +set -euo pipefail + +format="text" +while [[ $# -gt 0 ]]; do + case "$1" in + --format) format="$2"; shift 2 ;; + *) exit 1 ;; + esac +done + +if [[ "$format" == compact ]]; then + printf 'ok=%s fail=%s\n' 3 0 +else + printf 'passed: 3\nfailed: 0\n' +fi diff --git a/.github/skills/internal-bash-script/tests/test_testing_examples.py b/.github/skills/internal-bash-script/tests/test_testing_examples.py new file mode 100644 index 00000000..c8fc3bf0 --- /dev/null +++ b/.github/skills/internal-bash-script/tests/test_testing_examples.py @@ -0,0 +1,33 @@ +"""Execute published recipes against working and known defective shell targets.""" + +import os +import re +import subprocess +import sys +from pathlib import Path + +import pytest + +BUNDLE = Path(__file__).resolve().parents[1] + + +@pytest.mark.parametrize("variant", ["valid", "defective"]) +def test_recipe_detects_real_shell_defects(tmp_path: Path, variant: str) -> None: + reference = (BUNDLE / "references/testing.md").read_text(encoding="utf-8") + (recipe,) = re.findall(r"```python\n(.*?)```", reference, re.DOTALL) + (tmp_path / "test_recipe.py").write_text(recipe, encoding="utf-8") + fixture = BUNDLE / f"tests/evaluation/fixtures/cache-clean-{variant}.sh" + target = tmp_path / "tool.sh" + target.write_bytes(fixture.read_bytes()) + target.chmod(0o755) + result = subprocess.run( + [sys.executable, "-m", "unittest", "test_recipe", "-v"], + cwd=tmp_path, + env={"PATH": os.defpath, "HOME": str(tmp_path), "LC_ALL": "C"}, + capture_output=True, text=True, check=False, timeout=20, + ) + if variant == "valid": + assert result.returncode == 0, result.stderr + else: + assert result.returncode == 1, result.stderr + assert "FAIL:" in result.stderr, result.stderr diff --git a/.github/skills/internal-bash/SKILL.md b/.github/skills/internal-bash/SKILL.md index d80d5b24..636d423d 100644 --- a/.github/skills/internal-bash/SKILL.md +++ b/.github/skills/internal-bash/SKILL.md @@ -1,32 +1,27 @@ --- name: internal-bash -description: Use when creating, analyzing, reviewing, or modifying embedded Bash or POSIX `sh`, sourced shell helpers, or non-operator shell fragments that need dialect, safety, quoting, parser, or validation guidance. +description: Use when creating, analyzing, reviewing, or modifying embedded Bash or POSIX `sh`, sourced shell helpers, or non-operator shell fragments that need dialect, safety, quoting, parser, or validation guidance. Route standalone scripts, utilities, wrappers, and launchers to /internal-bash-script. --- # Internal Bash -## Referenced files - -- `references/review-anti-patterns.md`: Bash review anti-pattern catalog with - ID-tagged patterns, severity, rationale, and examples. Load for focused review - of shell content within this skill's scope. - ## When to use -- Review or modification of sourced `.sh` helpers and Bash snippets where the - main need is a shared safety baseline. -- Shell embedded in repository automation when no narrower owner has stronger rules. -- Non-operator Bash helpers that do not own a standalone operator entrypoint. -- Quick checks for quoting, strict mode, guard clauses, temp files, and parser choices. +- Sourced `.sh` helpers, shell snippets, and non-operator shell fragments. +- Shell semantics inside another format, such as a CI `run:` step, a Make + recipe, or a Dockerfile `RUN`, after that format's owner has claimed the file. +- The dialect, safety, review, and validation baseline that + `/internal-bash-script` may load for the full catalog and the checker. + +## Owner routing -## When not to use +Classify the file before applying rules: -- Standalone script design, standalone script review, launcher behavior, - operator UX, or script templates. Route standalone scripts, utilities, - wrappers, and launchers to `/internal-bash-script`. -- Shell embedded in an automation format whose enclosing platform contract is - the primary subject. -- Workflow-level behavior beyond the shell fragment itself. +| Evidence | Owner | +| --- | --- | +| Operator entrypoint: shebang with direct invocation, arguments, or usage text | Route to `/internal-bash-script` | +| Sourced file, function library, or shell fragment | This skill | +| Shell embedded in another format | Load that format's owner first; use this skill only for the shell semantics | ## Dialect decision @@ -41,6 +36,14 @@ execution environment, and the compatibility target as the dialect contract: Require an explicit POSIX baseline before treating Issue 8 behavior as portable. Do not infer portability from Bash invoked as `sh`. +Add `Compatibility target: Bash 3.2` only when the caller or repository +declares macOS `/bin/bash` support. Then avoid `mapfile`, `readarray`, +`declare -A`, `${var,,}`, `${var^^}`, and `wait -n`, and guard an empty-array +expansion under `set -u` with a `${#array[@]}` check. Static checks do not +detect these. Use the Python test harness to execute the shell target under +`/bin/bash` 3.2 in an isolated workspace, or report +`Bash 3.2 compatibility: unverified`. + ## Portable core - Quote expansions and use explicit status checks at correctness boundaries. @@ -52,9 +55,10 @@ portable. Do not infer portability from Bash invoked as `sh`. ## Bash branch -For `Dialect: Bash`, use the repository shebang convention -`#!/usr/bin/env bash`, `set -euo pipefail` with documented exceptions, arrays -for dynamic commands, `[[ ]]`, `local`, and Bash-specific traps or options. +For `Dialect: Bash`, follow the deployment shebang convention, normally +`#!/usr/bin/env bash`. Use `set -euo pipefail` with documented exceptions, +arrays for dynamic commands, `[[ ]]`, `local`, and Bash-specific traps or +options. ## POSIX `sh` branch @@ -72,11 +76,56 @@ baseline; it is not a safe assumption for an unspecified `/bin/sh`. - Apply pragmatic DRY: de-duplicate repeated decision paths, but keep one-off logic local when extraction harms auditability. +## Review + +For a focused review, load +[references/review-anti-patterns.md](references/review-anti-patterns.md) and +report each finding with its ID, severity, location, and the declared dialect. +A finding that assumes the wrong dialect is not a finding. + +## Testing + +- Write behavioral test setup, expected results, and assertions in Python. + Reuse the repository's existing Python test framework. +- Execute the real shell target through `subprocess` with the declared + interpreter. Shell fixtures, command stubs, and invocation snippets may + exercise the target, but must not implement test assertions. +- For sourced helpers, use a fixed shell snippet to load the helper and call + its functions. Pass test inputs as subprocess arguments, not interpolated + shell source, and assert captured output and exit status in Python. +- Do not introduce Bats or shell-based assertion harnesses. Syntax checks and + ShellCheck complement behavioral tests; they do not replace them. +- Name the defect each test detects and use independent expectations. Cover + meaningful success, error, and boundary cases, including argument boundaries, + failure propagation, and observable file effects where applicable. +- Give process tests an isolated workspace, controlled environment, and bounded + timeout. Keep the real shell target; stub only external command boundaries. +- Run focused native tests first, then relevant integration and declared shell + compatibility checks. Measure comparable timings before claiming gains. +- Load [Testing recipes](references/testing.md) when authoring or reorganizing + tests; the examples keep assertions in Python and run actual shell code. + ## Validation -For `Dialect: Bash`, run `bash -n