Skip to content

feat: V28082026 improve Refactor skills and enhance validation processes - #91

Open
diegoitaliait wants to merge 150 commits into
mainfrom
v28082026-improve
Open

diegoitaliait wants to merge 150 commits into
mainfrom
v28082026-improve

Conversation

@diegoitaliait

Copy link
Copy Markdown
Contributor

Description

Refactor and remove deprecated skills while introducing new skills and improving validation processes. This update enhances the test suite and validation scripts to ensure consistency and correctness across the skill framework.

Change Type

  • Agent update
  • Instruction update
  • Skill update
  • AI policy update (AGENTS.md, copilot-instructions.md)
  • Workflow/automation update
  • Verification/automation update
  • Documentation/configuration only
  • Scripts/Utils update
  • Other

List of changes

  • Removed deprecated skills and configurations.
  • Added new skills and improved their validation.
  • Enhanced test coverage for skills and validation scripts.

Consumer Impact

  • Affected customization assets: Skills and validation scripts.
  • Backward compatibility notes: Deprecated skills removed.
  • Rollout or migration notes (if any): Ensure to update any references to removed skills.

Testing Instructions

  1. Run the updated test suite to validate all skills and configurations.
  2. Verify the functionality of newly added skills.

Validation Evidence

  • Repository-level checks result: All checks passed.
  • Stack-specific checks/tests result: Tests executed successfully.
  • Security/supply-chain checks result: No vulnerabilities found.

Breaking Changes

  • This PR contains breaking changes
  • Deprecated skills have been removed, requiring updates to any dependent configurations.

Checklist

  • No secrets/tokens/credentials committed
  • Repository artifacts remain in English
  • Available repository checks executed as applicable
  • Workflow action pins use full SHA with adjacent release/tag reference
  • References remain repository-agnostic and reusable

- Deleted the `mattpocock-writing-great-skills` skill and its associated files, including SKILL.md and openai.yaml.
- Updated tests to remove references to the deleted skill and adjusted assertions accordingly.
- Added new skills `grill-me` and `grilling` to the resource management.
- Enhanced the normalization tests to ensure proper handling of legacy paths and invocation policies.
- Improved the test suite to validate the presence and correctness of new skills and their configurations.
- Ensured that the internal skill creator uses current entry points and does not reference removed skills.
…tions

- Introduced tests for local sync repos agent metadata and bundle paths.
- Enhanced code analysis workflow to run pytest on skill tests.
- Updated Makefile to include skill test paths in linting and testing.
- Documented the validation surface for contract checks in tech documentation.
- Configured pytest to discover tests in both root and skill directories.
- Implemented a scoring mechanism for codebase-improvement gateway evaluations.
- Added fixtures for passing and failing runs in gateway evaluation tests.
- Created comprehensive tests for gateway evaluation scoring logic.
- Ensured Makefile does not reference skill runtime scripts directly.
- Validated the layout of skill tests to ensure they are co-located with their bundles.
grill-me was blocked in both runtimes and the sync manifest regenerated
the flags on every refresh, so any skill mandating /grill-me could not
load it and silently improvised instead of interviewing the user.

- Remove disable-model-invocation from the grill-me bundle, drop
  policy.allow_implicit_invocation from its Codex metadata, and delete
  the invocation_policy entry that regenerated both.
- Update the sync contract split from 14/11 to 13/12 user-invoked and
  model-invoked Matt skills.
- Drop /superpowers-brainstorming from the internal-gateway-idea
  forbidden pre-acceptance routes; a user-invoked-only skill is not an
  agent route before or after acceptance.
- Route material internal-agent-creator questions through /grill-me
  instead of an ad-hoc prompt.
- Added `validate-skill-change-scope.py` to enforce strict allowlisting for protected external skill changes.
- Updated `validate_internal_skills.py` to align with new structure and import paths.
- Refactored various scripts and documentation to reflect changes in script naming conventions (e.g., `validate_internal_skills.py` to `validate-internal-skills.py`).
- Modified Makefile and workflow configurations to accommodate new script names and paths.
- Enhanced documentation to clarify usage and structure of scripts related to skill validation and catalog consistency.
- Removed obsolete test for subprocess timeout handling.
- Updated tests to ensure compatibility with new script paths and naming conventions.
- Removed direct calls to `validate-internal-skills.py` in various SKILL.md files, replacing them with a unified command using `run.sh` for consistency.
- Added new reference files for decision ledger and candidate persistence to streamline decision-making processes.
- Updated manifest references to clarify execution and authoring responsibilities.
- Improved report formatting guidelines and shapes for findings, residuals, and open questions.
- Enhanced command portability documentation to ensure clarity on executable commands and validation processes.
- Adjusted test scripts to align with new directory structure and validation methods.
- Introduced `test-rules.py` to validate skill prose assertions and internal skill findings.
- Created `test-scope.py` to ensure proper classification of protected skill bundles and validate allowlist functionality.
- Added `test-runner.py` to verify script resolution for various tools and ensure proper handling of protected skill scope validators.
- Implemented `test-rule-coverage.py` to check coverage of token rules and ensure expected findings are produced.
- Updated `test_makefile_skill_script_references.py` to reflect changes in script runner references.
- Enhanced `test_repository_test_layout_contract.py` to validate the structure of GitHub automation layout and ensure test paths are correctly identified.
@diegoitaliait
diegoitaliait requested review from a team and GNuccio96 as code owners August 29, 2026 17:50
…sts for missing references in OpenAI metadata
…hance file matching patterns and documentation generation
- Implemented tests for the knowledge configuration contract in `test_config.py`, ensuring valid configurations and error handling for unknown rules.
- Created tests for inventory discovery in `test_inventory.py`, validating component detection and reporting capabilities.
- Added behavior tests for operating modes in `test_modes.py`, covering audit and update functionalities.
- Introduced neutral tests for deterministic checks in `test_check.py`, verifying the behavior of the knowledge check process.
…on, inventory, and modes

- Deleted test_check.py, test_config.py, test_inventory.py, and test_modes.py
- These files contained neutral tests for various functionalities related to knowledge checks, configuration validation, inventory discovery, and operating modes.
- The removal is part of a cleanup effort to streamline the testing suite and remove outdated or redundant tests.
…ication; add guidance on invocation boundaries and deferring lessons
…age guidance, new references, and maintainability checks
…ance reference and update related descriptions
…rkflow, diagram disposition, and ownership evidence
…rity on evidence handling and domain promotion
…source synchronization; enhance clarity on asset management and synchronization processes
…p and refresh modes with setup and sync, introduce bucket mechanism, and enhance evaluation scenarios
…aid documentation; clarify styling and compatibility considerations
…c; clarify plan authoring readiness and separate implementation permissions
…lan authoring readiness and permissions for implementation
…nning; enhance inventory.py to exclude manifest support files from imported skills in tests
…cated code

- Removed unused imports and constants related to protocol markers and expected events.
- Simplified the extraction of metadata and evaluation cases from YAML and JSON files.
- Consolidated test cases to focus on new authoring contracts and metadata validation.
- Ensured retired gateway protocol files and routes are absent from the repository.
- Updated tests to reflect changes in the protocol structure and expectations.
…ferences to manage output paths and instruction overrides, ensuring compliance with legacy entries and improving artifact handling in tmp directory.
- Updated knowledge-navigation.md to clarify routing to maintained owners and ownership reporting.
- Expanded knowledge-report.md with a new section for reporting problems found, including a structured table format.
- Revised knowledge-scope.md to specify requirements for authored deletions and enforcement gap reporting.
- Improved knowledge-topology.md by refining definitions of common core and lanes, and added claim classification guidance.
- Adjusted managerial-maintenance.md to clarify roadmap creation criteria and ownership linking.
- Modified project-memory-maintenance.md to emphasize the separation of execution tasks from orientation entries.
- Enhanced readme-maintenance.md to specify conditions for including a table of contents and operational detail relocation.
- Updated standards-maintenance.md to clarify reporting gaps and avoid creating empty entries.
- Refined evaluation tests in bindings.json, compatibility.json, evals.json, and others to align with new documentation standards and ensure proper coverage.
- Added new tests for problem reporting, deletion approval, and knowledge layer preservation in test_eval_graders.py and test_eval_pack_integrity.py.
- Improved runtime view tests to ensure correct copying of public projections and selected references.
…r generated plans and enhance documentation for clarity
…irements for clarity and consistency in execution processes
…zation to enforce single implementation plans and retire deprecated skills
…ance and personal study repositories

- Created F7.tree.json and F8.tree.json to represent a non-code community grant governance repository and a personal study repository, respectively.
- Added corresponding gold files F7.gold.json and F8.gold.json to define expected outcomes for the new fixtures.
- Enhanced test coverage in test_bundle_contract.py to ensure knowledge types and alignment references are comprehensive.
- Updated test_eval_graders.py to remove obsolete lifecycle checks and ensure proper bindings for evaluation scenarios.
- Improved test_eval_pack_integrity.py to validate alignment and general repository scenarios against required cases.
- Modified test_runtime_view.py to ensure retired references are excluded from the runtime view.
- Introduced new testing recipes and guidelines for Python, shell, and mixed environments to improve test design, speed, and proof of concept.
- Added new evaluation criteria for test authoring, focusing on observable behavior and meaningful cases.
- Implemented new test cases for validating identifiers and handling output files, ensuring real behavior is tested without reliance on mocks.
- Updated existing tests to ensure they adhere to the new guidelines, including isolation of mutable state and explicit subprocess management.
- Enhanced documentation to clarify testing workflows and expectations, including the use of pytest and maintaining a consistent framework.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant