feat: V28082026 improve Refactor skills and enhance validation processes - #91
Open
diegoitaliait wants to merge 150 commits into
Open
diegoitaliait wants to merge 150 commits into
diegoitaliait wants to merge 150 commits into
Conversation
- Deleted the `mattpocock-writing-great-skills` skill and its associated files, including SKILL.md and openai.yaml. - Updated tests to remove references to the deleted skill and adjusted assertions accordingly. - Added new skills `grill-me` and `grilling` to the resource management. - Enhanced the normalization tests to ensure proper handling of legacy paths and invocation policies. - Improved the test suite to validate the presence and correctness of new skills and their configurations. - Ensured that the internal skill creator uses current entry points and does not reference removed skills.
…reserves_declared_scope
…n in test_manifest_references
…tions - Introduced tests for local sync repos agent metadata and bundle paths. - Enhanced code analysis workflow to run pytest on skill tests. - Updated Makefile to include skill test paths in linting and testing. - Documented the validation surface for contract checks in tech documentation. - Configured pytest to discover tests in both root and skill directories. - Implemented a scoring mechanism for codebase-improvement gateway evaluations. - Added fixtures for passing and failing runs in gateway evaluation tests. - Created comprehensive tests for gateway evaluation scoring logic. - Ensured Makefile does not reference skill runtime scripts directly. - Validated the layout of skill tests to ensure they are co-located with their bundles.
grill-me was blocked in both runtimes and the sync manifest regenerated the flags on every refresh, so any skill mandating /grill-me could not load it and silently improvised instead of interviewing the user. - Remove disable-model-invocation from the grill-me bundle, drop policy.allow_implicit_invocation from its Codex metadata, and delete the invocation_policy entry that regenerated both. - Update the sync contract split from 14/11 to 13/12 user-invoked and model-invoked Matt skills. - Drop /superpowers-brainstorming from the internal-gateway-idea forbidden pre-acceptance routes; a user-invoked-only skill is not an agent route before or after acceptance. - Route material internal-agent-creator questions through /grill-me instead of an ad-hoc prompt.
- Added `validate-skill-change-scope.py` to enforce strict allowlisting for protected external skill changes. - Updated `validate_internal_skills.py` to align with new structure and import paths. - Refactored various scripts and documentation to reflect changes in script naming conventions (e.g., `validate_internal_skills.py` to `validate-internal-skills.py`). - Modified Makefile and workflow configurations to accommodate new script names and paths. - Enhanced documentation to clarify usage and structure of scripts related to skill validation and catalog consistency. - Removed obsolete test for subprocess timeout handling. - Updated tests to ensure compatibility with new script paths and naming conventions.
- Removed direct calls to `validate-internal-skills.py` in various SKILL.md files, replacing them with a unified command using `run.sh` for consistency. - Added new reference files for decision ledger and candidate persistence to streamline decision-making processes. - Updated manifest references to clarify execution and authoring responsibilities. - Improved report formatting guidelines and shapes for findings, residuals, and open questions. - Enhanced command portability documentation to ensure clarity on executable commands and validation processes. - Adjusted test scripts to align with new directory structure and validation methods.
- Introduced `test-rules.py` to validate skill prose assertions and internal skill findings. - Created `test-scope.py` to ensure proper classification of protected skill bundles and validate allowlist functionality. - Added `test-runner.py` to verify script resolution for various tools and ensure proper handling of protected skill scope validators. - Implemented `test-rule-coverage.py` to check coverage of token rules and ensure expected findings are produced. - Updated `test_makefile_skill_script_references.py` to reflect changes in script runner references. - Enhanced `test_repository_test_layout_contract.py` to validate the structure of GitHub automation layout and ensure test paths are correctly identified.
…f /grill-me and related processes
…sts for missing references in OpenAI metadata
…hance file matching patterns and documentation generation
- Implemented tests for the knowledge configuration contract in `test_config.py`, ensuring valid configurations and error handling for unknown rules. - Created tests for inventory discovery in `test_inventory.py`, validating component detection and reporting capabilities. - Added behavior tests for operating modes in `test_modes.py`, covering audit and update functionalities. - Introduced neutral tests for deterministic checks in `test_check.py`, verifying the behavior of the knowledge check process.
…mentation and inventory
…on, inventory, and modes - Deleted test_check.py, test_config.py, test_inventory.py, and test_modes.py - These files contained neutral tests for various functionalities related to knowledge checks, configuration validation, inventory discovery, and operating modes. - The removal is part of a cleanup effort to streamline the testing suite and remove outdated or redundant tests.
…s for octal and truthy values
…ication; add guidance on invocation boundaries and deferring lessons
…age guidance, new references, and maintainability checks
…ance reference and update related descriptions
…rkflow, diagram disposition, and ownership evidence
…rity on evidence handling and domain promotion
…source synchronization; enhance clarity on asset management and synchronization processes
…p and refresh modes with setup and sync, introduce bucket mechanism, and enhance evaluation scenarios
…aid documentation; clarify styling and compatibility considerations
…c; clarify plan authoring readiness and separate implementation permissions
…lan authoring readiness and permissions for implementation
… and question formatting
…just asset counts and references
…nning; enhance inventory.py to exclude manifest support files from imported skills in tests
…lement and mattpocock-to-spec
…cated code - Removed unused imports and constants related to protocol markers and expected events. - Simplified the extraction of metadata and evaluation cases from YAML and JSON files. - Consolidated test cases to focus on new authoring contracts and metadata validation. - Ensured retired gateway protocol files and routes are absent from the repository. - Updated tests to reflect changes in the protocol structure and expectations.
…ferences to manage output paths and instruction overrides, ensuring compliance with legacy entries and improving artifact handling in tmp directory.
- Updated knowledge-navigation.md to clarify routing to maintained owners and ownership reporting. - Expanded knowledge-report.md with a new section for reporting problems found, including a structured table format. - Revised knowledge-scope.md to specify requirements for authored deletions and enforcement gap reporting. - Improved knowledge-topology.md by refining definitions of common core and lanes, and added claim classification guidance. - Adjusted managerial-maintenance.md to clarify roadmap creation criteria and ownership linking. - Modified project-memory-maintenance.md to emphasize the separation of execution tasks from orientation entries. - Enhanced readme-maintenance.md to specify conditions for including a table of contents and operational detail relocation. - Updated standards-maintenance.md to clarify reporting gaps and avoid creating empty entries. - Refined evaluation tests in bindings.json, compatibility.json, evals.json, and others to align with new documentation standards and ensure proper coverage. - Added new tests for problem reporting, deletion approval, and knowledge layer preservation in test_eval_graders.py and test_eval_pack_integrity.py. - Improved runtime view tests to ensure correct copying of public projections and selected references.
…ove readability of assertions
…r generated plans and enhance documentation for clarity
…irements for clarity and consistency in execution processes
…zation to enforce single implementation plans and retire deprecated skills
…kills from common skills list
…files for improved clarity and security
…ance and personal study repositories - Created F7.tree.json and F8.tree.json to represent a non-code community grant governance repository and a personal study repository, respectively. - Added corresponding gold files F7.gold.json and F8.gold.json to define expected outcomes for the new fixtures. - Enhanced test coverage in test_bundle_contract.py to ensure knowledge types and alignment references are comprehensive. - Updated test_eval_graders.py to remove obsolete lifecycle checks and ensure proper bindings for evaluation scenarios. - Improved test_eval_pack_integrity.py to validate alignment and general repository scenarios against required cases. - Modified test_runtime_view.py to ensure retired references are excluded from the runtime view.
…lls and remove redundant settings
…nd improve shell command execution
- Introduced new testing recipes and guidelines for Python, shell, and mixed environments to improve test design, speed, and proof of concept. - Added new evaluation criteria for test authoring, focusing on observable behavior and meaningful cases. - Implemented new test cases for validating identifiers and handling output files, ensuring real behavior is tested without reliance on mocks. - Updated existing tests to ensure they adhere to the new guidelines, including isolation of mutable state and explicit subprocess management. - Enhanced documentation to clarify testing workflows and expectations, including the use of pytest and maintaining a consistent framework.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
Refactor and remove deprecated skills while introducing new skills and improving validation processes. This update enhances the test suite and validation scripts to ensure consistency and correctness across the skill framework.
Change Type
List of changes
Consumer Impact
Testing Instructions
Validation Evidence
Breaking Changes
Checklist