diff --git a/CLAUDE.md b/CLAUDE.md index 96631ea..5285b85 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -9,6 +9,12 @@ the full roadmap and agent breakdown lives in [docs/PLAN.md](docs/PLAN.md) and **Tool-Semantics — Behavioral regression testing for MCP and AI-agent interfaces.** Know when an MCP change breaks the agent, not just the schema. +**Schema-valid ≠ agent-safe.** + +Product framing: **DETECT** (capture/compare) → **TEST** (probes) → **PROTECT** (CI). +Beginner path: capture → compare + probe (unified `eval` is #76). +Docs: [docs/simple-explanation.md](docs/simple-explanation.md), +[docs/index.md](docs/index.md). It snapshots a tool interface, diffs baseline vs. candidate, and reports risk across five compatibility layers: @@ -38,7 +44,8 @@ Streamable HTTP, and legacy SSE. - `src/tool_semantics/cli.py` — Typer CLI entrypoint - `tests/` — pytest suite, one file per module area - `examples/` — demo MCP-style manifests (GitHub server v1/v2, weather) -- `docs/` — architecture, config, change-codes, github-action, publishing, adapters +- `docs/` — [index](docs/index.md), simple-explanation, concepts, architecture, + config, change-codes, probes, github-action, publishing, adapters ## Dev commands @@ -53,6 +60,7 @@ ruff format --check . tool-semantics capture examples/github_server_v1.json -o .tool-semantics/v1.json tool-semantics compare .tool-semantics/v1.json .tool-semantics/v2.json --markdown-output report.md +tool-semantics probe .tool-semantics/v2.json --probes examples/probes/github_v1_offline.json tool-semantics capture-mcp -o snap.json https://example.com/mcp ``` diff --git a/README.md b/README.md index 22a7b2e..eee6663 100644 --- a/README.md +++ b/README.md @@ -9,7 +9,8 @@

- Behavioral compatibility testing for MCP tools and AI-agent interfaces. + Behavioral compatibility testing for MCP tools and AI-agent interfaces.
+ Schema-valid ≠ agent-safe.

@@ -23,14 +24,29 @@ ## Why Tool-Semantics? -AI agents do not call tools the way typed clients do. They choose tools from **descriptions**, invent **arguments** from schemas, and infer **side effects** from naming and prose. A change that remains JSON-Schema-valid can still: +AI agents do not call tools the way typed clients do. They choose tools from +**descriptions**, invent **arguments** from schemas, and infer **side effects** +from naming and prose. A change that remains JSON-Schema-valid can still: - steer the model toward the wrong tool - drop a required argument the model used to omit - rename enums the model still emits - quietly escalate from read-only to write/destructive behavior -**Tool-Semantics** captures normalized tool-interface snapshots and diffs them for structural *and* semantic risk — so teams can gate MCP and tool-API changes before agents ship broken workflows. +**Schema-valid ≠ agent-safe.** Tool-Semantics captures normalized tool-interface +snapshots and evaluates them for structural *and* semantic risk — so teams can +gate MCP and tool-API changes before agents ship broken workflows. + +New here? Read the [simple explanation](docs/simple-explanation.md) and +[concepts](docs/concepts.md). Full doc index: [docs/index.md](docs/index.md). + +## DETECT · TEST · PROTECT + +| | Job | Commands | +| --- | --- | --- | +| **DETECT** | Snapshot interfaces; see what changed | `capture`, `capture-mcp`, `compare` | +| **TEST** | Check tool selection / args / risk expectations | `probe` (offline or `--model`) | +| **PROTECT** | Fail CI when policy says so | exit codes, Action, [config](docs/config.md) | ## Compatibility layers @@ -42,7 +58,9 @@ AI agents do not call tools the way typed clients do. They choose tools from **d | 4. Execution | Do calls still succeed with prior argument patterns? | | 5. Intent / side effects | Did risk, confirmation needs, or outcomes change? | -The MVP implements deterministic interface snapshots and structural comparison (layers 1–2, with warnings that point at 3–5), plus live MCP capture over stdio, Streamable HTTP, and legacy SSE. Model-based behavioral testing is available as an opt-in library; see the [roadmap](ROADMAP.md). +Layers 1–2 are CI-gateable today. Layers 3–5 are covered by offline probes plus +**opt-in** model-backed probes. Live capture supports stdio, Streamable HTTP, +and legacy SSE. See the [roadmap](ROADMAP.md). ## How it works @@ -63,22 +81,30 @@ flowchart LR E --> F[CLI / CI exit codes] ``` -## Quick start +## Quick start — capture → evaluate ```bash python -m venv .venv source .venv/bin/activate pip install -e ".[dev]" -# Capture two interface versions +# 1. Capture baseline and candidate tool-semantics capture examples/github_server_v1.json -o .tool-semantics/v1.json tool-semantics capture examples/github_server_v2.json -o .tool-semantics/v2.json -# Compare — exits 1 on breaking/critical changes +# 2. Evaluate — structural diff + behavioral probes tool-semantics compare .tool-semantics/v1.json .tool-semantics/v2.json \ --markdown-output .tool-semantics/report.md +tool-semantics probe .tool-semantics/v2.json \ + --probes examples/probes/github_v1_offline.json + +# 3. Protect — non-zero exit fails CI (see docs/github-action.md) ``` +A unified `tool-semantics eval` command will combine compare + probes into one +beginner entrypoint ([#76](https://github.com/askmy-stack/tool-semantics/issues/76)). +Until then, use **capture → compare + probe** as the evaluate path. + ### Demo

@@ -132,6 +158,11 @@ print("compatible:", report.is_compatible) ## CLI reference +**Beginner path:** `capture` → `compare` + `probe` (evaluate) → CI protect. + +**Advanced / secondary** (linked below): provenance, `--config`, `--model`, +`--trials`, verbose logs, library APIs. + ```bash tool-semantics --version tool-semantics capture [-o .tool-semantics/snapshot.json] \ @@ -172,12 +203,23 @@ The optional provenance sidecar records capture context and a digest of the snapshot. It is separate from the snapshot and never affects compatibility comparisons. -JSON reports include `changes`, `is_compatible`, and `counts` by severity. -Change-code catalog: [docs/change-codes.md](docs/change-codes.md). -Ignore-config schema: [docs/config.md](docs/config.md). -GitHub Action: [docs/github-action.md](docs/github-action.md). -Publishing: [docs/publishing.md](docs/publishing.md). -Migration adapters: [docs/adapters.md](docs/adapters.md). +JSON reports include `changes`, `is_compatible`, and `counts` by severity. + +### Docs + +| | | +| --- | --- | +| Index | [docs/index.md](docs/index.md) | +| Simple explanation | [docs/simple-explanation.md](docs/simple-explanation.md) | +| Concepts | [docs/concepts.md](docs/concepts.md) | +| Change codes | [docs/change-codes.md](docs/change-codes.md) | +| Architecture | [docs/architecture.md](docs/architecture.md) | +| Config / policy | [docs/config.md](docs/config.md) | +| Probes | [docs/probes.md](docs/probes.md) | +| GitHub Action | [docs/github-action.md](docs/github-action.md) | +| Publishing | [docs/publishing.md](docs/publishing.md) | +| Adapters | [docs/adapters.md](docs/adapters.md) | +| Milestone 7+ plan | [docs/PLAN.md](docs/PLAN.md), [docs/AGENT_EXECUTION.md](docs/AGENT_EXECUTION.md) | ### Optional `risk` field @@ -212,19 +254,20 @@ assert report.passed ## Project layout ```text -src/tool_semantics/ # scanner, models, diff engine, report, CLI -examples/ # demo MCP-style manifests +src/tool_semantics/ # scanner, models, diff engine, probes, report, CLI +examples/ # demo MCP-style manifests + probe fixtures tests/ # pytest suite +docs/ # index, concepts, architecture, change-codes, … docs/assets/ # README visuals ``` ## Roadmap -See [ROADMAP.md](ROADMAP.md) for milestones. Live MCP supports stdio, Streamable -HTTP, and legacy SSE ([docs/mcp-versions.md](docs/mcp-versions.md)). -Model-backed probes / metrics / stability are available via the library -API (`docs/probes.md`). Downstream consumers (myelinmesh, dogfood capture): -[`docs/downstream.md`](docs/downstream.md). +See [ROADMAP.md](ROADMAP.md) and [docs/PLAN.md](docs/PLAN.md) for Milestone 7+. +Live MCP supports stdio, Streamable HTTP, and legacy SSE +([docs/mcp-versions.md](docs/mcp-versions.md)). Model-backed probes: +[docs/probes.md](docs/probes.md). Downstream consumers: +[docs/downstream.md](docs/downstream.md). ## Contributing diff --git a/ROADMAP.md b/ROADMAP.md index 608ab8c..1bbd153 100644 --- a/ROADMAP.md +++ b/ROADMAP.md @@ -96,6 +96,7 @@ - [ ] Corpus / mutations / research harness — [#92](https://github.com/askmy-stack/tool-semantics/issues/92)–[#95](https://github.com/askmy-stack/tool-semantics/issues/95) **P2/P3** - [ ] Confidence intervals + dev/test/verified split — [#115](https://github.com/askmy-stack/tool-semantics/issues/115), [#116](https://github.com/askmy-stack/tool-semantics/issues/116) **P3** - [ ] DX: init/doctor/lint/audit/generate-probes — [#96](https://github.com/askmy-stack/tool-semantics/issues/96)–[#98](https://github.com/askmy-stack/tool-semantics/issues/98) **P2** +- [x] Docs / README capture+eval positioning — [#101](https://github.com/askmy-stack/tool-semantics/issues/101) **P2** - [x] ROADMAP/PLAN + agent execution spec — [#103](https://github.com/askmy-stack/tool-semantics/issues/103) - [ ] Apply GitHub priority labels — [#118](https://github.com/askmy-stack/tool-semantics/issues/118) **P0** - Operational spec: [docs/AGENT_EXECUTION.md](docs/AGENT_EXECUTION.md) diff --git a/docs/PLAN.md b/docs/PLAN.md index a6dfe72..aa9bf31 100644 --- a/docs/PLAN.md +++ b/docs/PLAN.md @@ -1,7 +1,9 @@ # Course of action (Milestones 7+) Primary operational spec: [AGENT_EXECUTION.md](AGENT_EXECUTION.md). -Product positioning: behavioral regression for MCP / AI-agent interfaces. +Product positioning: **schema-valid ≠ agent-safe** — behavioral regression for +MCP / AI-agent interfaces ([simple-explanation.md](simple-explanation.md), +[concepts.md](concepts.md), [index.md](index.md)). Shipped through **v0.4.0**: Milestones 0–6 (structural detect, stdio/SSE capture, offline + library model probes, policy, Action, adapters). @@ -13,6 +15,9 @@ offline + library model probes, policy, Action, adapters). 3. **Evaluate real behavior** — collision/rename, traces, discovery, safety/output (#79–#91) 4. **Prove** — corpus, mutations, research harness, DX/docs (#92–#103) +Beginner UX today: **capture → compare + probe**. Unified `eval` (#76) becomes +the single evaluate entrypoint when merged. + ## Immediate next PRs | Order | Issues | Priority | Outcome | diff --git a/docs/architecture.md b/docs/architecture.md index acceb6c..9069f6c 100644 --- a/docs/architecture.md +++ b/docs/architecture.md @@ -2,6 +2,8 @@ Tool-Semantics separates **transport**, **normalization**, and **compatibility analysis** so the same diff engine can run on static manifests and live MCP servers. +Doc index / vocabulary: [index.md](index.md), [concepts.md](concepts.md). + ## Pipeline ```mermaid diff --git a/docs/concepts.md b/docs/concepts.md new file mode 100644 index 0000000..cd99ba0 --- /dev/null +++ b/docs/concepts.md @@ -0,0 +1,77 @@ +# Concepts and shared vocabulary + +Short definitions used across Tool-Semantics docs and reports. + +## Snapshots + +A normalized **InterfaceSnapshot** of an MCP / tool catalog: tool names, +descriptions, parameters, optional output schemas, optional `risk`, plus +protocol/server metadata. Snapshots are the **canonical compare input** — +not live traffic. + +- Capture from a JSON manifest: `capture` +- Capture from a live server: `capture-mcp` (stdio / Streamable HTTP / SSE) + +See [architecture.md](architecture.md), [mcp-versions.md](mcp-versions.md). + +## Diff / change codes + +Deterministic comparison of baseline vs candidate snapshots. Each finding has a +stable **code** (e.g. `tool.removed`), **severity** (`info` → `critical`), +subject, and message. + +Catalog: [change-codes.md](change-codes.md). +Semantics note: description / rename warnings matter because models route on +prose, not only types. + +## Semantics (agent behavior) + +Beyond schema validity: + +| Idea | Meaning | +| --- | --- | +| Tool selection | Will the model still pick the intended tool? | +| Argument patterns | Will prior argument shapes still validate? | +| Risk / confirmation | Did side-effect expectations change? | + +## Probes + +Behavioral checks against a snapshot. + +- **Offline** — no LLM; assert expected tools / params / risk exist. +- **Model-backed** — opt-in runner selects a tool for an intent (`approved` + probes by default). + +See [probes.md](probes.md). + +## Traces + +Recorded agent trajectories (tool calls over time) for replay and regression. +Trace schema / replay land in Milestone work (#83–#84); not required for the +beginner capture → evaluate path. + +## Safety + +Optional tool-level `risk` (`read_only`, `external_write`, `destructive`, +`unknown`) and confirmation expectations. Missing risk defaults to `unknown` +(no false escalation). Broader safety/scope diffs are roadmap items (#87). + +## Policies + +Release knobs that map severities to CI failure +(`compatible` / `breaking` / `critical-only` / …). Config: +[config.md](config.md). + +## Benchmarks / research + +Corpus cases, splits (DEV/TEST/VERIFIED), mutations, and sensitivity harnesses +prove detectors without contaminating published scores. See ROADMAP Milestone +12+ and [AGENT_EXECUTION.md](AGENT_EXECUTION.md). + +## DETECT · TEST · PROTECT + +Product framing used in the README and [simple-explanation.md](simple-explanation.md): + +1. **DETECT** — snapshots + structural diff +2. **TEST** — probes (and later traces / eval) +3. **PROTECT** — policies, Action, exit codes diff --git a/docs/index.md b/docs/index.md new file mode 100644 index 0000000..02a2159 --- /dev/null +++ b/docs/index.md @@ -0,0 +1,49 @@ +# Documentation index + +Start here if you are new: [simple-explanation.md](simple-explanation.md) → +[concepts.md](concepts.md) → beginner path on the [README](../README.md). + +## By topic + +| Topic | Doc | +| --- | --- | +| Plain-language overview | [simple-explanation.md](simple-explanation.md) | +| Shared vocabulary | [concepts.md](concepts.md) | +| Architecture / pipeline | [architecture.md](architecture.md) | +| Snapshots & live MCP capture | [mcp-versions.md](mcp-versions.md), [adr-live-mcp-capture.md](adr-live-mcp-capture.md) | +| Structural diff & change codes | [change-codes.md](change-codes.md) | +| Semantics & agent risk | [concepts.md](concepts.md)#semantics-agent-behavior | +| Probes (offline / model) | [probes.md](probes.md) | +| Traces (roadmap) | [AGENT_EXECUTION.md](AGENT_EXECUTION.md), ROADMAP #83–#84 | +| Safety / risk field | [concepts.md](concepts.md)#safety, [probes.md](probes.md) | +| Policies & ignore config | [config.md](config.md) | +| Benchmarks / research | [AGENT_EXECUTION.md](AGENT_EXECUTION.md), [PLAN.md](PLAN.md) | +| GitHub Action (protect) | [github-action.md](github-action.md) | +| Migration adapters | [adapters.md](adapters.md) | +| Downstream consumers | [downstream.md](downstream.md) | +| Publishing | [publishing.md](publishing.md) | + +## Beginner vs advanced CLI + +**Beginner (capture → evaluate → protect)** + +1. `capture` / `capture-mcp` +2. `compare` + `probe` +3. CI Action / non-zero exit + +**Advanced / secondary** + +- `--config` ignore rules, provenance sidecars, verbose stderr logs +- `--model` / `--trials` model-backed probes +- Library APIs for custom runners, adapters, policies + +Unified `eval` ([#76](https://github.com/askmy-stack/tool-semantics/issues/76)) +will become the single evaluate entrypoint; `compare` and `probe` remain for +advanced workflows. + +## Milestone 7+ plan + +- [PLAN.md](PLAN.md) — course of action +- [AGENT_EXECUTION.md](AGENT_EXECUTION.md) — binding priority tables +- [../ROADMAP.md](../ROADMAP.md) — checklist +- [../CLAUDE.md](../CLAUDE.md) — short agent guidance diff --git a/docs/simple-explanation.md b/docs/simple-explanation.md new file mode 100644 index 0000000..5a5182f --- /dev/null +++ b/docs/simple-explanation.md @@ -0,0 +1,54 @@ +# Tool-Semantics in plain language + +**Schema-valid is not the same as agent-safe.** + +Typed clients fail loudly when a field disappears. Language-model agents often +**guess**: they pick tools from descriptions, invent arguments from schemas, and +infer side effects from names. A change that still validates as JSON Schema can +still break the agent. + +Tool-Semantics answers one question: + +> If we ship this MCP / tool-interface change, will agents still behave? + +## Three jobs: DETECT · TEST · PROTECT + +| Job | What you do | Typical command | +| --- | --- | --- | +| **DETECT** | Snapshot the interface; see what changed | `capture`, `capture-mcp`, `compare` | +| **TEST** | Check whether agents still pick tools / args correctly | `probe` (offline or `--model`) | +| **PROTECT** | Gate merges in CI with policies and reports | GitHub Action, exit codes, config | + +## Beginner path + +1. **Capture** a baseline and a candidate snapshot. +2. **Evaluate** with structural compare plus behavioral probes. +3. **Protect** the merge when the report fails policy. + +```bash +# 1. Capture +tool-semantics capture examples/github_server_v1.json -o .tool-semantics/v1.json +tool-semantics capture examples/github_server_v2.json -o .tool-semantics/v2.json + +# 2. Evaluate (structural + behavioral) +tool-semantics compare .tool-semantics/v1.json .tool-semantics/v2.json \ + --markdown-output .tool-semantics/report.md +tool-semantics probe .tool-semantics/v2.json \ + --probes examples/probes/github_v1_offline.json + +# 3. Protect — non-zero exit fails CI (see docs/github-action.md) +``` + +A unified `eval` command ([#76](https://github.com/askmy-stack/tool-semantics/issues/76)) +will combine compare + probes into one beginner entrypoint; until then, use the +two commands above. + +## What “breaking” means here + +- **Breaking / critical** structural codes fail `compare` by default + ([change-codes.md](change-codes.md)). +- **Warnings** (e.g. description drift) flag selection risk without always + failing CI — agents may re-route even when schemas remain valid. +- **Probes** catch behavioral regressions that pure schema diffs miss. + +Next: [concepts.md](concepts.md) · [index.md](index.md) · [README](../README.md) diff --git a/tests/test_docs_positioning.py b/tests/test_docs_positioning.py new file mode 100644 index 0000000..0e4dbb7 --- /dev/null +++ b/tests/test_docs_positioning.py @@ -0,0 +1,37 @@ +"""Smoke tests for docs positioning (#101).""" + +from __future__ import annotations + +from pathlib import Path + +ROOT = Path(__file__).resolve().parents[1] + + +def test_docs_index_and_beginner_pages_exist() -> None: + for relative in ( + "docs/index.md", + "docs/simple-explanation.md", + "docs/concepts.md", + ): + path = ROOT / relative + assert path.is_file(), relative + text = path.read_text(encoding="utf-8") + assert len(text) > 100 + + +def test_readme_positions_capture_evaluate_and_schema_valid() -> None: + readme = (ROOT / "README.md").read_text(encoding="utf-8") + assert "Schema-valid ≠ agent-safe" in readme or "schema-valid ≠ agent-safe" in readme.lower() + assert "DETECT" in readme and "TEST" in readme and "PROTECT" in readme + assert "capture → evaluate" in readme.lower() or "capture → evaluate" in readme + assert "docs/simple-explanation.md" in readme + assert "docs/index.md" in readme + + +def test_claude_and_plan_point_at_milestone7_docs() -> None: + claude = (ROOT / "CLAUDE.md").read_text(encoding="utf-8") + plan = (ROOT / "docs/PLAN.md").read_text(encoding="utf-8") + assert "docs/PLAN.md" in claude + assert "AGENT_EXECUTION.md" in claude + assert "simple-explanation.md" in plan + assert "capture → compare" in plan or "capture → compare + probe" in plan