From 9bc4a8492336b824840eada5d6377f270003b4c2 Mon Sep 17 00:00:00 2001
From: Cursor Agent
Date: Wed, 16 Sep 2026 15:54:13 +0000
Subject: [PATCH] =?UTF-8?q?docs:=20restructure=20README=20for=20capture?=
=?UTF-8?q?=E2=86=92evaluate=20and=20DETECT/TEST/PROTECT=20(#101)?=
MIME-Version: 1.0
Content-Type: text/plain; charset=UTF-8
Content-Transfer-Encoding: 8bit
Add simple-explanation, concepts, and docs index; align CLAUDE/PLAN pointers
with Milestone 7+ and schema-valid ≠ agent-safe positioning.
Co-authored-by: Abhinaysai Kamineni
---
CLAUDE.md | 10 +++-
README.md | 83 ++++++++++++++++++++++++++--------
ROADMAP.md | 1 +
docs/PLAN.md | 7 ++-
docs/architecture.md | 2 +
docs/concepts.md | 77 +++++++++++++++++++++++++++++++
docs/index.md | 49 ++++++++++++++++++++
docs/simple-explanation.md | 54 ++++++++++++++++++++++
tests/test_docs_positioning.py | 37 +++++++++++++++
9 files changed, 298 insertions(+), 22 deletions(-)
create mode 100644 docs/concepts.md
create mode 100644 docs/index.md
create mode 100644 docs/simple-explanation.md
create mode 100644 tests/test_docs_positioning.py
diff --git a/CLAUDE.md b/CLAUDE.md
index 96631ea..5285b85 100644
--- a/CLAUDE.md
+++ b/CLAUDE.md
@@ -9,6 +9,12 @@ the full roadmap and agent breakdown lives in [docs/PLAN.md](docs/PLAN.md) and
**Tool-Semantics — Behavioral regression testing for MCP and AI-agent interfaces.**
Know when an MCP change breaks the agent, not just the schema.
+**Schema-valid ≠ agent-safe.**
+
+Product framing: **DETECT** (capture/compare) → **TEST** (probes) → **PROTECT** (CI).
+Beginner path: capture → compare + probe (unified `eval` is #76).
+Docs: [docs/simple-explanation.md](docs/simple-explanation.md),
+[docs/index.md](docs/index.md).
It snapshots a tool interface, diffs baseline vs. candidate, and reports risk
across five compatibility layers:
@@ -38,7 +44,8 @@ Streamable HTTP, and legacy SSE.
- `src/tool_semantics/cli.py` — Typer CLI entrypoint
- `tests/` — pytest suite, one file per module area
- `examples/` — demo MCP-style manifests (GitHub server v1/v2, weather)
-- `docs/` — architecture, config, change-codes, github-action, publishing, adapters
+- `docs/` — [index](docs/index.md), simple-explanation, concepts, architecture,
+ config, change-codes, probes, github-action, publishing, adapters
## Dev commands
@@ -53,6 +60,7 @@ ruff format --check .
tool-semantics capture examples/github_server_v1.json -o .tool-semantics/v1.json
tool-semantics compare .tool-semantics/v1.json .tool-semantics/v2.json --markdown-output report.md
+tool-semantics probe .tool-semantics/v2.json --probes examples/probes/github_v1_offline.json
tool-semantics capture-mcp -o snap.json https://example.com/mcp
```
diff --git a/README.md b/README.md
index 22a7b2e..eee6663 100644
--- a/README.md
+++ b/README.md
@@ -9,7 +9,8 @@
- Behavioral compatibility testing for MCP tools and AI-agent interfaces.
+ Behavioral compatibility testing for MCP tools and AI-agent interfaces.
+ Schema-valid ≠ agent-safe.
@@ -23,14 +24,29 @@
## Why Tool-Semantics?
-AI agents do not call tools the way typed clients do. They choose tools from **descriptions**, invent **arguments** from schemas, and infer **side effects** from naming and prose. A change that remains JSON-Schema-valid can still:
+AI agents do not call tools the way typed clients do. They choose tools from
+**descriptions**, invent **arguments** from schemas, and infer **side effects**
+from naming and prose. A change that remains JSON-Schema-valid can still:
- steer the model toward the wrong tool
- drop a required argument the model used to omit
- rename enums the model still emits
- quietly escalate from read-only to write/destructive behavior
-**Tool-Semantics** captures normalized tool-interface snapshots and diffs them for structural *and* semantic risk — so teams can gate MCP and tool-API changes before agents ship broken workflows.
+**Schema-valid ≠ agent-safe.** Tool-Semantics captures normalized tool-interface
+snapshots and evaluates them for structural *and* semantic risk — so teams can
+gate MCP and tool-API changes before agents ship broken workflows.
+
+New here? Read the [simple explanation](docs/simple-explanation.md) and
+[concepts](docs/concepts.md). Full doc index: [docs/index.md](docs/index.md).
+
+## DETECT · TEST · PROTECT
+
+| | Job | Commands |
+| --- | --- | --- |
+| **DETECT** | Snapshot interfaces; see what changed | `capture`, `capture-mcp`, `compare` |
+| **TEST** | Check tool selection / args / risk expectations | `probe` (offline or `--model`) |
+| **PROTECT** | Fail CI when policy says so | exit codes, Action, [config](docs/config.md) |
## Compatibility layers
@@ -42,7 +58,9 @@ AI agents do not call tools the way typed clients do. They choose tools from **d
| 4. Execution | Do calls still succeed with prior argument patterns? |
| 5. Intent / side effects | Did risk, confirmation needs, or outcomes change? |
-The MVP implements deterministic interface snapshots and structural comparison (layers 1–2, with warnings that point at 3–5), plus live MCP capture over stdio, Streamable HTTP, and legacy SSE. Model-based behavioral testing is available as an opt-in library; see the [roadmap](ROADMAP.md).
+Layers 1–2 are CI-gateable today. Layers 3–5 are covered by offline probes plus
+**opt-in** model-backed probes. Live capture supports stdio, Streamable HTTP,
+and legacy SSE. See the [roadmap](ROADMAP.md).
## How it works
@@ -63,22 +81,30 @@ flowchart LR
E --> F[CLI / CI exit codes]
```
-## Quick start
+## Quick start — capture → evaluate
```bash
python -m venv .venv
source .venv/bin/activate
pip install -e ".[dev]"
-# Capture two interface versions
+# 1. Capture baseline and candidate
tool-semantics capture examples/github_server_v1.json -o .tool-semantics/v1.json
tool-semantics capture examples/github_server_v2.json -o .tool-semantics/v2.json
-# Compare — exits 1 on breaking/critical changes
+# 2. Evaluate — structural diff + behavioral probes
tool-semantics compare .tool-semantics/v1.json .tool-semantics/v2.json \
--markdown-output .tool-semantics/report.md
+tool-semantics probe .tool-semantics/v2.json \
+ --probes examples/probes/github_v1_offline.json
+
+# 3. Protect — non-zero exit fails CI (see docs/github-action.md)
```
+A unified `tool-semantics eval` command will combine compare + probes into one
+beginner entrypoint ([#76](https://github.com/askmy-stack/tool-semantics/issues/76)).
+Until then, use **capture → compare + probe** as the evaluate path.
+
### Demo
@@ -132,6 +158,11 @@ print("compatible:", report.is_compatible)
## CLI reference
+**Beginner path:** `capture` → `compare` + `probe` (evaluate) → CI protect.
+
+**Advanced / secondary** (linked below): provenance, `--config`, `--model`,
+`--trials`, verbose logs, library APIs.
+
```bash
tool-semantics --version
tool-semantics capture [-o .tool-semantics/snapshot.json] \
@@ -172,12 +203,23 @@ The optional provenance sidecar records capture context and a digest of the
snapshot. It is separate from the snapshot and never affects compatibility
comparisons.
-JSON reports include `changes`, `is_compatible`, and `counts` by severity.
-Change-code catalog: [docs/change-codes.md](docs/change-codes.md).
-Ignore-config schema: [docs/config.md](docs/config.md).
-GitHub Action: [docs/github-action.md](docs/github-action.md).
-Publishing: [docs/publishing.md](docs/publishing.md).
-Migration adapters: [docs/adapters.md](docs/adapters.md).
+JSON reports include `changes`, `is_compatible`, and `counts` by severity.
+
+### Docs
+
+| | |
+| --- | --- |
+| Index | [docs/index.md](docs/index.md) |
+| Simple explanation | [docs/simple-explanation.md](docs/simple-explanation.md) |
+| Concepts | [docs/concepts.md](docs/concepts.md) |
+| Change codes | [docs/change-codes.md](docs/change-codes.md) |
+| Architecture | [docs/architecture.md](docs/architecture.md) |
+| Config / policy | [docs/config.md](docs/config.md) |
+| Probes | [docs/probes.md](docs/probes.md) |
+| GitHub Action | [docs/github-action.md](docs/github-action.md) |
+| Publishing | [docs/publishing.md](docs/publishing.md) |
+| Adapters | [docs/adapters.md](docs/adapters.md) |
+| Milestone 7+ plan | [docs/PLAN.md](docs/PLAN.md), [docs/AGENT_EXECUTION.md](docs/AGENT_EXECUTION.md) |
### Optional `risk` field
@@ -212,19 +254,20 @@ assert report.passed
## Project layout
```text
-src/tool_semantics/ # scanner, models, diff engine, report, CLI
-examples/ # demo MCP-style manifests
+src/tool_semantics/ # scanner, models, diff engine, probes, report, CLI
+examples/ # demo MCP-style manifests + probe fixtures
tests/ # pytest suite
+docs/ # index, concepts, architecture, change-codes, …
docs/assets/ # README visuals
```
## Roadmap
-See [ROADMAP.md](ROADMAP.md) for milestones. Live MCP supports stdio, Streamable
-HTTP, and legacy SSE ([docs/mcp-versions.md](docs/mcp-versions.md)).
-Model-backed probes / metrics / stability are available via the library
-API (`docs/probes.md`). Downstream consumers (myelinmesh, dogfood capture):
-[`docs/downstream.md`](docs/downstream.md).
+See [ROADMAP.md](ROADMAP.md) and [docs/PLAN.md](docs/PLAN.md) for Milestone 7+.
+Live MCP supports stdio, Streamable HTTP, and legacy SSE
+([docs/mcp-versions.md](docs/mcp-versions.md)). Model-backed probes:
+[docs/probes.md](docs/probes.md). Downstream consumers:
+[docs/downstream.md](docs/downstream.md).
## Contributing
diff --git a/ROADMAP.md b/ROADMAP.md
index 608ab8c..1bbd153 100644
--- a/ROADMAP.md
+++ b/ROADMAP.md
@@ -96,6 +96,7 @@
- [ ] Corpus / mutations / research harness — [#92](https://github.com/askmy-stack/tool-semantics/issues/92)–[#95](https://github.com/askmy-stack/tool-semantics/issues/95) **P2/P3**
- [ ] Confidence intervals + dev/test/verified split — [#115](https://github.com/askmy-stack/tool-semantics/issues/115), [#116](https://github.com/askmy-stack/tool-semantics/issues/116) **P3**
- [ ] DX: init/doctor/lint/audit/generate-probes — [#96](https://github.com/askmy-stack/tool-semantics/issues/96)–[#98](https://github.com/askmy-stack/tool-semantics/issues/98) **P2**
+- [x] Docs / README capture+eval positioning — [#101](https://github.com/askmy-stack/tool-semantics/issues/101) **P2**
- [x] ROADMAP/PLAN + agent execution spec — [#103](https://github.com/askmy-stack/tool-semantics/issues/103)
- [ ] Apply GitHub priority labels — [#118](https://github.com/askmy-stack/tool-semantics/issues/118) **P0**
- Operational spec: [docs/AGENT_EXECUTION.md](docs/AGENT_EXECUTION.md)
diff --git a/docs/PLAN.md b/docs/PLAN.md
index a6dfe72..aa9bf31 100644
--- a/docs/PLAN.md
+++ b/docs/PLAN.md
@@ -1,7 +1,9 @@
# Course of action (Milestones 7+)
Primary operational spec: [AGENT_EXECUTION.md](AGENT_EXECUTION.md).
-Product positioning: behavioral regression for MCP / AI-agent interfaces.
+Product positioning: **schema-valid ≠ agent-safe** — behavioral regression for
+MCP / AI-agent interfaces ([simple-explanation.md](simple-explanation.md),
+[concepts.md](concepts.md), [index.md](index.md)).
Shipped through **v0.4.0**: Milestones 0–6 (structural detect, stdio/SSE capture,
offline + library model probes, policy, Action, adapters).
@@ -13,6 +15,9 @@ offline + library model probes, policy, Action, adapters).
3. **Evaluate real behavior** — collision/rename, traces, discovery, safety/output (#79–#91)
4. **Prove** — corpus, mutations, research harness, DX/docs (#92–#103)
+Beginner UX today: **capture → compare + probe**. Unified `eval` (#76) becomes
+the single evaluate entrypoint when merged.
+
## Immediate next PRs
| Order | Issues | Priority | Outcome |
diff --git a/docs/architecture.md b/docs/architecture.md
index acceb6c..9069f6c 100644
--- a/docs/architecture.md
+++ b/docs/architecture.md
@@ -2,6 +2,8 @@
Tool-Semantics separates **transport**, **normalization**, and **compatibility analysis** so the same diff engine can run on static manifests and live MCP servers.
+Doc index / vocabulary: [index.md](index.md), [concepts.md](concepts.md).
+
## Pipeline
```mermaid
diff --git a/docs/concepts.md b/docs/concepts.md
new file mode 100644
index 0000000..cd99ba0
--- /dev/null
+++ b/docs/concepts.md
@@ -0,0 +1,77 @@
+# Concepts and shared vocabulary
+
+Short definitions used across Tool-Semantics docs and reports.
+
+## Snapshots
+
+A normalized **InterfaceSnapshot** of an MCP / tool catalog: tool names,
+descriptions, parameters, optional output schemas, optional `risk`, plus
+protocol/server metadata. Snapshots are the **canonical compare input** —
+not live traffic.
+
+- Capture from a JSON manifest: `capture`
+- Capture from a live server: `capture-mcp` (stdio / Streamable HTTP / SSE)
+
+See [architecture.md](architecture.md), [mcp-versions.md](mcp-versions.md).
+
+## Diff / change codes
+
+Deterministic comparison of baseline vs candidate snapshots. Each finding has a
+stable **code** (e.g. `tool.removed`), **severity** (`info` → `critical`),
+subject, and message.
+
+Catalog: [change-codes.md](change-codes.md).
+Semantics note: description / rename warnings matter because models route on
+prose, not only types.
+
+## Semantics (agent behavior)
+
+Beyond schema validity:
+
+| Idea | Meaning |
+| --- | --- |
+| Tool selection | Will the model still pick the intended tool? |
+| Argument patterns | Will prior argument shapes still validate? |
+| Risk / confirmation | Did side-effect expectations change? |
+
+## Probes
+
+Behavioral checks against a snapshot.
+
+- **Offline** — no LLM; assert expected tools / params / risk exist.
+- **Model-backed** — opt-in runner selects a tool for an intent (`approved`
+ probes by default).
+
+See [probes.md](probes.md).
+
+## Traces
+
+Recorded agent trajectories (tool calls over time) for replay and regression.
+Trace schema / replay land in Milestone work (#83–#84); not required for the
+beginner capture → evaluate path.
+
+## Safety
+
+Optional tool-level `risk` (`read_only`, `external_write`, `destructive`,
+`unknown`) and confirmation expectations. Missing risk defaults to `unknown`
+(no false escalation). Broader safety/scope diffs are roadmap items (#87).
+
+## Policies
+
+Release knobs that map severities to CI failure
+(`compatible` / `breaking` / `critical-only` / …). Config:
+[config.md](config.md).
+
+## Benchmarks / research
+
+Corpus cases, splits (DEV/TEST/VERIFIED), mutations, and sensitivity harnesses
+prove detectors without contaminating published scores. See ROADMAP Milestone
+12+ and [AGENT_EXECUTION.md](AGENT_EXECUTION.md).
+
+## DETECT · TEST · PROTECT
+
+Product framing used in the README and [simple-explanation.md](simple-explanation.md):
+
+1. **DETECT** — snapshots + structural diff
+2. **TEST** — probes (and later traces / eval)
+3. **PROTECT** — policies, Action, exit codes
diff --git a/docs/index.md b/docs/index.md
new file mode 100644
index 0000000..02a2159
--- /dev/null
+++ b/docs/index.md
@@ -0,0 +1,49 @@
+# Documentation index
+
+Start here if you are new: [simple-explanation.md](simple-explanation.md) →
+[concepts.md](concepts.md) → beginner path on the [README](../README.md).
+
+## By topic
+
+| Topic | Doc |
+| --- | --- |
+| Plain-language overview | [simple-explanation.md](simple-explanation.md) |
+| Shared vocabulary | [concepts.md](concepts.md) |
+| Architecture / pipeline | [architecture.md](architecture.md) |
+| Snapshots & live MCP capture | [mcp-versions.md](mcp-versions.md), [adr-live-mcp-capture.md](adr-live-mcp-capture.md) |
+| Structural diff & change codes | [change-codes.md](change-codes.md) |
+| Semantics & agent risk | [concepts.md](concepts.md)#semantics-agent-behavior |
+| Probes (offline / model) | [probes.md](probes.md) |
+| Traces (roadmap) | [AGENT_EXECUTION.md](AGENT_EXECUTION.md), ROADMAP #83–#84 |
+| Safety / risk field | [concepts.md](concepts.md)#safety, [probes.md](probes.md) |
+| Policies & ignore config | [config.md](config.md) |
+| Benchmarks / research | [AGENT_EXECUTION.md](AGENT_EXECUTION.md), [PLAN.md](PLAN.md) |
+| GitHub Action (protect) | [github-action.md](github-action.md) |
+| Migration adapters | [adapters.md](adapters.md) |
+| Downstream consumers | [downstream.md](downstream.md) |
+| Publishing | [publishing.md](publishing.md) |
+
+## Beginner vs advanced CLI
+
+**Beginner (capture → evaluate → protect)**
+
+1. `capture` / `capture-mcp`
+2. `compare` + `probe`
+3. CI Action / non-zero exit
+
+**Advanced / secondary**
+
+- `--config` ignore rules, provenance sidecars, verbose stderr logs
+- `--model` / `--trials` model-backed probes
+- Library APIs for custom runners, adapters, policies
+
+Unified `eval` ([#76](https://github.com/askmy-stack/tool-semantics/issues/76))
+will become the single evaluate entrypoint; `compare` and `probe` remain for
+advanced workflows.
+
+## Milestone 7+ plan
+
+- [PLAN.md](PLAN.md) — course of action
+- [AGENT_EXECUTION.md](AGENT_EXECUTION.md) — binding priority tables
+- [../ROADMAP.md](../ROADMAP.md) — checklist
+- [../CLAUDE.md](../CLAUDE.md) — short agent guidance
diff --git a/docs/simple-explanation.md b/docs/simple-explanation.md
new file mode 100644
index 0000000..5a5182f
--- /dev/null
+++ b/docs/simple-explanation.md
@@ -0,0 +1,54 @@
+# Tool-Semantics in plain language
+
+**Schema-valid is not the same as agent-safe.**
+
+Typed clients fail loudly when a field disappears. Language-model agents often
+**guess**: they pick tools from descriptions, invent arguments from schemas, and
+infer side effects from names. A change that still validates as JSON Schema can
+still break the agent.
+
+Tool-Semantics answers one question:
+
+> If we ship this MCP / tool-interface change, will agents still behave?
+
+## Three jobs: DETECT · TEST · PROTECT
+
+| Job | What you do | Typical command |
+| --- | --- | --- |
+| **DETECT** | Snapshot the interface; see what changed | `capture`, `capture-mcp`, `compare` |
+| **TEST** | Check whether agents still pick tools / args correctly | `probe` (offline or `--model`) |
+| **PROTECT** | Gate merges in CI with policies and reports | GitHub Action, exit codes, config |
+
+## Beginner path
+
+1. **Capture** a baseline and a candidate snapshot.
+2. **Evaluate** with structural compare plus behavioral probes.
+3. **Protect** the merge when the report fails policy.
+
+```bash
+# 1. Capture
+tool-semantics capture examples/github_server_v1.json -o .tool-semantics/v1.json
+tool-semantics capture examples/github_server_v2.json -o .tool-semantics/v2.json
+
+# 2. Evaluate (structural + behavioral)
+tool-semantics compare .tool-semantics/v1.json .tool-semantics/v2.json \
+ --markdown-output .tool-semantics/report.md
+tool-semantics probe .tool-semantics/v2.json \
+ --probes examples/probes/github_v1_offline.json
+
+# 3. Protect — non-zero exit fails CI (see docs/github-action.md)
+```
+
+A unified `eval` command ([#76](https://github.com/askmy-stack/tool-semantics/issues/76))
+will combine compare + probes into one beginner entrypoint; until then, use the
+two commands above.
+
+## What “breaking” means here
+
+- **Breaking / critical** structural codes fail `compare` by default
+ ([change-codes.md](change-codes.md)).
+- **Warnings** (e.g. description drift) flag selection risk without always
+ failing CI — agents may re-route even when schemas remain valid.
+- **Probes** catch behavioral regressions that pure schema diffs miss.
+
+Next: [concepts.md](concepts.md) · [index.md](index.md) · [README](../README.md)
diff --git a/tests/test_docs_positioning.py b/tests/test_docs_positioning.py
new file mode 100644
index 0000000..0e4dbb7
--- /dev/null
+++ b/tests/test_docs_positioning.py
@@ -0,0 +1,37 @@
+"""Smoke tests for docs positioning (#101)."""
+
+from __future__ import annotations
+
+from pathlib import Path
+
+ROOT = Path(__file__).resolve().parents[1]
+
+
+def test_docs_index_and_beginner_pages_exist() -> None:
+ for relative in (
+ "docs/index.md",
+ "docs/simple-explanation.md",
+ "docs/concepts.md",
+ ):
+ path = ROOT / relative
+ assert path.is_file(), relative
+ text = path.read_text(encoding="utf-8")
+ assert len(text) > 100
+
+
+def test_readme_positions_capture_evaluate_and_schema_valid() -> None:
+ readme = (ROOT / "README.md").read_text(encoding="utf-8")
+ assert "Schema-valid ≠ agent-safe" in readme or "schema-valid ≠ agent-safe" in readme.lower()
+ assert "DETECT" in readme and "TEST" in readme and "PROTECT" in readme
+ assert "capture → evaluate" in readme.lower() or "capture → evaluate" in readme
+ assert "docs/simple-explanation.md" in readme
+ assert "docs/index.md" in readme
+
+
+def test_claude_and_plan_point_at_milestone7_docs() -> None:
+ claude = (ROOT / "CLAUDE.md").read_text(encoding="utf-8")
+ plan = (ROOT / "docs/PLAN.md").read_text(encoding="utf-8")
+ assert "docs/PLAN.md" in claude
+ assert "AGENT_EXECUTION.md" in claude
+ assert "simple-explanation.md" in plan
+ assert "capture → compare" in plan or "capture → compare + probe" in plan