Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
10 changes: 9 additions & 1 deletion CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -9,6 +9,12 @@ the full roadmap and agent breakdown lives in [docs/PLAN.md](docs/PLAN.md) and
**Tool-Semantics — Behavioral regression testing for MCP and AI-agent interfaces.**

Know when an MCP change breaks the agent, not just the schema.
**Schema-valid ≠ agent-safe.**

Product framing: **DETECT** (capture/compare) → **TEST** (probes) → **PROTECT** (CI).
Beginner path: capture → compare + probe (unified `eval` is #76).
Docs: [docs/simple-explanation.md](docs/simple-explanation.md),
[docs/index.md](docs/index.md).

It snapshots a tool interface, diffs baseline vs. candidate, and reports risk
across five compatibility layers:
Expand Down Expand Up @@ -38,7 +44,8 @@ Streamable HTTP, and legacy SSE.
- `src/tool_semantics/cli.py` — Typer CLI entrypoint
- `tests/` — pytest suite, one file per module area
- `examples/` — demo MCP-style manifests (GitHub server v1/v2, weather)
- `docs/` — architecture, config, change-codes, github-action, publishing, adapters
- `docs/` — [index](docs/index.md), simple-explanation, concepts, architecture,
config, change-codes, probes, github-action, publishing, adapters

## Dev commands

Expand All @@ -53,6 +60,7 @@ ruff format --check .

tool-semantics capture examples/github_server_v1.json -o .tool-semantics/v1.json
tool-semantics compare .tool-semantics/v1.json .tool-semantics/v2.json --markdown-output report.md
tool-semantics probe .tool-semantics/v2.json --probes examples/probes/github_v1_offline.json
tool-semantics capture-mcp -o snap.json https://example.com/mcp
```

Expand Down
83 changes: 63 additions & 20 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -9,7 +9,8 @@
</p>

<p align="center">
Behavioral compatibility testing for MCP tools and AI-agent interfaces.
Behavioral compatibility testing for MCP tools and AI-agent interfaces.<br />
<em>Schema-valid ≠ agent-safe.</em>
</p>

<p align="center">
Expand All @@ -23,14 +24,29 @@

## Why Tool-Semantics?

AI agents do not call tools the way typed clients do. They choose tools from **descriptions**, invent **arguments** from schemas, and infer **side effects** from naming and prose. A change that remains JSON-Schema-valid can still:
AI agents do not call tools the way typed clients do. They choose tools from
**descriptions**, invent **arguments** from schemas, and infer **side effects**
from naming and prose. A change that remains JSON-Schema-valid can still:

- steer the model toward the wrong tool
- drop a required argument the model used to omit
- rename enums the model still emits
- quietly escalate from read-only to write/destructive behavior

**Tool-Semantics** captures normalized tool-interface snapshots and diffs them for structural *and* semantic risk — so teams can gate MCP and tool-API changes before agents ship broken workflows.
**Schema-valid ≠ agent-safe.** Tool-Semantics captures normalized tool-interface
snapshots and evaluates them for structural *and* semantic risk — so teams can
gate MCP and tool-API changes before agents ship broken workflows.

New here? Read the [simple explanation](docs/simple-explanation.md) and
[concepts](docs/concepts.md). Full doc index: [docs/index.md](docs/index.md).

## DETECT · TEST · PROTECT

| | Job | Commands |
| --- | --- | --- |
| **DETECT** | Snapshot interfaces; see what changed | `capture`, `capture-mcp`, `compare` |
| **TEST** | Check tool selection / args / risk expectations | `probe` (offline or `--model`) |
| **PROTECT** | Fail CI when policy says so | exit codes, Action, [config](docs/config.md) |

## Compatibility layers

Expand All @@ -42,7 +58,9 @@ AI agents do not call tools the way typed clients do. They choose tools from **d
| 4. Execution | Do calls still succeed with prior argument patterns? |
| 5. Intent / side effects | Did risk, confirmation needs, or outcomes change? |

The MVP implements deterministic interface snapshots and structural comparison (layers 1–2, with warnings that point at 3–5), plus live MCP capture over stdio, Streamable HTTP, and legacy SSE. Model-based behavioral testing is available as an opt-in library; see the [roadmap](ROADMAP.md).
Layers 1–2 are CI-gateable today. Layers 3–5 are covered by offline probes plus
**opt-in** model-backed probes. Live capture supports stdio, Streamable HTTP,
and legacy SSE. See the [roadmap](ROADMAP.md).

## How it works

Expand All @@ -63,22 +81,30 @@ flowchart LR
E --> F[CLI / CI exit codes]
```

## Quick start
## Quick start — capture → evaluate

```bash
python -m venv .venv
source .venv/bin/activate
pip install -e ".[dev]"

# Capture two interface versions
# 1. Capture baseline and candidate
tool-semantics capture examples/github_server_v1.json -o .tool-semantics/v1.json
tool-semantics capture examples/github_server_v2.json -o .tool-semantics/v2.json

# Compare — exits 1 on breaking/critical changes
# 2. Evaluate — structural diff + behavioral probes
tool-semantics compare .tool-semantics/v1.json .tool-semantics/v2.json \
--markdown-output .tool-semantics/report.md
tool-semantics probe .tool-semantics/v2.json \
--probes examples/probes/github_v1_offline.json

# 3. Protect — non-zero exit fails CI (see docs/github-action.md)
```

A unified `tool-semantics eval` command will combine compare + probes into one
beginner entrypoint ([#76](https://github.com/askmy-stack/tool-semantics/issues/76)).
Until then, use **capture → compare + probe** as the evaluate path.

### Demo

<p align="center">
Expand Down Expand Up @@ -132,6 +158,11 @@ print("compatible:", report.is_compatible)

## CLI reference

**Beginner path:** `capture` → `compare` + `probe` (evaluate) → CI protect.

**Advanced / secondary** (linked below): provenance, `--config`, `--model`,
`--trials`, verbose logs, library APIs.

```bash
tool-semantics --version
tool-semantics capture <manifest.json> [-o .tool-semantics/snapshot.json] \
Expand Down Expand Up @@ -172,12 +203,23 @@ The optional provenance sidecar records capture context and a digest of the
snapshot. It is separate from the snapshot and never affects compatibility
comparisons.

JSON reports include `changes`, `is_compatible`, and `counts` by severity.
Change-code catalog: [docs/change-codes.md](docs/change-codes.md).
Ignore-config schema: [docs/config.md](docs/config.md).
GitHub Action: [docs/github-action.md](docs/github-action.md).
Publishing: [docs/publishing.md](docs/publishing.md).
Migration adapters: [docs/adapters.md](docs/adapters.md).
JSON reports include `changes`, `is_compatible`, and `counts` by severity.

### Docs

| | |
| --- | --- |
| Index | [docs/index.md](docs/index.md) |
| Simple explanation | [docs/simple-explanation.md](docs/simple-explanation.md) |
| Concepts | [docs/concepts.md](docs/concepts.md) |
| Change codes | [docs/change-codes.md](docs/change-codes.md) |
| Architecture | [docs/architecture.md](docs/architecture.md) |
| Config / policy | [docs/config.md](docs/config.md) |
| Probes | [docs/probes.md](docs/probes.md) |
| GitHub Action | [docs/github-action.md](docs/github-action.md) |
| Publishing | [docs/publishing.md](docs/publishing.md) |
| Adapters | [docs/adapters.md](docs/adapters.md) |
| Milestone 7+ plan | [docs/PLAN.md](docs/PLAN.md), [docs/AGENT_EXECUTION.md](docs/AGENT_EXECUTION.md) |

### Optional `risk` field

Expand Down Expand Up @@ -212,19 +254,20 @@ assert report.passed
## Project layout

```text
src/tool_semantics/ # scanner, models, diff engine, report, CLI
examples/ # demo MCP-style manifests
src/tool_semantics/ # scanner, models, diff engine, probes, report, CLI
examples/ # demo MCP-style manifests + probe fixtures
tests/ # pytest suite
docs/ # index, concepts, architecture, change-codes, …
docs/assets/ # README visuals
```

## Roadmap

See [ROADMAP.md](ROADMAP.md) for milestones. Live MCP supports stdio, Streamable
HTTP, and legacy SSE ([docs/mcp-versions.md](docs/mcp-versions.md)).
Model-backed probes / metrics / stability are available via the library
API (`docs/probes.md`). Downstream consumers (myelinmesh, dogfood capture):
[`docs/downstream.md`](docs/downstream.md).
See [ROADMAP.md](ROADMAP.md) and [docs/PLAN.md](docs/PLAN.md) for Milestone 7+.
Live MCP supports stdio, Streamable HTTP, and legacy SSE
([docs/mcp-versions.md](docs/mcp-versions.md)). Model-backed probes:
[docs/probes.md](docs/probes.md). Downstream consumers:
[docs/downstream.md](docs/downstream.md).

## Contributing

Expand Down
1 change: 1 addition & 0 deletions ROADMAP.md
Original file line number Diff line number Diff line change
Expand Up @@ -96,6 +96,7 @@
- [ ] Corpus / mutations / research harness — [#92](https://github.com/askmy-stack/tool-semantics/issues/92)–[#95](https://github.com/askmy-stack/tool-semantics/issues/95) **P2/P3**
- [ ] Confidence intervals + dev/test/verified split — [#115](https://github.com/askmy-stack/tool-semantics/issues/115), [#116](https://github.com/askmy-stack/tool-semantics/issues/116) **P3**
- [ ] DX: init/doctor/lint/audit/generate-probes — [#96](https://github.com/askmy-stack/tool-semantics/issues/96)–[#98](https://github.com/askmy-stack/tool-semantics/issues/98) **P2**
- [x] Docs / README capture+eval positioning — [#101](https://github.com/askmy-stack/tool-semantics/issues/101) **P2**
- [x] ROADMAP/PLAN + agent execution spec — [#103](https://github.com/askmy-stack/tool-semantics/issues/103)
- [ ] Apply GitHub priority labels — [#118](https://github.com/askmy-stack/tool-semantics/issues/118) **P0**
- Operational spec: [docs/AGENT_EXECUTION.md](docs/AGENT_EXECUTION.md)
7 changes: 6 additions & 1 deletion docs/PLAN.md
Original file line number Diff line number Diff line change
@@ -1,7 +1,9 @@
# Course of action (Milestones 7+)

Primary operational spec: [AGENT_EXECUTION.md](AGENT_EXECUTION.md).
Product positioning: behavioral regression for MCP / AI-agent interfaces.
Product positioning: **schema-valid ≠ agent-safe** — behavioral regression for
MCP / AI-agent interfaces ([simple-explanation.md](simple-explanation.md),
[concepts.md](concepts.md), [index.md](index.md)).

Shipped through **v0.4.0**: Milestones 0–6 (structural detect, stdio/SSE capture,
offline + library model probes, policy, Action, adapters).
Expand All @@ -13,6 +15,9 @@ offline + library model probes, policy, Action, adapters).
3. **Evaluate real behavior** — collision/rename, traces, discovery, safety/output (#79–#91)
4. **Prove** — corpus, mutations, research harness, DX/docs (#92–#103)

Beginner UX today: **capture → compare + probe**. Unified `eval` (#76) becomes
the single evaluate entrypoint when merged.

## Immediate next PRs

| Order | Issues | Priority | Outcome |
Expand Down
2 changes: 2 additions & 0 deletions docs/architecture.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,8 @@

Tool-Semantics separates **transport**, **normalization**, and **compatibility analysis** so the same diff engine can run on static manifests and live MCP servers.

Doc index / vocabulary: [index.md](index.md), [concepts.md](concepts.md).

## Pipeline

```mermaid
Expand Down
77 changes: 77 additions & 0 deletions docs/concepts.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,77 @@
# Concepts and shared vocabulary

Short definitions used across Tool-Semantics docs and reports.

## Snapshots

A normalized **InterfaceSnapshot** of an MCP / tool catalog: tool names,
descriptions, parameters, optional output schemas, optional `risk`, plus
protocol/server metadata. Snapshots are the **canonical compare input** —
not live traffic.

- Capture from a JSON manifest: `capture`
- Capture from a live server: `capture-mcp` (stdio / Streamable HTTP / SSE)

See [architecture.md](architecture.md), [mcp-versions.md](mcp-versions.md).

## Diff / change codes

Deterministic comparison of baseline vs candidate snapshots. Each finding has a
stable **code** (e.g. `tool.removed`), **severity** (`info` → `critical`),
subject, and message.

Catalog: [change-codes.md](change-codes.md).
Semantics note: description / rename warnings matter because models route on
prose, not only types.

## Semantics (agent behavior)

Beyond schema validity:

| Idea | Meaning |
| --- | --- |
| Tool selection | Will the model still pick the intended tool? |
| Argument patterns | Will prior argument shapes still validate? |
| Risk / confirmation | Did side-effect expectations change? |

## Probes

Behavioral checks against a snapshot.

- **Offline** — no LLM; assert expected tools / params / risk exist.
- **Model-backed** — opt-in runner selects a tool for an intent (`approved`
probes by default).

See [probes.md](probes.md).

## Traces

Recorded agent trajectories (tool calls over time) for replay and regression.
Trace schema / replay land in Milestone work (#83–#84); not required for the
beginner capture → evaluate path.

## Safety

Optional tool-level `risk` (`read_only`, `external_write`, `destructive`,
`unknown`) and confirmation expectations. Missing risk defaults to `unknown`
(no false escalation). Broader safety/scope diffs are roadmap items (#87).

## Policies

Release knobs that map severities to CI failure
(`compatible` / `breaking` / `critical-only` / …). Config:
[config.md](config.md).

## Benchmarks / research

Corpus cases, splits (DEV/TEST/VERIFIED), mutations, and sensitivity harnesses
prove detectors without contaminating published scores. See ROADMAP Milestone
12+ and [AGENT_EXECUTION.md](AGENT_EXECUTION.md).

## DETECT · TEST · PROTECT

Product framing used in the README and [simple-explanation.md](simple-explanation.md):

1. **DETECT** — snapshots + structural diff
2. **TEST** — probes (and later traces / eval)
3. **PROTECT** — policies, Action, exit codes
49 changes: 49 additions & 0 deletions docs/index.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,49 @@
# Documentation index

Start here if you are new: [simple-explanation.md](simple-explanation.md) →
[concepts.md](concepts.md) → beginner path on the [README](../README.md).

## By topic

| Topic | Doc |
| --- | --- |
| Plain-language overview | [simple-explanation.md](simple-explanation.md) |
| Shared vocabulary | [concepts.md](concepts.md) |
| Architecture / pipeline | [architecture.md](architecture.md) |
| Snapshots & live MCP capture | [mcp-versions.md](mcp-versions.md), [adr-live-mcp-capture.md](adr-live-mcp-capture.md) |
| Structural diff & change codes | [change-codes.md](change-codes.md) |
| Semantics & agent risk | [concepts.md](concepts.md)#semantics-agent-behavior |
| Probes (offline / model) | [probes.md](probes.md) |
| Traces (roadmap) | [AGENT_EXECUTION.md](AGENT_EXECUTION.md), ROADMAP #83–#84 |
| Safety / risk field | [concepts.md](concepts.md)#safety, [probes.md](probes.md) |
| Policies & ignore config | [config.md](config.md) |
| Benchmarks / research | [AGENT_EXECUTION.md](AGENT_EXECUTION.md), [PLAN.md](PLAN.md) |
| GitHub Action (protect) | [github-action.md](github-action.md) |
| Migration adapters | [adapters.md](adapters.md) |
| Downstream consumers | [downstream.md](downstream.md) |
| Publishing | [publishing.md](publishing.md) |

## Beginner vs advanced CLI

**Beginner (capture → evaluate → protect)**

1. `capture` / `capture-mcp`
2. `compare` + `probe`
3. CI Action / non-zero exit

**Advanced / secondary**

- `--config` ignore rules, provenance sidecars, verbose stderr logs
- `--model` / `--trials` model-backed probes
- Library APIs for custom runners, adapters, policies

Unified `eval` ([#76](https://github.com/askmy-stack/tool-semantics/issues/76))
will become the single evaluate entrypoint; `compare` and `probe` remain for
advanced workflows.

## Milestone 7+ plan

- [PLAN.md](PLAN.md) — course of action
- [AGENT_EXECUTION.md](AGENT_EXECUTION.md) — binding priority tables
- [../ROADMAP.md](../ROADMAP.md) — checklist
- [../CLAUDE.md](../CLAUDE.md) — short agent guidance
Loading