Your agent says done. Make it prove it.
testcase is a set of agent skills that take a requirement all the way to code that ships: pin down what the spec asks for, turn it into real tests, then build until every check passes. Every step leaves proof you can rerun.
You ask an AI agent to build a feature. It says done. You check — a requirement is missing, a test was never written, the edge case from the third comment on the ticket got ignored. You ask again. It finds more. Repeat.
testcase ends that loop.
Every requirement counted. Each one gets an id, a yes/no question and the line it came from. A requirement can't quietly drop out.
Tests before trust. Requirements become test cases, and the test cases become real tests in your repo's own framework — not a checklist in a chat window.
Done means exit 0. goalrun builds phase by phase and won't say done until the script passes on the whole ledger. Not "done pending X". Not "I checked by hand".
The core pipeline, one handoff per stage. Use any stage alone, or chain them and hand off the whole job:
| Step | Skill | You get |
|---|---|---|
| 1. Require | docs-review |
Atomic REQ- ids, each a yes/no, each citing its source line |
| 2. Prove | testcase |
TC- ids traced to a REQ-, implemented as tests in the repo's own framework |
| 3. Build | goalrun |
The implementation, phase by phase, until the ledger of checks exits 0 |
Four diagrams of the pipeline, and of how goalrun decides a row and proves a check: How these skills work.
- Install the plugin (Claude Code shown; other hosts below):
claude plugin marketplace add xtieume/testcase claude plugin install testcase@testcase-marketplace
- Point it at a spec: "write test cases for spec.md".
- Hand off the build: "/goalrun build the export feature and don't stop until it's done".
Tip
On Claude Code, start with /goal. /goal sets a condition that is checked after every turn, and Claude keeps working until it holds — so the run doesn't stop halfway to ask whether to continue:
/goal /goalrun build the export feature — done when goalrun.py exits 0 on the whole ledger
Tip
Already think it's finished? Ask "is this actually done?" — goalrun audits the work against its ledger and tells you what's still red.
Claude Code — marketplace install, gets every skill in the repo:
claude plugin marketplace add xtieume/testcase
claude plugin install testcase@testcase-marketplaceZCode — same flow, reading .zcode-plugin/:
zcode plugin marketplace add xtieume/testcase
zcode plugin install testcase@testcase-marketplaceCursor / Antigravity — both read .agents/skills/ natively. Clone once, then symlink it into a project or copy it globally:
git clone https://github.com/xtieume/testcase.git
# project level (either editor)
ln -s "$(pwd)/testcase/.agents/skills" .agents/skills
# global
cp -R testcase/.agents/skills/* ~/.cursor/skills/ # Cursor
cp -R testcase/.agents/skills/* ~/.gemini/antigravity/skills/ # AntigravityAny host, one skill only — copy the folder you want:
cp -R .agents/skills/<name> ~/.claude/skills/<name>Skills trigger on natural language, or explicitly as /<name>.
Each name links to its SKILL.md, which is the reference for that skill — triggers, workflow, flags, scripts.
| Skill | Does | Trigger |
|---|---|---|
🎯 goalrun |
Turns a goal into a ledger of checks, builds each phase through an independent subagent, and will not say done until the script exits 0. Also audits work someone else called finished | "build X and don't stop until it's done", "is this actually done?" |
✍️ normalize |
Rewrites a rambling everyday prompt into a compact engineering prompt — same requirements, imperative form, explicit escalation and done criteria | "rewrite this prompt for /goalrun" |
| Skill | Does | Trigger |
|---|---|---|
🧪 testcase |
Test cases from a requirement or a Figma design, attacks its own output for missed cases, then implements the automatable ones as runnable tests and files what fails | "write test cases for…" |
📋 docs-review |
Audits docs against a spec: required vs actually written, with a citation per verdict | "review the docs against spec.md" |
| Skill | Does | Trigger |
|---|---|---|
📥 playwright-cdp |
Notion pages and Slack threads → markdown through a logged-in browser: body, every comment, and the attachments downloaded | "read this Notion page", "pull this Slack thread" |
Conventions every skill here follows, so a new one is predictable before you open it:
SKILL.mdstays small. Detail goes inreferences/, loaded only when the task needs it.- Review before returning. Skills that produce a deliverable end with an independent subagent pass that re-derives the work and attacks it, repeating until a round adds nothing.
- Scripts are stdlib. Python or Node, no install step unless the skill says otherwise; they lint the output rather than trusting it.
- Verdicts carry evidence. Any claim about a document or requirement cites the line it came from.
Setup beyond cloning, where a skill needs it:
cd .agents/skills/playwright-cdp/scripts && npm install # onceDrop a folder into .agents/skills/<name>/ with a SKILL.md (frontmatter name + description), plus optional references/ and scripts/. No manifest edit: both plugin.json files point at the directory, not a list. Add a row to the table above. Versions are not edited by hand: merging to main bumps all four manifests, tags the commit and publishes a release — feat: gives a minor bump, ! or BREAKING CHANGE a major one, anything else a patch.
.agents/skills/<name>/ SKILL.md + references/ + scripts/
.claude-plugin/ Claude Code manifests
.zcode-plugin/ ZCode manifests
MIT