A pre-work gate for Claude. Before it starts building, it lists the assumptions it was about to make on your behalf, asks about the ones the deliverable rests on, and waits for an answer.
The problem it solves isn't ambiguity — it's silent inference. An expert's prompt is compressed. What you didn't say isn't absence of intent, it's unstated intent, and a model will fill it from its priors without ever mentioning that it did.
- Scans your request against the spec a complete version of that request would have contained — purpose, audience, format, constraints, done-ness — plus the forks it would hit while doing the work.
- Ledgers every gap it was about to fill itself, with what it would have filled it with and where that guess came from.
- Bands each one: settled (the words admit only one reading, or a standing answer covers it) or open (anything else that shapes the deliverable).
- Asks every open one, through the question tool rather than buried in prose. Likely answers are offered as one-click options; facts only you hold are asked open-ended, never invented or placeholdered.
- Shows its work: every open item — including ones queued for the next round — is named in the reply, along with anything it found while reading that a question depends on.
- Repeats until nothing load-bearing is left to guess. No round cap; you end it early by saying "just go".
- Gates on an explicit yes, with the full ledger shown so you can audit every inference it was about to make. Silence is not consent.
The goal is residual inference near zero. A "sensible default" is still an inference, so there is no defaults band: what gets minimised is the effort per question, not the number of questions.
1.x declared "defaulted" items inline and capped questioning at three per round. In testing, that band was where silent inference hid: defaults stayed in private notes, and interpretations were presented as the user's own rules. 2.0 removes the band and the cap, names every queued item, requires quoting the user in full, and explicitly checks for reference points (what a check is measured against) and for questions the user's own constraints already settle.
Claude Code, via marketplace:
/plugin marketplace add bryanthood-wph/clarify
/plugin install clarify@bryanthood-wph
Claude Code, manual:
git clone https://github.com/bryanthood-wph/clarify.git
cp -r clarify/skills/clarify ~/.claude/skills/claude.ai — download the release zip, then Customize → Skills → Add. Requires code execution to be enabled. See the note below first.
It fires when you ask for it in any of the usual ways — check your
assumptions, gate check this, don't assume, ask me, make sure you
understand before you start, interview me. You can also call it directly as
/clarify.
It also fires unprompted when a short prompt leaves open a decision that would be expensive to undo if guessed wrong: architecture, data or state design, money or pricing, anything published or customer-facing at launch, or multi-hour work.
It's built not to fire on quick, cheaply-redone drafts, when you say just go or no questions, for red-teaming an existing plan, or when "clarify" points backward at something Claude already produced (an explanation request, not a gate).
claude.ai enforces a hard 200-character limit on skill descriptions. The description this skill ships with is longer than that, because the long version buys precision by naming its exclusions explicitly — which is what keeps it from firing on near-miss prompts.
So the shipped description works in Claude Code and must be shortened for claude.ai. Shortening it is a real tradeoff and not one to guess at: the description is the entire triggering mechanism, and cutting the exclusions is what most likely breaks it.
evals/ exists for exactly this. Test a candidate before you trust it.
evals/ has a labelled set of 20 prompts and a script that measures how a
candidate description performs against them.
python evals/run_trigger_eval.py --skill skills/clarifyThere's a real gotcha documented in evals/README.md: if clarify is already
installed on the account you run as, Claude invokes the installed one instead of
your candidate and every candidate scores zero regardless of what it says. Run
the evals somewhere the skill isn't installed.
These were measured by the author on one machine with Claude Opus 5.5, with the skill installed under its real name (so they test the installed description, not a candidate), using simulated users. Treat them as indicative, not as a benchmark.
- Triggering: 20/20 on each of two labelled sets (40 prompts, 3 runs each). One set came from the author's own project; the other from unrelated fields (business, databases, writing, personal tasks), with near-misses such as "clarify what you meant", "just draft it, no questions", and red-team requests.
- Behaviour: 6 scenarios × 3 runs each, comparing 2.0, 1.x, and no skill. Scenarios covered a compressed business request, a near-complete coding request, a personal plan with a hidden conflict, and three article-writing tasks (from scratch, mid-conversation, and conflicting with standing rules). Graded on residual inference: load-bearing decisions the reply made itself rather than taking from the user or asking.
| Version | Residual inferences per reply (mean) |
|---|---|
| 2.0 | 0.4 |
| 1.x | 4.3 |
| No skill | 2.3 |
No skill matched 2.0 on straightforward article prompts and on obvious conflicts with standing rules. The gap was largest where memory facts or earlier answers invited assumptions. 1.x scored below no skill in most scenarios, which is why its defaults band was removed.
The bands, the stop rule, the question quality bar and the exclusions are all tuned to how I work. They probably shouldn't be tuned to how you work. Fork it, rewrite the questions to fit the gaps you already suspect you have, and if the evals say your version is better, I'd like to see it.
MIT. See LICENSE.