Skip to content

Tell the week the watermark arrived, including the part that went wrong - #63

Merged
peopleworks merged 2 commits into
mainfrom
watermark-article
Aug 20, 2026
Merged

Tell the week the watermark arrived, including the part that went wrong#63
peopleworks merged 2 commits into
mainfrom
watermark-article

Conversation

@peopleworks

Copy link
Copy Markdown
Owner

The eighth article, bilingual, plus its cover and the publication copy.

The angle

Not "the watermark doesn't threaten us". The article says I set out to prove that, could not, and found a worse fault in my own tool on the way. Three results, in this order:

  1. Rewriting to strip a watermark does not change whether we flag a passage. Five crossed, two crossed back, McNemar exact p = 0.453. The hypothesis is withdrawn in bold.
  2. Length does. Same 32 documents: 0/32 whole documents flagged, 13/89 four-hundred-word windows flagged, and eleven of thirty documents flagged at one position in the text and not at another. That is The verdict boundary has no length condition, and short passages cross it #59.
  3. Three independent reviewers found three false sentences of mine, one of them describing my own code wrongly.

The order matters and should survive editing. A reader who leaves with "this tool catches AI-paraphrased work" has taken away the one thing the study could not show, so the negation leads — in the headline, in the article, and in the social copy in PUBLICACION.md.

Verified before generating the HTML

Both languages pass their own linter: ES 7/100 · burstiness 0.66 · ~1,850 words, EN 7/100 · burstiness 0.66 · ~1,780 words. The English draft tripped our own em-dash rule (19 in 1,775 words) and was rewritten to clear it, the same gesture the calibration article made.

Cover

social/watermark-story-cover.png, hand-authored SVG. Wordless apart from two scores, so one image serves both editions. It argues rather than decorates: one page of writing, three windows over it, three different verdicts — and the flagged stretch is the one whose line lengths run even, which is the mechanism visible without being named.

Also opened alongside this work: #61 (five Spanish rules firing on ordinary formal Spanish) and #62 (the character scanner reading Turkish dotless ı as an artifact).

🤖 Generated with Claude Code

peopleworks and others added 2 commits August 19, 2026 20:03
The eighth article, in both languages. The news is days old — Claude has been
marking its output since the second — so the timing is the best of the eight
and the story is the least comfortable.

It does not argue that the watermark leaves us untouched. It says I set out to
prove that, could not, and found a worse fault in my own tool on the way:
rewriting to strip a mark does not change whether we flag a passage (p =
0.453, hypothesis withdrawn), while cutting the same writing to four hundred
words does — eleven of thirty documents flagged at one position in the text
and not another.

The reviewers are in it by name of finding: three independent readers, three
false sentences of mine, one of them describing my own code wrongly. A project
that has spent seven articles asking other people to publish their errors does
not get to leave that out.

The order of the three results is deliberate and should survive editing. A
reader who leaves with "this tool catches AI-paraphrased work" has taken away
the one thing the study could not show, so the negation leads and the social
copy leads with it too.

Both pass their own linter at 7/100 with burstiness 0.66. The English draft
tripped our own em-dash rule at 19 in 1,775 words and was fixed before the
HTML was generated, which is the same gesture the calibration article made.

The cover is wordless apart from two scores, so one image serves both
editions: one page, three windows, three verdicts, and the flagged stretch is
the one whose lines run even.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015PEbbiYSNPw7jE3LrPNhyF
AI_Tasks/ is where the state of play sits between sessions: what is half
finished, what only Pedro can do, briefs for the other models, and the
committee's verdicts. Adapted from the ERP's folder of the same name, with one
difference that governs the rest — that project has no issue tracker and this
one does, so nothing here duplicates a GitHub issue. The test is whether a
stranger could act on it: if they could, it is an issue.

Committed as an ignore rule rather than left loose on purpose. An ignore that
exists only in somebody's working copy is one `git add -A` away from putting
unpublished drafts and unfixed defects on the public default branch, which has
happened here once already with temp/.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015PEbbiYSNPw7jE3LrPNhyF
@peopleworks

Copy link
Copy Markdown
Owner Author

Added one more commit on this branch: an ignore rule for AI_Tasks/, the working desk this session set up outside the repository (status, Pedro's manual publication queue, task briefs for the other models, committee verdicts).

It rides along here rather than in its own PR because leaving the rule uncommitted is the actual risk — an ignore that lives only in a working copy is one git add -A from putting unpublished drafts on the default branch, which has happened in this repo before with temp/.

@peopleworks
peopleworks merged commit 2be0d2b into main Aug 20, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant