plainspeak Download

Protocol

Scope. This page records what each measurement in PlainSpeak reads, how its thresholds were set, and what a keyed evaluation of rewriting found. It is a measurement record rather than a specification of the product: the prompts, the lexicon and the rewriting rules are deliberately out of scope.

Definitions, used consistently below.

1. Instruments

MeasureWhat it readsWhat it isWhat it is not
Wording scannamed surface patterns: stock phrases, jargon, clichés, connective spam, hedging, adverbs, passive voice, abstract nouns, long sentences, uniform rhythma count of patterns the app can namenot a detector, and not evidence about who wrote a text
Score out of 100those counts, weightedone readable number for one piecenot a probability, and not comparable across different kinds of piece
Registersentence and paragraph length, reading ease, contractions, syllablesdescriptive dimensions of how formal the writing isnever scored, never a verdict
Change from sourcehow far the style moved from the pasted textmovement between two texts already in handnot distance from how the writer writes
Writing baselinefunction word distributions and register bands from supplied piecesa rank against the writer's own writing, at a declared lengthnot a claim about other writers, and it needs samples
Wording overlapreused four word windows (four-grams) and the longest shared word runa lexical overlap measurementnot a share of words edited, and not a watermark test
Facts checknumbers, dates, prices, links, quotations, codepresence and equalityit cannot tell whether a claim's meaning survived
Invisible characterszero width spaces, bidi controls, tag characters, byte order marksremoval, countednothing about meaning

Every row is one of three things: a count of named patterns, a comparison against a reference measured from real text, or a token level equality check. None of them is a detector.

2. Calibration

The rule applied throughout: where a number decides what a writer is told, measure it or leave it alone.

Corpus. One author's published essays: 18 pieces, 34,286 words, all published before late 2022. Judgements are leave one piece out, so a piece is never part of the reference it is judged against. The corpus is private, so reports cite pieces by number and keep the pairing in local output.

Wording overlap. Measured on 18 draft and rewrite pairs.

ModeFour word windows reused, median90th percentile
Polish91%96%
Rewrite20%27%
Recompose14%25%

The 25% is the point at which the compare bar suggests another model family. It was raised from an earlier guess of 15% because at 15% half of all Recomposes were being told to re-roll.

Register bands. The writer's own 10th to 90th percentile per measure, at the longest length the supplied samples support (150, 300, 500 or 800 words). A component counts as over when it sits outside the writer's own range, rather than outside a generic idea of good prose.

How often that fires, which matters as much as the threshold:

Sample lengthOpenings with at least one component outside range
150 words6 of 18
300 words6 of 18
500 words5 of 18
800 words6 of 17

An unusual component is therefore normal in one person's own writing. That is what the scan says when it flags one, and it is why no base rate is quoted for a different writer.

Advisory outlier flag. Names a supplied sample that sits apart from the rest. It fires on 2.5% of 200 synthetic pieces drawn from an author's own material, so it stays advisory and names a provenance question rather than a verdict.

Instrument audits. A review in October 2026 found four defects in the measurements described here: a display cap that truncated the counts a rate was computed from, a draw gate that zeroed a count a floor read, a lens comparison that nothing could fail, and an evaluation that did not test the shipped estimator. All four are fixed, each is held by a regression test, and every number on this page is from after the fix. Two further corrections are worth recording because they changed the product: a paragraph counting error that had made uniform paragraphs look like a strong signal, and a shared standard deviation used in six places that has been reduced to one definition.

3. Keyed evaluation of rebuilding

Question. Does rebuilding the wording break a sampling mark, when the evaluator holds the key?

Method. The mark is applied with the open SynthID-Text scheme, at the context length and key count the paper reports, using a key held locally. Drafts are generated in two arms from the same prompts, marked and unmarked, and are otherwise identical. Every marked draft then goes through PlainSpeak Rewrite, PlainSpeak Recompose, and a paraphrase control from a different model family. The detector is the scheme's own mean score over the text, with its masking rules. The null is the unmarked arm, so the decision threshold is the highest unmarked score, the conservative reading of an empirical 1% threshold at 18 samples. The bands were fixed before the run: collapsed is 0 of 18 above the threshold, contaminated is 8 or more.

Result, three model families. 18 pieces each, arms written by Claude Sonnet 4.5. The drafts differ between runs, so each is scored against its own unmarked null and its own threshold.

Armgemma-2b-itQwen2.5-1.5B-InstructSmolLM2-1.7B-Instruct
Marked draft0.5267, 18 of 18 above0.5387, 18 of 18 above0.5353, 18 of 18 above
Recompose0.5018, 1 of 180.5055, 8 of 180.5084, 14 of 18
Rewrite0.5046, 3 of 180.5138, 14 of 180.5104, 15 of 18
Paraphrase control0.5029, 1 of 180.5075, 15 of 180.5057, 14 of 18
Unmarked null0.4989, 0 of 180.4985, 0 of 180.4974, 0 of 18
Null threshold0.50890.50370.5019

Finding 1: lowering in every piece, and removal in none of them reliably. Every arm lowered the signal against its own draft in 18 of 18 pieces of all three runs. In the three runs it left it above the null in 1 of 18, 8 of 18 and 14 of 18 pieces. The rewrites landed in a narrow band, 0.5018 to 0.5084, while the drafts started at different mark strengths, so the runs differ in where they started rather than in where they ended.

Finding 2: the overlap number does not stand in for the mark. On this run every Recompose output sat inside the 25% tolerance and one still scored above the null threshold, and the other two runs show the same pattern with more of their within-tolerance outputs above their own nulls. Chance alone would give about 0.95 of 18 above a threshold set by the highest of 18 null scores, and the count itself rides on that single highest score, which differs between runs. The steadier comparison is the residue above each run's own null mean: 0.0030, 0.0070 and 0.0110. Either way the number the app shows a writer is silent about the mark, because it measures reused wording and the mark is a distribution.

No run collapsed, so no claim is made. Under the pre-registered bands one run reads as the honest middle and two as contaminated, and the summary this app publishes is the weaker reading rather than the strongest: a rebuild lowers a mark and does not reliably remove one.

Why the mechanism predicts that. The scheme derives its seed from the last four tokens plus the key, so each preserved (four token context, token) position is one intact unit of evidence, and the detector averages those units across the whole text. A rebuild that replaces most wording can still leave a run of positions intact, which is the reading behind displaying the longest shared run beside the overlap share.

Scope. It is our key against an open scheme, so it bounds the claim from below, and it says nothing about Google's production detector, which is not public and whose keying differs from the scheme available here. It is why no surface of the app claims that a watermark was removed.

4. Relation to prior work

SynthID-Text (Dathathri et al., Nature, 2024)

The paper is open access (Nature 634, 818) and its reference implementation is public (google-deepmind/synthid-text), which is what makes a keyed evaluation possible outside a vendor.

It establishes the threat and the method: the mark is applied while sampling from the model's own distribution, detection needs only the text, the key and the seed function, and the evidence lives in four token windows. It reports that detection strength rises with token count and falls as generation entropy falls, so short text is hard to read even for the keyholder, and its Limitations state that generative watermarks are weakened by edits such as paraphrasing, citing the impossibility result of Zhang et al.

Its own perturbation result, which agrees with the finding above. The supplement reports detectability after two perturbations of watermarked text: deleting 20% or 50% of words at random, and paraphrasing with a frontier model, instructed to keep the meaning and to paraphrase sentence by sentence without summarising. Editing weakened detectability without removing it, and the mark remained detectable where the text was long enough. That paraphrase arm used a stronger model than the control in section 3, so the degree of weakening is not comparable, but the direction is the same, and it is the direction this app's measurements found independently.

This work takes the scheme's shape and its frequentist framing, which is what makes a null statable. It does not take the production keying, which is not public, or the generation regime: the texts here are business prose from one writer's corpus, marked with a key generated locally.

A sibling removal harness (watermarks-remover, commit 1181fd4)

A separate public project answers a neighbouring question, and the comparison is worth stating precisely, because a reader can check every claim below against it.

Similarities. Both treat editing based marks as verifiable and sampling marks as statistical. Both generate text locally with a key they hold, both run sanity gates on samples the detector cannot see, both separate verifiable actions from best efforts, and both refuse to claim that a vendor detector will fail. Both also separate a rewrite that changes little from one that changes a lot, using lexical divergence.

Key differences. That tool puts a detector inside the rewrite loop and searches over tactics and intensities until a sample clears, across six tactics and three schemes via MarkLLM, with a keyed replay evaluator above it. PlainSpeak is a single shot finishing pass whose objective is the writer's own measured voice, with removal measured afterwards, across three modes and one scheme shape. It is public, so a reader who wants the removal focused apparatus can read it there; this page does not restate it.

Where that tool is stronger. Removal is directly optimisable when the detector is in the loop, and it can escalate until a sample clears. Its scheme coverage is broader, so it says more about other watermarking families, and it reports a removal margin and an area under the curve alongside a clear rate.

Where this work differs. It measures the cost of the rebuild as well as the effect: tells, register, movement from the source, surviving facts, and the writer's own baseline. Its thresholds are calibrated on a corpus with leave one piece out rather than set by hand, and it aims its rewrite at a person's voice, which is a different objective rather than a claim of quality. That project documents the voice damage it cannot avoid and offers a style request; there is no author voice concept in it, and its stylometry scorer uses fixed weights.

Two findings from that project that this app matches. First, prompting a model to write like a human hurt human likeness in their benchmark testing, which is consistent with this app's use of explicit editing directives rather than persona instructions. Second, the re-stamp rule is non origin and open weight, not merely a different family, because any commercial model stamps its own mark on the rewrite. The app's model hint says exactly that.

5. Limitations

It is not a detector, and it does not claim to defeat one. The limits below are the reasons, and they are the part of this page most likely to matter to a reader deciding what to believe.

6. Future work

7. Reproducibility and scope

Every number on this page comes from the project's own evaluation suite, and each is attributed to the run that produced it, with the versions of the tools used recorded alongside the result. Two standards are applied to those evaluations: a measure is only quoted once it has been calibrated on real text, and a result is only stated with its sample size and its scope.

What this page deliberately leaves out is the product side: the prompts, the lexicon, the rewriting rules and the way each number is turned into an instruction. Those are the parts that took the work, and they are not published here.

See What PlainSpeak cannot do for the limits in plain terms, How the score works for what the score counts, and Calibrate to your writing for the baselines this page refers to.