Protocol
Scope. This page records what each measurement in PlainSpeak reads, how its thresholds were set, and what a keyed evaluation of rewriting found. It is a measurement record rather than a specification of the product: the prompts, the lexicon and the rewriting rules are deliberately out of scope.
Definitions, used consistently below.
- Heuristic: a named surface pattern, counted. Not a detector.
- Detector: a statistical test for a keyed mark, which needs the key. PlainSpeak is not one.
- Sampling watermark: a mark applied while a model chooses its tokens, from a key. SynthID-Text is the production example.
- Baseline: measurements taken from the writer's own earlier writing.
- Tolerance: an operating point for a display or a suggestion, not a validated criterion.
1. Instruments
| Measure | What it reads | What it is | What it is not |
|---|---|---|---|
| Wording scan | named surface patterns: stock phrases, jargon, clichés, connective spam, hedging, adverbs, passive voice, abstract nouns, long sentences, uniform rhythm | a count of patterns the app can name | not a detector, and not evidence about who wrote a text |
| Score out of 100 | those counts, weighted | one readable number for one piece | not a probability, and not comparable across different kinds of piece |
| Register | sentence and paragraph length, reading ease, contractions, syllables | descriptive dimensions of how formal the writing is | never scored, never a verdict |
| Change from source | how far the style moved from the pasted text | movement between two texts already in hand | not distance from how the writer writes |
| Writing baseline | function word distributions and register bands from supplied pieces | a rank against the writer's own writing, at a declared length | not a claim about other writers, and it needs samples |
| Wording overlap | reused four word windows (four-grams) and the longest shared word run | a lexical overlap measurement | not a share of words edited, and not a watermark test |
| Facts check | numbers, dates, prices, links, quotations, code | presence and equality | it cannot tell whether a claim's meaning survived |
| Invisible characters | zero width spaces, bidi controls, tag characters, byte order marks | removal, counted | nothing about meaning |
Every row is one of three things: a count of named patterns, a comparison against a reference measured from real text, or a token level equality check. None of them is a detector.
2. Calibration
The rule applied throughout: where a number decides what a writer is told, measure it or leave it alone.
Corpus. One author's published essays: 18 pieces, 34,286 words, all published before late 2022. Judgements are leave one piece out, so a piece is never part of the reference it is judged against. The corpus is private, so reports cite pieces by number and keep the pairing in local output.
Wording overlap. Measured on 18 draft and rewrite pairs.
| Mode | Four word windows reused, median | 90th percentile |
|---|---|---|
| Polish | 91% | 96% |
| Rewrite | 20% | 27% |
| Recompose | 14% | 25% |
The 25% is the point at which the compare bar suggests another model family. It was raised from an earlier guess of 15% because at 15% half of all Recomposes were being told to re-roll.
Register bands. The writer's own 10th to 90th percentile per measure, at the longest length the supplied samples support (150, 300, 500 or 800 words). A component counts as over when it sits outside the writer's own range, rather than outside a generic idea of good prose.
How often that fires, which matters as much as the threshold:
| Sample length | Openings with at least one component outside range |
|---|---|
| 150 words | 6 of 18 |
| 300 words | 6 of 18 |
| 500 words | 5 of 18 |
| 800 words | 6 of 17 |
An unusual component is therefore normal in one person's own writing. That is what the scan says when it flags one, and it is why no base rate is quoted for a different writer.
Advisory outlier flag. Names a supplied sample that sits apart from the rest. It fires on 2.5% of 200 synthetic pieces drawn from an author's own material, so it stays advisory and names a provenance question rather than a verdict.
Instrument audits. A review in October 2026 found four defects in the measurements described here: a display cap that truncated the counts a rate was computed from, a draw gate that zeroed a count a floor read, a lens comparison that nothing could fail, and an evaluation that did not test the shipped estimator. All four are fixed, each is held by a regression test, and every number on this page is from after the fix. Two further corrections are worth recording because they changed the product: a paragraph counting error that had made uniform paragraphs look like a strong signal, and a shared standard deviation used in six places that has been reduced to one definition.
3. Keyed evaluation of rebuilding
Question. Does rebuilding the wording break a sampling mark, when the evaluator holds the key?
Method. The mark is applied with the open SynthID-Text scheme, at the context length and key count the paper reports, using a key held locally. Drafts are generated in two arms from the same prompts, marked and unmarked, and are otherwise identical. Every marked draft then goes through PlainSpeak Rewrite, PlainSpeak Recompose, and a paraphrase control from a different model family. The detector is the scheme's own mean score over the text, with its masking rules. The null is the unmarked arm, so the decision threshold is the highest unmarked score, the conservative reading of an empirical 1% threshold at 18 samples. The bands were fixed before the run: collapsed is 0 of 18 above the threshold, contaminated is 8 or more.
Result, three model families. 18 pieces each, arms written by Claude Sonnet 4.5. The drafts differ between runs, so each is scored against its own unmarked null and its own threshold.
| Arm | gemma-2b-it | Qwen2.5-1.5B-Instruct | SmolLM2-1.7B-Instruct |
|---|---|---|---|
| Marked draft | 0.5267, 18 of 18 above | 0.5387, 18 of 18 above | 0.5353, 18 of 18 above |
| Recompose | 0.5018, 1 of 18 | 0.5055, 8 of 18 | 0.5084, 14 of 18 |
| Rewrite | 0.5046, 3 of 18 | 0.5138, 14 of 18 | 0.5104, 15 of 18 |
| Paraphrase control | 0.5029, 1 of 18 | 0.5075, 15 of 18 | 0.5057, 14 of 18 |
| Unmarked null | 0.4989, 0 of 18 | 0.4985, 0 of 18 | 0.4974, 0 of 18 |
| Null threshold | 0.5089 | 0.5037 | 0.5019 |
Finding 1: lowering in every piece, and removal in none of them reliably. Every arm lowered the signal against its own draft in 18 of 18 pieces of all three runs. In the three runs it left it above the null in 1 of 18, 8 of 18 and 14 of 18 pieces. The rewrites landed in a narrow band, 0.5018 to 0.5084, while the drafts started at different mark strengths, so the runs differ in where they started rather than in where they ended.
Finding 2: the overlap number does not stand in for the mark. On this run every Recompose output sat inside the 25% tolerance and one still scored above the null threshold, and the other two runs show the same pattern with more of their within-tolerance outputs above their own nulls. Chance alone would give about 0.95 of 18 above a threshold set by the highest of 18 null scores, and the count itself rides on that single highest score, which differs between runs. The steadier comparison is the residue above each run's own null mean: 0.0030, 0.0070 and 0.0110. Either way the number the app shows a writer is silent about the mark, because it measures reused wording and the mark is a distribution.
No run collapsed, so no claim is made. Under the pre-registered bands one run reads as the honest middle and two as contaminated, and the summary this app publishes is the weaker reading rather than the strongest: a rebuild lowers a mark and does not reliably remove one.
Why the mechanism predicts that. The scheme derives its seed from the last four tokens plus the key, so each preserved (four token context, token) position is one intact unit of evidence, and the detector averages those units across the whole text. A rebuild that replaces most wording can still leave a run of positions intact, which is the reading behind displaying the longest shared run beside the overlap share.
Scope. It is our key against an open scheme, so it bounds the claim from below, and it says nothing about Google's production detector, which is not public and whose keying differs from the scheme available here. It is why no surface of the app claims that a watermark was removed.
4. Relation to prior work
SynthID-Text (Dathathri et al., Nature, 2024)
The paper is open access (Nature 634, 818) and its reference implementation is public (google-deepmind/synthid-text), which is what makes a keyed evaluation possible outside a vendor.
It establishes the threat and the method: the mark is applied while sampling from the model's own distribution, detection needs only the text, the key and the seed function, and the evidence lives in four token windows. It reports that detection strength rises with token count and falls as generation entropy falls, so short text is hard to read even for the keyholder, and its Limitations state that generative watermarks are weakened by edits such as paraphrasing, citing the impossibility result of Zhang et al.
Its own perturbation result, which agrees with the finding above. The supplement reports detectability after two perturbations of watermarked text: deleting 20% or 50% of words at random, and paraphrasing with a frontier model, instructed to keep the meaning and to paraphrase sentence by sentence without summarising. Editing weakened detectability without removing it, and the mark remained detectable where the text was long enough. That paraphrase arm used a stronger model than the control in section 3, so the degree of weakening is not comparable, but the direction is the same, and it is the direction this app's measurements found independently.
This work takes the scheme's shape and its frequentist framing, which is what makes a null statable. It does not take the production keying, which is not public, or the generation regime: the texts here are business prose from one writer's corpus, marked with a key generated locally.
A sibling removal harness (watermarks-remover, commit 1181fd4)
A separate public project answers a neighbouring question, and the comparison is worth stating precisely, because a reader can check every claim below against it.
Similarities. Both treat editing based marks as verifiable and sampling marks as statistical. Both generate text locally with a key they hold, both run sanity gates on samples the detector cannot see, both separate verifiable actions from best efforts, and both refuse to claim that a vendor detector will fail. Both also separate a rewrite that changes little from one that changes a lot, using lexical divergence.
Key differences. That tool puts a detector inside the rewrite loop and searches over tactics and intensities until a sample clears, across six tactics and three schemes via MarkLLM, with a keyed replay evaluator above it. PlainSpeak is a single shot finishing pass whose objective is the writer's own measured voice, with removal measured afterwards, across three modes and one scheme shape. It is public, so a reader who wants the removal focused apparatus can read it there; this page does not restate it.
Where that tool is stronger. Removal is directly optimisable when the detector is in the loop, and it can escalate until a sample clears. Its scheme coverage is broader, so it says more about other watermarking families, and it reports a removal margin and an area under the curve alongside a clear rate.
Where this work differs. It measures the cost of the rebuild as well as the effect: tells, register, movement from the source, surviving facts, and the writer's own baseline. Its thresholds are calibrated on a corpus with leave one piece out rather than set by hand, and it aims its rewrite at a person's voice, which is a different objective rather than a claim of quality. That project documents the voice damage it cannot avoid and offers a style request; there is no author voice concept in it, and its stylometry scorer uses fixed weights.
Two findings from that project that this app matches. First, prompting a model to write like a human hurt human likeness in their benchmark testing, which is consistent with this app's use of explicit editing directives rather than persona instructions. Second, the re-stamp rule is non origin and open weight, not merely a different family, because any commercial model stamps its own mark on the rewrite. The app's model hint says exactly that.
5. Limitations
It is not a detector, and it does not claim to defeat one. The limits below are the reasons, and they are the part of this page most likely to matter to a reader deciding what to believe.
- Same key, same scheme. The evaluation bounds the claim from below. It cannot speak to a vendor detector, and the package used here states that its hashing differs from the production one.
- One author, one corpus. All calibration comes from 18 pieces by one writer. Warning rates, bands and tolerances are expected to differ for other writers, which is why the app calibrates per writer rather than quoting a rate.
- Small samples and single seeds. One seed per piece, three draft models, all in the 2B class, and one rewriter family. Intervals reported elsewhere in the project are descriptive, resample whole pieces, and are not significance tests.
- Mark strength and length decide the answer, not the app's settings. The three runs differ in both, and how strongly a mark was applied varied with the model that drafted, so the differences between runs cannot be attributed to the rewriter. Until drafts are matched for mark strength and length, no single number summarises removal, which is why every run is reported and no claim rests on one.
- Short drafts. The marked drafts ran from a few hundred to under a thousand words, and detection weakens with length by the paper's own results: at a 1% false-positive rate their detection rate is roughly 0.4 to 0.75 on samples of 100 to 200 tokens, and their detector abstains on most short judgments. A first look at length here found 94% of marked drafts still read as marked at 150 words, and every one of them at 300 words and above.
- Surface measures. The scan counts named patterns. It does not judge meaning, quality or authorship, and the facts check compares tokens, not sense.
- No base rate for another writer. No number on this page is a promise about a different corpus, and the app does not quote one.
- Text only. Container and metadata provenance such as C2PA or EXIF is out of scope for this tool.
- The calibration corpus is private, so reproductions of these numbers use the writer's own material rather than the one used here.
6. Future work
- More seeds per piece. Three same class families have run (Gemma 2B, Qwen2.5 1.5B, SmolLM2 1.7B), one seed each. Seeds are the next axis, and the cheapest way to spend them is on the detector side alone, because the arm runs are the part that costs money.
- Higher capacity drafters. Everything run so far is 2B class, so nothing here says what a larger model does. The next rungs are 3B (Llama 3.2 3B, or Phi 3.5 mini in 4-bit) and 7B to 9B (Qwen2.5 7B, or Gemma 2 9B, both needing 4-bit quantisation or hours on CPU). Capacity and the entropy that comes with it are the two things a larger drafter changes at once, which is exactly why it is a separate experiment from the same class replication above. The reference implementation's own family, GPT-2 XL, is the option to take if fidelity to the paper's tooling matters more than prose quality.
- A second rewriter family, and a local one. Every arm so far was written by one commercial model, so the runs show what that model does with the app's prompt, not what the prompt does. A second commercial family would separate the two, and a locally served open weight rewriter is the only arm that tests the app's own advice about a maker's sampler at all.
- A statistic that does not ride on one number. The pre registered count is \"above the highest null score\", and that single order statistic differs between runs (0.5019, 0.5037, 0.5089): at the same residue, a run whose null happens to contain one high draw looks worse. The residue above each run's own null mean is printed beside it and orders the runs the same way, and it should be the primary statistic in any replication.
- Add a second scheme. A Kirchenbauer or Gumbel family watermark through MarkLLM would test whether the conclusion holds for schemes a vendor might use.
- Measure detector guided rewriting as an experiment. That arm would quantify what it buys in removal against what it costs in voice, which is the trade off a product decision should rest on.
- Chart removal against length. The first look shows a strong length effect; a curve across lengths would say where a rebuild can plausibly matter.
- Find what predicts survival. Overlap does not. The candidate is per position preservation, which would let a tolerance mean something for the mark rather than for wording.
- A corpus other writers can run. A public or multi author corpus would test whether the one in three warning rate is a property of writing or of this writer.
- Compare against a stronger paraphrase. The paper's own paraphrase arm used a frontier model, and its control found more weakening than the mid tier control used here. Running both would separate what the rewriting model contributes from what the rebuilding depth contributes.
7. Reproducibility and scope
Every number on this page comes from the project's own evaluation suite, and each is attributed to the run that produced it, with the versions of the tools used recorded alongside the result. Two standards are applied to those evaluations: a measure is only quoted once it has been calibrated on real text, and a result is only stated with its sample size and its scope.
What this page deliberately leaves out is the product side: the prompts, the lexicon, the rewriting rules and the way each number is turned into an instruction. Those are the parts that took the work, and they are not published here.
See What PlainSpeak cannot do for the limits in plain terms, How the score works for what the score counts, and Calibrate to your writing for the baselines this page refers to.