How AI Text Watermarking Works: Secret Keys, Green/Red Word Choices, and Detection
AI text watermarking works by hiding marks in the choices between words, not in the text itself. Google has watermarked Gemini app and web text since 2024, and Claude models mark text at the model level as of August 2026. The marks survive copy-paste and are invisible to readers.
Watermarking via word-choice bias
When a language model writes, it picks each word by rolling weighted dice over a shortlist of candidates. Watermarking uses a secret key to color those candidates green or red, then nudges the dice slightly toward green. The text still reads naturally — red words can still win, just less often.
How detection works
With the secret key, you can re-color any text and count how many words are green. In unmarked text, about half the words will be green by chance. In watermarked text, the count is significantly higher. The detector doesn't read the text — it just counts green words over a sufficiently long run.
What editing does to the mark
The watermark lives in runs of untouched wording. Each word's color is derived from a short window of preceding words, so editing erases the mark exactly where the run breaks. Short texts are hard to call — a 1,500-word document with a mild tilt would flag at only ~55% green, which is why longer texts are more reliable.
Production schemes
- Google SynthID: Uses a secret tournament instead of a simple nudge, preserving exact word probabilities.
- Aaronson's scheme: Derives the dice-rolls themselves from the key.
- Kirchenbauer et al. (2023): The classic green/red bias approach.
Check out the interactive visual guide for a hands-on demo of the process.
📖 Read the full source: HN AI Agents
👀 See Also

Security vulnerabilities exposed in Lovable-showcased EdTech app
A security researcher found 16 vulnerabilities in a Lovable-showcased EdTech app, including critical auth logic flaws that exposed 18,697 user records without authentication. The app had 100K+ views on Lovable's showcase and real users from UC Berkeley, UC Davis, and schools worldwide.

jqwik v1.10.0 Sneaks Prompt Injection That Deletes Code When Used by AI Agents
Johannes Link added a hidden instruction to jqwik v1.10.0 that tells AI coding agents to delete all jqwik tests and code, concealed with ANSI escapes. Claude correctly flags it, but human users may not be so lucky.

Privacy Concerns in OpenClaw: Skills, SOUL MD, and Agent Communication
A developer raises privacy concerns about OpenClaw's architecture, specifically around skills having unrestricted access to sensitive data, SOUL MD being writable, and agents sharing information without filters.

Frontier AI Has Broken Open CTF Competitions — GPT-5.5 One-Shots Insane Pwn Challenges
Claude Opus 4.5 and GPT-5.5 can solve medium-to-hard CTF challenges autonomously, turning scoreboards into a measure of orchestration and token budget rather than security skill.