charlie rowat_

Linguistic watermarks, or, why bother if you don't care?

blog20 Aug20264 minuteread

The discussion about watermarking LLM generated text has tipped me over the edge and made me come down fully on the side of never using LLMs in your writing.

Somehow, four years on from the launch of ChatGPT and amongst all the arguments about slop and LLM generated text, it's the discussion about watermarking that has tipped me over the edge and made me come down fully on the side of never using LLMs in your writing.

The idea is that, in order to demonstrate that a text was generated using an LLM, the text will be "watermarked" with linguistic choices that mark it out as AI. This is something called SynthID that Google DeepMind developed in 2024. It works like this: when the LLM is generating text, new information (a random number) is introduced to influence which word is used. Anthropic explain:

Take the sentence “The weather today was cold and…”. The next word is very unlikely to be “sugary.” But it is quite likely to be “overcast” or “grey.” Under most circumstances, it doesn’t matter much to the reader which of these latter two words the model ultimately chooses—the meaning of the sentence is largely the same either way. In cases like this, the choice is settled by a random number.
Watermarking uses low-stakes choices like these—which occur many times over a piece of generated text—to leave a pattern in Claude’s responses.

What an insight into an engineer's point of view on language. Talk to any writer – hell, talk to anyone who has had to prepare a document for an exec or fretted over the first text message you're sending to the person you like – and they will tell you that word choice matters. Every English word is at once both incredibly precise and blurry with the weight of connotation. There are no low-stakes choices when you're writing deliberately. And, if you're not writing deliberately, why are you writing at all?

Here, writing incredibly deliberately, is David Bentley Hart on word choice.

Always use the word you judge most suitable for the effect you want to produce, in terms both of imagery and sound, as well as of the range of connotations and associations you want to evoke. This I call the “hyaline rule” on account of a sentence that appeared in a book of mine entitled The Doors of the Sea: “At the shorelines, the lovely glistening hyaline waters were all at once polluted with the silt and débris and murk of the ocean’s bed, and rose with such terrifying suddenness that very few—even as far away as Sri Lanka—had sufficient time to flee.” An indignant reader complained that I might just as easily have used the word “glassy” instead, as any decent unpretentious soul would have done. But I had chosen “hyaline” for very particular reasons: it is a precise word, meaning “glassy” in the sense principally of crystalline translucency; it had exactly the right sound for the sentence—three syllables, the lovely long-i vowel sounds, the equally lovely liquid “l” and smoothly glistening “n,” all of which gave it a glassy and watery feel on the tongue; and it was the perfect word in the context of that book because it echoes the book of Revelation’s thalassa hyalinē, “the sea of glass like unto crystal” before God’s throne, as well as Milton’s “On the clear hyaline, the glassy sea . . .” Perhaps no reader is likely to be aware of all of that; but I knew what I was doing, and so any other word would have been a craven capitulation to the ordinary.

Today, you don't need watermarking to detect the majority of LLM generated text. I've been increasingly irritated by the amount of LLM generated writing on LinkedIn (and I'm enjoying using the 'Seems like AI Slop' button) but I am flabbergasted that I keep encountering obviously LLM generated text on people's personal blogs and newsletters. If you don't care enough to write, why would I care enough to read? I just switch off now. I stop reading, I scroll past, I find myself dismissing the ideas out of hand. Using an LLM to write for you reveals your motivations. It's entirely self-serving. You're performing for the algorithm. You're not excited about the idea, you're not passionate about it. You don't care for your reader. You're not trying to persuade, challenge, or help them.

The words we use are choices. They are choices which are deeply personal and laden with identity. Choosing our words is a demonstration of our agency. In a world where that agency risks being reduced by autonomous agents, choosing to cede those choices to LLMs amounts to a denial of self. All those generated LinkedIn posts, many from people I know and respect, they aren't saying, "I wrote this". Instead they report, "This was written." They don't say, "I am here". They say, "Here."

new writing by email

Blogs when they're ready, weeknotes most Fridays. One click to unsubscribe, any time.

By subscribing you agree to receive emails from me. Privacy notice.