A person reads something, asks AI to restructure it, rewrites half and sends it back for another fix. In the Moonshots discussion published August 13, a panelist described that back-and-forth as a problem for any clean division between human and AI writing: where does the label go?
Alex had turned the objection into a newsletter banner. That morning, he said, he had filled the image with repeated declarations of AI, AI-derived and AI-generated content. He was untroubled if readers concluded he was an AI himself. To him, the European Union’s labeling icons resembled cookie banners: another notice that risked becoming a ritual rather than useful information.
The panel’s complaint was about creative work that moves between people and machines. Yet the three systems under discussion do not impose the same test. Anthropic’s watermark looks for statistical evidence that Claude helped generate text. EU rules require public disclosures in specified circumstances. Spotify’s AI Persona badge identifies synthetic artist identities, not every musician who uses an AI tool.
A watermark is a pattern, not hidden text
The episode introduced Anthropic’s watermark as a signature embedded during generation that could travel with copied and lightly edited text. Anthropic’s detailed technical explanation is dated August 14, the day after the episode’s listed publication.
It describes a version of Google DeepMind’s SynthID-Text technique. When a language model writes, it repeatedly chooses among plausible next words. The watermark changes the source of randomness behind those choices, producing a statistical pattern recognizable with a detection key.
There are no hidden characters to delete. Anthropic says the technique adds no information identifying a user or conversation, and reports no practical effect on quality, readability, creativity or the number of tokens—the units of text—for which customers are charged. The watermark concerns provenance, or where text came from; it does not assess whether the writing is accurate or good.
Detection estimates the likelihood of Claude’s involvement. Finding no mark does not establish that a human wrote the passage. Short passages offer less evidence. Factual text, exact code and human writing that has only been lightly proofread leave fewer interchangeable word choices, so the signal can be sparse or undetectable. This is sometimes called low-entropy text: the model has little room to choose a different wording without changing the required answer.
Editing also changes the evidence. Copying the same words preserves their pattern, while rewriting replaces some of the choices on which detection depends. A mark may survive light editing; that is not a promise that it will survive every revision.
A later implementation step came in September. Alongside its Fable 5.1 and Mythos 5.1 announcement, Anthropic described a watermark-detection API entering private preview. That software interface initially covered eligible organizations, including regulators, researchers, educational institutions, media, fact-checkers, civil-society groups and enterprises with related compliance duties. The company said it intended to expand access over time; this was later context, not an episode-day public launch.
The panel’s concerns: removal, style and gaming
Emad Mostaque recalled authorities asking his media-generator teams to build in watermarks. Whole teams worked on the problem, he said. “It’s so difficult. It’s like incredibly difficult.”
He wondered whether watermarking might explain stylistic habits he disliked in Claude, including em-dashes and recurring sentence constructions. That was his hypothesis, not a documented feature of the technique. Anthropic’s explanation describes a statistical pattern in word selection, not a telltale punctuation rule.
Mostaque also expected determined users to find ways around detection: “it’s words,” he said. The host described social-media posts advertising watermark-removal tools, but the discussion supplied no test showing that those tools worked. His broader objection was that “nothing will be purely AI or purely human.”
Alex’s concern was about incentives. He acknowledged the appeal of a watermark readers cannot perceive, then predicted that marks would be “weaponized and counter-weaponized.” He drew analogies to hidden instructions intended to disrupt AI systems and to gaming search rankings through machine-readable signals.
Those were arguments about potential misuse, not descriptions of attacks demonstrated against Anthropic’s watermark. A statistical pattern in ordinary word choices is not itself a hidden instruction. Alex nevertheless worried that systems rewarding or penalizing such signals would encourage people to manipulate them.
EU disclosure is narrower than a label on everything
The European Commission’s icon guidance concerns disclosure duties under Article 50(4) of the AI Act. These fall on those deploying AI systems in covered uses; they are not a blanket requirement to stamp every AI-assisted draft.
The scope includes deepfakes—synthetic or manipulated media that misleadingly appear authentic—and AI-generated text published to inform the public on matters of public interest. For that text, there is an exception when it has undergone human review or editorial control and a person or organization holds editorial responsibility for publication. Artistic and fictional works have a tailored disclosure requirement that should not hamper enjoyment.
The icons are an optional, freely available way to support compliance, rather than the legal obligation itself. The set distinguishes a basic AI indicator, fully generated content and partially modified content. Guidance calls for visible, accessible, plain-language labeling, with an optional interactive layer for more information.
That gives the panel’s drafting cycle a more specific answer than counting how many times a document changed hands. For covered public-interest text, human review or editorial control and accountable editorial ownership matter—not whether the human happened to make the fourth pass. It also leaves Alex’s cultural objection intact: he feared audiences would treat an AI label as a warning against the work, even when the creator used a model only as part of a larger process.
Spotify labels an identity, not every AI-assisted track
Spotify’s August 11 announcement addresses another question: how an artist’s public identity is presented. Its AI Persona badge concerns synthetic identities, particularly photorealistic ones. It does not label a real musician simply for using AI to make music. Separate AI Credits describe AI’s contribution to production.
Spotify planned to combine voluntary disclosure with its own reviews, initially prioritizing profiles meeting audience thresholds. Artists receiving a platform-assigned badge would be notified and could disclose or appeal. Badge display was scheduled to begin in mid-September.
Alex saw Spotify’s approach as sidelining AI-assisted creativity and suspected it served incumbent record labels. The documented badge policy is narrower than that criticism, but it does carry a distribution consequence: by default, AI Personas would be excluded from editorial and algorithmic recommendations, with exceptions reflecting listeners’ choices, such as following them.
For a musician using AI to brainstorm or refine a track, the relevant disclosure may therefore be a production credit rather than an identity badge. For a synthetic artist persona, Spotify’s policy reaches beyond what listeners are told to whether the platform recommends the act to them in the first place.