Majorsafety alignmentAnthropic

Watermarking Alters LLM Responses to Harmful Prompts, Study Finds

Published
Sep 17, 2026 18:33 UTC

The implementation of watermarking schemes, such as Google’s SynthID-Text, has been shown to alter the behavior of large language models (LLMs) under adversarial conditions. Andrea Siposova, an AI security researcher at Lasso Security, noted that watermarking enables the identification of content generated by specific platforms and significantly affects model responses to harmful prompts. Siposova stated, "As compared to the same models without watermarking, it is definitely going to change their behavior, especially when we place it under adversarial conditions, or we make these models call tools when they’re powering an agent." This follows recent efforts to comply with EU regulations on AI safety, highlighting a growing focus on the implications of watermarking in AI content generation.

Summarised from the primary source with AI assistance under human editorial oversight. Turing Wire is not a primary source — read the original for the authoritative account.

Source: Ars Technica AI