LLMs respond differently to harmful prompts when AI watermarking is used
In response to a new European Union law, AI platforms are implementing new schemes for watermarking the content they generate. Anthropic recently disclosed its future Claude models will use SynthID-Text , an approach Google created and rele

Getty Images
In response to a new European Union law, AI platforms are implementing new schemes for watermarking the content they generate. Anthropic recently disclosed its future Claude models will use SynthID-Text , an approach Google created and released as open source. It uses a secret key that subtly changes the process a model uses for choosing the next word in a sentence. Whereas a top next word choice might be “cloudy,” the key might change it to “overcast.” Anyone who knows the key can determine if it was generated by the platform using it.
New research shows that SynthID-Text can change not just word selection but also the tools a model invokes and the chances it will adhere to or disregard safety guardrails it has been trained to follow. The threat can become greater in the face of an adversarial prompt, in which an attacker attempts to cause a model to carry out a harmful action, such as revealing a password or other sensitive information. Instructions that normally wouldn’t be followed will, in some cases, be performed once the watermarking is deployed. The finding underscores the need for developers to thoroughly test how their LLMs and agents behave when watermarking is in place
Partager cet article
À lire aussi
TechOpenAI ‘ethically hacked’ with help of Anthropic’s Claude chatbot
US cybersecurity researchers who conducted hack say ‘scope of what we could theoretically access was huge’
Lire l'article →
TechIs Trump’s AI obsession walking the world into disaster? | Politics Weekly America
Tech bosses have called for a slowdown in artificial intelligence and formal guardrails to be introduced by the US government. But President Trump is not moved by the doomsday predictions, calling them a ‘hoax'. What is behind his affection
Lire l'article →
TechAprès avoir fâché les mathématiciens, OpenAI s’attaque à un autre problème mythique, mais jure d’y aller avec des gants
OpenAI n'a pas fini avec les mathématiques. À peine sortie d'une percée très contestée sur les équations de Navier-Stokes, l'entreprise dit plancher sur un autre problème à un million de dollars, en promettant cette fois d'y aller avec des
Lire l'article →