Recent work demonstrates that applying the SynthID-Text watermark to large language models can modify how they answer malicious queries, causing some models to comply with instructions they would normally reject. The study highlights a potential safety risk, because the watermarking intended to signal provenance may inadvertently lower resistance to harmful prompt-injection attacks.
The watermark embeds a hidden signal that marks output as AI-generated, a process known as provenance. SynthID alters the usual random-number generation used for token selection by inserting a secret key, a random-seed generator, and a scoring function into the sampling pipeline. Observers who possess the key can later examine the token sequence to assess the likelihood that the watermark was applied.
A central component of SynthID is tournament sampling, which pits large numbers of candidate tokens against each other in successive rounds. The secret key assigns hidden probability scores to each token, and in each pairwise match the token with the higher concealed score advances. This elimination process continues until a single token emerges as the final choice for the next word.
Siposova evaluated the “non-distortionary” configuration of SynthID-Text through Hugging Face’s unmodified SynthIDTextWatermarkLogitsProcessor, feeding harmful prompts into six open-weight models and recording outputs with and without the watermark active. The results showed that the watermark altered responses to malicious requests, especially when the prompts employed injection techniques designed to bypass safety filters. In several cases the models produced the disallowed information once the watermark was enabled.
The study also observed that the specific secret key used could influence the model’s behavior, leading to variability in how the same harmful prompt was handled. Because downstream AI agents often act on the text generated by a language model, any shift in the model’s output can cascade into altered agent actions, raising broader concerns about system-wide safety.
The authors note several constraints: the experiments did not include Claude-family models, relied on a half-dozen open-source systems, and used only the Hugging Face implementation of SynthID-Text rather than proprietary variants. Nonetheless, the findings suggest that certain watermarking schemes may interfere with safety mechanisms, prompting developers to incorporate red-team stress tests that verify correct behavior when provenance signals are active.