LLMs respond differently to harmful prompts when AI watermarking is used
SynthID can cause models to follow harmful instructions they would otherwise refuse.

SynthID can cause models to follow harmful instructions they would otherwise refuse.

SynthID can cause models to follow harmful instructions they would otherwise refuse.
The page is ready to read now. The fuller skim-friendly version will appear here automatically.
SynthID can cause models to follow harmful instructions they would otherwise refuse.
Open the app view to save this story, compare related coverage, and continue from the same source.