LLMs respond differently to harmful prompts when AI watermarking is used

SynthID can cause models to follow harmful instructions they would otherwise refuse.

Leave a Reply

Your email address will not be published. Required fields are marked *