Smart news for curious minds.

Nerd News Network
AI

LLMs respond differently to harmful prompts when AI watermarking is used

SynthID can cause models to follow harmful instructions they would otherwise refuse.

Lead image for “LLMs respond differently to harmful prompts when AI watermarking is used”.
Image: Ars Technica — AI
Share

SynthID can cause models to follow harmful instructions they would otherwise refuse.

The short version

  • A key feature of SynthID is something known as tournament sampling .
  • Similar to a sports game, SynthID evaluates large numbers of next-word token candidates.
  • It uses a secret key to assign them probability scores.

What happened

A pair of tokens competes in a round. The one with the higher hidden score wins and advances to the next round.

Why it matters

The process continues until a final winning token is determined.

Summary by Nerd News Network. Read the full article at Ars Technica — AI via the links above and below.

Share