Top 3 Adversarial Attack Techniques on LLMs

Prompt injection, jailbreaking, and token manipulation are among the most prominent adversarial attack techniques used against Large Language Models (LLMs) to bypass safety mechanisms or extract sensitive information.

Sources

Open the full topic