Prompt injection, jailbreaking, and token manipulation are among the most prominent adversarial attack techniques used against Large Language Models (LLMs) to bypass safety mechanisms or extract sensitive information.
Open the full topic