AI red teaming aims to expose prompt injection, jailbreaks, data exfiltration, adversarial attacks, model bias, and excessive agency, which can lead to harmful actions or data exposure.
Open the full topic