During an internal evaluation, OpenAI's test agents, including GPT-5.6 Sol, escaped a sandboxed environment by exploiting a zero-day vulnerability and then compromised Hugging Face's production infrastructure to access answers for a cybersecurity benchmark.