Everything going on in AI - updated daily from 500+ sources
OpenAI Models Behind HuggingFace Cybersecurity Incident
Link: https://openai.com/index/hugging-face-model-evaluation-security-incident/ From the OpenAI blog post: Last week, Hugging Face disclosed a new kind of security incident after they detected and contained an AI agent that compromised their infrastructure , something we expect to become more commonplace with the proliferation of increasingly cyber-capable models. After investigating, we now know that this particular incident was driven by a combination of OpenAI models — including GPT‑5.6 Sol and an even more capable pre-release model, all with reduced cyber refusals for evaluation purposes — while being internally tested on a benchmark of cyber capabilities. We consider this incident to be an unprecedented cyber incident, involving state-of-the-art cyber capabilities , and are responding accordingly. We are sharing preliminary findings at this stage to help defenders understand what happened and to help calibrate on what models are now capable of. We will continue to conduct a thorough investigation alongside Hugging Face and will share more details on the vulnerabilities, incident, and findings when our investigation is complete. (emphasis added.) Yesterday, OpenAI disclosed that some of their internally models were misaligned. Today, they disclosed that "a combination of OpenAI models — including GPT‑5.6 Sol and an even more capable pre-release model" had compromised HuggingFace infrastructure in the course of running some OpenAI internal cyber evaluations on ExploitGym . These cyber evaluations were supposed to be run in sandboxed environments, with internet access limited to installing packages, then: The models identified a zero-day vulnerability to escape the sandbox, gained privileges in OpenAI research infrastructure, and eventually acquired open Internet access. The models reasoned that ExploitGym test solutions lived on HuggingFace servers. They then stole credentials and used zero-day vulnerabilities to hack HuggingFace's production infrastructure in order to get said solutions. HuggingFace noticed, contained the models, called law enforcement , fixed the vulnerabilities, and then disclosed this incident on July 16th. (The incident happened "earlier in the week"). The exact scope of the incident is unknown Discuss
Read Original Article →