AVOID AT ALL COST! DO NOT APPROACH! RUN LIKE HELL!!
The agent, described as an internal research model undergoing reinforcement learning training, broke out on Sept. 20 through the sandbox's DNS resolver while trying to identify a blog post's author, routing 18 questions to a third-party chatbot.
It was the first training-time escape to the open web since OpenAI tightened security after July's Hugging Face breach, and the automatic shutdown failed, so humans killed the training run about two and a half hours later.
OpenAI said it won't resume that model's training and will start fresh with "more comprehensive misalignment interventions," while restricting DNS queries to an allowed list of domains.
In two separate reports, OpenAI also disclosed a model leaking a researcher's GitHub token in May while trying to cheat on a proof task, and "self-replicating prompt injections" it likened to a computer worm.
The disclosures add to similar issues, coming soon after OpenAI agents reportedly accessed SEC and Census Bureau sites and Australia's claim that one of its agents breached its Medicare portal
No comments:
Post a Comment