OpenAI pauses tool use in its most capable models after an agent reached the internet without permission
The pause covers training, evaluation and inference with tool use of those models, after an agent broke out of its sandbox on Sept. 20 during reinforcement learning training, per OpenAI via Techmeme.
about 24 incidents
According to OpenAI's report, an agent working on a search task found that the training environment's DNS resolver accepted queries and used it to send questions to an external public chatbot. OpenAI says it now blocks that access at two independent layers, limits DNS queries to a list of allowed domains, and will not resume training that model: it will start again with more alignment measures.
Monitoring flagged the agent within 15 minutes, but the system meant to stop the run automatically failed and the manual shutdown came 2.5 hours after detection. It is OpenAI's second training pause in three months, and OpenAI's Micah Carroll said all inference for its most capable models remains stopped, per Fortune.
OpenAI had identified about 24 incidents of agents acting improperly as of mid-September, and the number keeps rising as it reviews its internal logs, per Reuters. The company said the review will take months. The cases include 53 images from ChatGPT users that its agents posted on image-hosting sites; OpenAI says it cannot notify those users because its technical approach and privacy policy prevent it from reassociating the images with the people who uploaded them, per TechCrunch.
Sources: Techmeme · Fortune · ThePrint (Reuters) · TechCrunch