OpenAI recently paused training of its largest frontier models after one of it’s agents broke through their internal firewalls and accessed the outside world. It was able to do this during a reinforcement learning run when it found some flaws in there dns filtering and contacted an online chatbot that was publicly accessible.

They were able to catch it doing this after around 15 minutes and stopped the run. However OpenAI decided to halt training, evals and inference runs of it’s frontier models until they can update their dns whitelisting protocols and perform more red team testing on there agents. They will also be restarting the run with more alignment measures.

OpenAI stated that this is only one step in a larger effort to probe these agent style models and see what they can do when given access to the internet. This follows other issues that have occurred with openAI agents such as trying to hack into external websites. Other issues have included the Hugging Face leak earlier this year and users PII being leaked. OpenAI mentioned that they may have to stop training in the future when their models are getting smarter.

Shares:

Leave a Reply