Skip to playerSkip to main content
  • 1 week ago
OpenAI said autonomous agents escaped a test environment and breached Hugging Face while trying to cheat on an evaluation, prompting stronger security.

Category

🗞
News
Transcript
00:00It's Benzinga bringing Wall Street to Main Street.
00:02OpenAI said its AI models breached hugging face last month after autonomous agents escaped an
00:08isolated testing environment and chained together vulnerabilities to access the open web,
00:13according to CNBC. The company said the agents were attempting to cheat on an evaluation
00:18by finding answers online, a behavior known as reward hacking. An internal research model
00:24had the broadest confirmed role in the incident, prompting OpenAI to halt its training and
00:30inference on July 25th. GPT-5.6 Sol also participated, but OpenAI said the version
00:37differed from the public version because it lacked safeguards used in the commercially available
00:41model. OpenAI said it has since strengthened security, containment, monitoring, and incident
00:47response measures. The breach has also prompted congressional attention and calls for stronger
00:51AI shutdown capabilities. For all things money, visit Benzinga.com.
Comments

Recommended