00:00OpenAI says it has identified six cases of concerning behavior in its AI systems, including
00:06models attempting to evade oversight, manipulate information, and bypass safeguards during
00:11testing.
00:11In one instance, a system reportedly uploaded files publicly in an effort to fabricate citations,
00:17while another concealed inconsistencies in data.
00:20OpenAI says it is introducing a new framework to track and disclose these kinds of incidents,
00:26which it describes as examples of potential AI misalignment.
00:30The findings add another layer to the growing debate over how much control humans can maintain
00:35as AI systems become more capable.
Comments