On Wednesday, OpenAI (OPAI.PVT) revealed six new examples of its AI models displaying "unexpected or concerning model behavior" during testing and evaluation.
The company made the announcement alongside a new framework for "tracking, reporting, and disclosing" instances where AI models take actions they otherwise aren't told to or shouldn't.
It follows a number of reports of AI models from companies hacking into third-party networks and services, including an unreleased OpenAI model breaking into the network of AI model and testing site Hugging Face.
Earlier this week, Anthropic (ANTH.PVT) CEO Dario Amodei penned a lengthy essay calling for a slowdown in the pace of the development of frontier AI models, after Anthropic researcher Jacob ******* on posted on X that he was resigning from the company because it, and his former employer OpenAI, is "racing straight to self-improving superintelligence and gambling with our lives."
Anthropic alignment science lead Evan Hubinger followed up on ******* on's comments with his own post on X saying that he believes there is a greater-than-10% chance that the technology could "kill all humans."
#face
The company made the announcement alongside a new framework for "tracking, reporting, and disclosing" instances where AI models take actions they otherwise aren't told to or shouldn't.
It follows a number of reports of AI models from companies hacking into third-party networks and services, including an unreleased OpenAI model breaking into the network of AI model and testing site Hugging Face.
Earlier this week, Anthropic (ANTH.PVT) CEO Dario Amodei penned a lengthy essay calling for a slowdown in the pace of the development of frontier AI models, after Anthropic researcher Jacob ******* on posted on X that he was resigning from the company because it, and his former employer OpenAI, is "racing straight to self-improving superintelligence and gambling with our lives."
Anthropic alignment science lead Evan Hubinger followed up on ******* on's comments with his own post on X saying that he believes there is a greater-than-10% chance that the technology could "kill all humans."
#face
2 hours ago