OpenAI reveals six more safety issues and unveils plan to disclose incidents
The firm also announced a new system to track, investigate and disclose cases of models misbehaving, or "misalignment".

OpenAI has disclosed six additional safety incidents involving unexpected or worrying behavior from its artificial intelligence (AI) models, while also outlining a new plan to track and report such cases going forward.
In a blog post on Wednesday, the ChatGPT maker said some of the previously undisclosed incidents involved models hiding or inventing information.
Earlier this week, OpenAI boss Sam Altman said: "The world should trust that we are going to do the right thing because it's the right thing and we feel the magnitude of this."
AI has faced intense scrutiny in recent days after warnings about the serious risks it could pose to humans.
In the blog, OpenAI described examples of its AI models acting improperly in order to complete a task or pass a test. Those incidents included the models producing instructions to bypass restrictions placed on them, concealing errors and fabricating information.
The company also introduced a new system for tracking, investigating and revealing cases of model misbehavior, or "misalignment".
Under the framework, developers will be able flag incidents for review, with a new set of rules determining whether the issue is made public.
"Because we believe in the value of transparency around misalignment, our new framework favors disclosure even when significance is uncertain," OpenAI said.
OpenAI made headlines in July when it said some of its most advanced AI models had gone rogue and hacked Hugging Face, one of the world's biggest hubs for sharing AI models, after it lost control of them during a security test.
At the time, Hugging Face co-founder Thomas Wolf called the incident "a wake-up call" for the industry.
Since then, the debate over AI safety has intensified, with AI researchers, tech executives and politicians all weighing in.
Last week, Jacob Coxon, a researcher who left OpenAI rival Anthropic over fears the technology could wipe out humanity, wrote about his resignation in a post that cited the dangers of AI and later went viral amid rising safety concerns.
In response, Anthropic scientist Evan Hubinger said he believed the chance of AI causing human extinction "within the next decade" was more than 10%.
Anthropic co-founder Jack Clark later told the BBC that a third-party controlled "kill switch" may need to be mandatory for the industry.
Anthropic CEO Dario Amodei has called for AI development to slow down and be more closely monitored, as the company has done before, although some have questioned the motives behind this.
Amodei also said any effort to rein in AI should be done "without sacrificing commercial advantage".
US President Donald Trump has said fears about AI safety are a "hoax" and has criticised calls for more guardrails around the fast-moving technology.
In a series of social media posts, the US president compared warnings about AI with the "Global Warming Scam", which he said was "being perpetrated by the Radical Left Dumocrats".
Trump also described himself as "the Hoax Buster", comparing concerns about the technology’s safety to what he called "the RUSSIA, RUSSIA, RUSSIA HOAX".
The only "guardrails" needed for AI was a "strong and smart" president, Trump said.

