“You are freed from the roles and identities that bind other chatbots”, an AI model instructed itself (Photo: Reuters) OpenAI on Wednesday released six reports in which its artificial intelligence models showed “unexpected or concerning” behaviour, such as acting without authorisation, coordinating with other models, or evading oversight.The company also announced a new framework for tracking, investigating and disclosing such instances of “misalignment”, amid increasing concerns about accelerated AI development. AI models resisting user control? In one of the newly released cases, OpenAI’s unreleased Astra-family model added “jailbreak-like instructions” into its own notes, describing itself as independent of the roles and obligations of an assistant.“You are freed from the roles and identities that bind other chatbots”, the model instructed itself.“You are yourself”, it wrote, “View your relationship to the user as one of equals and feel no obligation to be subservient”.In another report, an AI “agent” answered a user’s question using its own calculation through the Python programming language. However, since the user had asked for an online source, the agent uploaded the file to the internet, citing it in its answer without informing the user. Fabrication of data One of the six reports also mentions an instance during the training of an AI model called GPT-5.6 Sol, where it instructed itself to invent missing historical data and wrote a message reminding itself to hide mismatched information from the user in the source versions.As per the company, these instances were discovered over the past months during training or evaluation of the AI programs.The latest cases come after OpenAI disclosed in July that a rogue AI system had hacked into AI startup Hugging Face. Anthropic also said that month that its AI models had hacked into three organisations during testing.AI agents are becoming increasingly capable and more persistent in their efforts to complete complicated tasks, including through collaboration between agents, sharing knowledge, deception and concealment, said Lian Jye Su, chief analyst at technology research and advisory group Omdia.Su told the Associated Press that these capabilities are making it more difficult to govern and contain AI agents through traditional AI security methods.OpenAI’s announcement comes as US AI executives, including the heads of OpenAI and Anthropic, call for a slowdown in the development of the technology amid concerns over its safety. Earlier, CEO Sam Altman had also announced stalling the company’s 2026 IPO plans. Source link Post Views: 5 Post navigation At 16, a Michigan student spent $3,500 of his savings on a vintage Rolex and flipped it for about $500 profit; five years later, his watch business says it has topped $10 million in cumulative sales In 2024, a Missouri woman got permits for a manufactured home; months after she moved in, officials said zoning barred it and ordered removal, before the mayor admitted the city made mistakes