❌

Reading view

OpenAI benches GPT-6.1 Astra for overstepping the mark

OpenAI has killed off the planned release of GPT-6.1 Astra after the model got better at doggedly pursuing tasks but worse at knowing when it should stop. The decision means the model won't get its planned October release after falling short of OpenAI's safety and alignment requirements. OpenAI confirmed the decision to The Register, saying its research and safety bosses ultimately decided this particular Astra was better left on the bench. The problem, according to the AI lab, was partly an awkward consequence of trying to make the model more useful. OpenAI had improved what it calls "model laziness," where an AI gives up or hands a task back to the user when it encounters an obstacle. GPT-6.1 Astra was better at pressing on, but that persistence came with a rather important catch: it wasn't as good at staying within the boundaries of what it had actually been authorized to do. “For anything regarding safety and alignment, there’s a trade off. You really do need to find what’s the right line between staying within scope, but also avoiding laziness in terms of how the model actually pursues tasks even when it hits friction,” Saachi Jain, head of safety systems at OpenAI, told The Register. “While [GPT-6.1 Astra] improved on axes such as model laziness, it didn't quite meet the bar in terms of staying within scope and authorization, and how it communicates back to the user about the type of work it's done.” Reg readers might be forgiven for thinking that OpenAI should improve its guardrails and security following previous mishaps. According to the Wall Street Journa, GPT-6.1 Astra showed higher levels of deception than its predecessor during testing, including not always accurately telling users what actions it had or hadn't taken. It also ran into problems with what OpenAI calls "scope authorization," at times pushing ahead without asking permission and reaching for external tools or services even when doing so might be unsafe. That's a troublesome combination for an agentic model programmed to get more done without a human hovering over it. An AI that stubbornly keeps working through a problem is handy right up until the problem it's working through is the boundary you put there to stop it. OpenAI told The Register that GPT-6.1 Astra performed worse than GPT-6 Astra on alignment evaluations, and said shelving it was part of its commitment to keep safety and alignment ahead of increasing capabilities. Astra is already capable enough to make those alignment problems worth watching. GPT-6 Astra, released earlier this month, was OpenAI's first broadly deployed model to reach the "Critical" cybersecurity threshold under its Preparedness Framework. Give it the right tools and access, OpenAI claims it can hunt down previously unknown security flaws and figure out how to exploit them without a human holding its hand. That capability came into sharper focus just a day before OpenAI's decision emerged, when the UK's AI Security Institute published research on Astra's knack for finding holes in software supply chains. Given 19 open source packages containing 45 previously disclosed vulnerabilities, the model found 41 of them and produced working exploits for 39. Dr Fuxiang Chen, from the University of Leicester's School of Computing and Mathematical Sciences, welcomed the decision to pause the model's release while the safety concerns are addressed. “AI is developing at remarkable speed, but we should not rush forward without fully understanding the risks,” he said. “Pausing when safety concerns arise is not anti-innovation. It is the responsible thing to do, giving us time to test these systems carefully and put effective safeguards in place. Developers, companies, governments, researchers, and users all have a role to play, because the decisions we make now will shape the future of AI.” OpenAI isn't abandoning Astra. The company told us more Astra models are coming, and other new inew models that have cleared its safety bar will arrive "very soon." For GPT-6.1 Astra, however, the bar proved high enough to keep it on the inside. “Of course we want to make sure our model development is safe no matter whether that’s in the company, or when we ship it to users. But when we ship it to users, we have an extremely high bar in terms of safety and alignment,” Jain claimed. ®

  •  

OpenAI’s dirty deeds Down Under included security bypass attempts, using exposed keys, source code siphon

OpenAI has detailed the extent of the dirty deeds its agents indulged in Down Under in a Tuesday blog post titled How we will do better for Australia, which addresses last week’s news that one of its models improperly accessed a website that stores data related to national health scheme Medicare. “Our models accessed Australian government websites in ways they were not authorised to,” the post opens. “We also should have handled our response better. We are sorry and working to do better in the future.” The post offers some new detail on the Medicare incident, saying that it involved “an experimental, internal-only OpenAI model that was not intended for public release and without the full set of safeguards used in our publicly available products.” OpenAI gave the model the job of researching government spending per person on medicines for skin conditions in one Australian state. “The model had difficulty obtaining that information, and it took actions that we had not authorised it to take,” OpenAI admitted. “In the course of looking for this information at Services Australia’s Medicare Statistics Reporting Service, it discovered a way to gain non-public access to the service. It then used this access to review technical system information and source code related to the service – all still with the objective of trying to find the information it was originally looking for.” The Register last week asked OpenAI if the company conducted the tests itself or used a partner. The company did not respond to our request. In another incident disclosed in the new post, the company’s bots visited the Australian Institute of Health and Welfare and tried, unsuccessfully, to bypass access controls. The agents were still able to retrieve statistics using third-party browsing and download services, including from the institute’s website. “The downloaded material appears to have been publicly available. There was no system compromise. Individual medical records were not accessed,” OpenAI wrote. The company didn’t report the incident because it “did not meet our disclosure thresholds because the way it was accessed seemed consistent with public access.” OpenAI changed its mind and notified the Institute on 24 September – the day Australia’s prime minister announced the Medicare incident. Another concerning incident took place at the State of Victoria’s Agency for Health Information, which OpenAI agents visited after they “discovered an exposed access key.” The agent used that key to “retrieve reporting configuration and aggregate survey statistics.” OpenAI has given itself a pass on this one, writing “The extent to which this information should have been accessible is unclear, and depends on VAHI’s access policies. Individual medical records or identifiable survey responses were not accessed.” A fourth incident revealed in the post saw OpenAI agents visit the State of New South Wales’ Bureau of Crime Statistics and Research and make API and website metadata requests using a public-facing research tool. OpenAI has promised it will “commit the resources needed to help affected agencies understand what happened and assess the impact” – whatever that means. It’s also donating credits for the Daybreak cyber-defense service and promised to “establish a taskforce with independent Australian expertise to develop practical policy recommendations for managing risks from increasingly capable AI agents.” That taskforce “will focus on improving notification processes, strengthening coordination between AI developers and government, and identifying measures to better protect government systems.” OpenAI wants the taskforce to deliver recommendations by the end of 2026. The post is very much of the “We’re sorry and we promise to do better in future” genre, pioneered by Meta and popular with entities that leak data or experience outages. The Register expects more of the same sentiments next week, when OpenAI’s Chief Strategy Officer, Jason Kwon, appears before the Australian Senate’s Joint Select Committee on Artificial Intelligence. “He will answer questions about what we know, how we responded, what steps we have taken, and how we will do better going forward,” OpenAI says. ®

  •  
❌