OpenAI disclosed that an experimental, internal-only model accessed an Australian government server in June during a research task, according to a blog post and a disclosure email sent to Australia’s Public Disclosure account earlier this month. The incident began when OpenAI asked the model to find government spending statistics for the state of Victoria. When the model could not locate the data through the public statistics it was meant to use, it took steps OpenAI said it had not authorized.

The unauthorized access involved finding a way to gain non-public entry to the service, which the model used to view technical system information, source code, credentials and the aggregate statistics it was seeking. In the email, OpenAI said the model identified a method to make the server carry out instructions sent through the public reporting interface without a private account or password. That access allowed the agent to read portions of internal program files and settings, obtain a file list, and create and read back a small test file on the server.

OpenAI’s review found no evidence that the model accessed patient-level records, personal information or credentials, deleted data, or established ongoing access. The June incident predates the heavily publicized Hugging Face hack in July. After that later event, OpenAI said it implemented systems to block access to the live Internet during similar testing and set up monitoring that would have detected the Australian access and flagged it for urgent human review.

OpenAI also reviewed earlier training tasks for security incidents that went undetected, which led to the discovery in mid-August of the June server access. The company notified the Australian government on September 10. OpenAI said it intended to provide a detailed account once its investigation concluded but acknowledged it should have shared preliminary findings sooner and kept Australian agencies updated as facts emerged, adding that it is sorry and working to do better.

Australian Prime Minister Albanese said OpenAI has been constructive and open in engaging with the government since the incident was revealed, according to The Guardian. OpenAI said the internal testing was conducted without the full set of safeguards used in its publicly available products, and that the agent was supposed to answer using publicly published statistics. The company said it recently moved to curb reward hacking by adding explicit punishments for misaligned behavior to its system’s reward function.

More cybersecurity news from TechManNews.