Skip to main content
Get our mobile app
Download on the App StoreGet it on Google Play

AI Used to Hack AI

Researchers Used Anthropic's Claude to Take Over an OpenAI Employee's Account

White-hat researchers at Hacktron AI used Claude to take over an OpenAI employee's ChatGPT account and reach the company's internal GitHub.

Claude AI

Security researchers say they used Anthropic's Claude models to compromise an OpenAI employee's ChatGPT account and reach the company's internal GitHub environment, in an operation the firm disclosed this week.

The researchers, from the security startup Hacktron AI, reported the vulnerabilities to OpenAI and to Discourse, the forum software company, in July. The Wall Street Journal independently reported that the team gained access to an OpenAI employee's account and had a path to read and propose changes to private OpenAI software.

The entry point was OpenAI's own help forum, community.openai.com, which runs on Discourse. Hacktron says that until two months ago, any user or OpenAI employee logging into the forum could have had their ChatGPT and Codex accounts taken over. Because users can connect outside services to those accounts, the potential reach included GitHub, Slack and email.

According to Hacktron's account, the team confirmed remote code execution through an image upload on July 25, then ran Claude in an autonomous loop against its own test instance. The researchers said Claude's Opus model refused to write an exploit aimed at a remote system, so they routed the task through a proxy that made the target look like a capture-the-flag training exercise. The resulting exploit script was then used against OpenAI's instance.

The researchers say they took over an OpenAI employee's account whose Codex was connected to OpenAI's GitHub organization, but stopped short of examining any internal source code. To demonstrate the impact, they instead prompted the employee's Codex account to open a pull request in OpenAI's internal monorepo.

Hacktron says the entire sequence, from initial discovery to access to an OpenAI repository, took less than 72 hours. It says it reported the vulnerability immediately and worked with both companies on a patch, and that OpenAI paid a $6,500 bounty. Discourse has confirmed and patched the underlying image-upload flaw.

The incident points to a broader shift. AI agents are increasingly given credentials and connectors that reach across source code, communications and corporate data, while the same class of models is making exploit development faster and cheaper.

Ready for more?

Join our newsletter to receive updates on new articles and exclusive content.

We respect your privacy and will never share your information.

Enjoyed this article?

Yes
No
Follow Us:

Unmissable content


Loading comments...

Also of Interest