Researchers Used Claude to Hack OpenAI — and It Shows How AI Is Changing Cybersecurity
By Stacked AI · Published September 18, 2026
AI Security · Part 1
One of the most interesting cybersecurity stories of the week involves two of the world's most prominent AI companies.
Cybersecurity researchers from Hacktron AI used Anthropic's Claude during security research that uncovered vulnerabilities affecting OpenAI's systems. According to reporting from The Wall Street Journal and Financial Times, the researchers were ultimately able to demonstrate access to an OpenAI employee's ChatGPT account and reach part of OpenAI's internal software-development environment.
The researchers were participating in legitimate security research rather than conducting a malicious attack. They stopped before accessing sensitive source code, reported the vulnerabilities to OpenAI, and received a $6,500 bug bounty.
That distinction is important.
But from a cybersecurity perspective, the most interesting part is not simply that OpenAI had vulnerabilities.
Every large technology platform eventually encounters security bugs.
What makes this case important is how the vulnerabilities were discovered and exploited.
The researchers used a competing frontier AI model as part of the vulnerability-research process.
That offers a glimpse of what offensive and defensive cybersecurity may increasingly look like.
From vulnerability discovery to an actual attack chain
Security vulnerabilities rarely exist in isolation.
An individual bug may appear relatively limited:
Web application flaw ↓ Session or authentication exposure ↓ Access to another service ↓ Privilege escalation ↓ Internal resources
The difficult part of penetration testing is often identifying how several apparently unrelated weaknesses can be combined.
According to the reports, Hacktron's research began around OpenAI's community forum infrastructure, which used the third-party Discourse platform.
The researchers discovered a weakness that enabled them to obtain authentication material associated with an OpenAI employee. They then demonstrated that the resulting access could extend further into OpenAI's systems.
Business Insider separately reports that the researchers demonstrated access involving an employee ChatGPT account and Codex before stopping the exercise and notifying OpenAI.
OpenAI subsequently revoked affected authentication tokens and sessions, according to the reporting.
Where Claude enters the picture
Hacktron used Anthropic's Claude as part of its vulnerability research.
This is significant because modern AI coding systems can potentially assist with many parts of traditional penetration testing:
reading unfamiliar source code, explaining authentication logic, analyzing API behavior, generating proof-of-concept code, interpreting application errors, modifying exploit scripts, analyzing responses, and identifying relationships between different vulnerabilities.
None of those capabilities automatically creates a successful hacker.
Understanding scope, recognizing exploitable conditions, designing an attack chain, and validating the impact still require substantial judgment.
But AI can reduce the amount of manual work involved in each step.
That changes the economics of security testing.
Verified facts
The following points are supported by the reporting:
Hacktron AI researchers conducted the research. Anthropic's Claude was used during the work. Vulnerabilities involving OpenAI infrastructure were successfully demonstrated. An OpenAI employee account was accessed. The researchers demonstrated access involving OpenAI's internal development environment. The researchers stopped their testing rather than accessing sensitive source code. OpenAI received the disclosure and paid a $6,500 bounty. Analysis
The broader conclusion — that AI could significantly accelerate offensive cybersecurity — is an interpretation of these facts rather than something this single incident proves by itself.
However, it fits a wider pattern being documented by AI companies.
Anthropic's September threat-intelligence reporting says it has observed threat actors using AI not simply as a chatbot but as part of workflows that automate reconnaissance, tool development, exploitation, persistence, infrastructure management, and data processing.
That makes the Hacktron case especially useful.
It shows the same underlying capability being applied by defenders under responsible-disclosure rules.
AI is compressing the penetration-testing workflow
Traditional penetration testing contains several phases.
A simplified workflow might look like:
Reconnaissance ↓ Attack-surface mapping ↓ Vulnerability discovery ↓ Exploit development ↓ Validation ↓ Privilege escalation ↓ Reporting
Historically, each stage required considerable manual effort.
A penetration tester might spend hours reading documentation, experimenting with request parameters, writing scripts, reviewing JavaScript bundles, examining authentication flows, and modifying payloads.
AI agents can increasingly assist with these repetitive activities.
Consider source-code analysis.
A security engineer can provide an AI coding agent with a repository and ask it to trace where user-controlled input enters an application.
The agent can search the codebase, identify data flows, inspect sanitization, and highlight places where validation appears inconsistent.
That does not mean its conclusions are automatically correct.
But it dramatically reduces the amount of code a human needs to inspect manually.
The same applies to API testing.
Instead of manually writing every request variation, an agent can help generate test cases and analyze responses.
For exploit development, it can explain the underlying vulnerability and help create a proof of concept.
Then it can iterate when the first attempt fails.
The important improvement is not necessarily that AI knows some secret hacking technique.
Much of cybersecurity knowledge is already publicly documented.
The advantage is iteration speed.
An AI system can continuously:
inspect → generate → execute → observe → modify → retry
That loop is exactly where agentic AI becomes different from an ordinary chatbot.
The attacker needs fewer repetitive skills
One implication is that the skill distribution in cybersecurity may change.
Previously, an attacker might need strong knowledge of several areas simultaneously:
web development, authentication, scripting, networking, operating systems, exploit development, cloud infrastructure, and API behavior.
AI does not remove the need for security expertise.
But it may allow someone who is strong in one part of the attack chain to compensate more effectively for gaps elsewhere.
For defenders, this matters because security assumptions often depend implicitly on attacker cost.
A vulnerability that is difficult to discover may historically have attracted fewer attackers.
A vulnerability requiring a custom exploit may have been less likely to be weaponized quickly.
AI can reduce both costs.
That means organizations may increasingly need to assume that vulnerabilities will move from disclosure to practical exploitation faster.
The defensive side may benefit even more
The same capabilities can be used defensively.
This is where the Hacktron incident is particularly instructive.
The researchers used AI in a controlled security context, found real weaknesses, reported them, and allowed the vendor to remediate the problems.
Organizations could apply similar workflows internally.
For example, imagine an AI-assisted DAST system operating against a staging application.
The system first maps the application's attack surface.
/login /api/users /api/orders/{id} /admin /upload
It then creates security hypotheses.
Could /api/orders/{id} contain IDOR?
Could /upload allow unsafe files?
Does /admin correctly enforce authorization?
Can login error messages leak account existence?
Instead of firing thousands of generic payloads, an AI-assisted system could use information about the application to determine which tests are most meaningful.
SAST tools could work similarly.
A traditional scanner may identify a potentially dangerous function.
An AI layer could inspect the surrounding code, trace whether the input is actually attacker-controlled, and determine whether the finding is realistically exploitable.
This could help address one of the oldest problems in application security:
false positives.
Security teams often receive far more findings than they have time to investigate.
A useful AI security system would not simply produce even more alerts.
It would prioritize:
Vulnerability + Reachability + Exploitability + Business impact + Evidence
That is substantially more valuable than a long list of CVEs.
This also creates new security risks
Giving an AI security agent additional autonomy introduces obvious danger.
A scanning system capable of sending HTTP requests is relatively constrained.
An autonomous security agent capable of:
running commands, modifying exploits, using credentials, navigating cloud infrastructure, executing browsers, and chaining vulnerabilities
has a very different risk profile.
Security architecture therefore becomes critical.
Organizations experimenting with agentic penetration testing should define:
Scope Exactly which systems may the agent test?
Authentication Which credentials may it use?
Network boundaries Which hosts can it communicate with?
Execution permissions Can it execute commands or only generate recommendations?
Destructive actions Can it delete, modify, or upload data?
Logging Can investigators reconstruct every action the agent performed?
Human approval Which high-risk actions require manual confirmation?
Without those controls, an AI security agent designed for defensive testing could itself become a security incident.
The important lesson is not "AI can hack"
That headline is tempting, but incomplete.
Security researchers have automated attacks for decades.
Metasploit, Burp Suite, Nmap, vulnerability scanners, fuzzers, exploit frameworks, and custom scripts already automate significant parts of offensive security.
AI adds something different.
It makes automation more adaptive.
Traditional security automation generally follows rules written in advance.
An AI agent can examine an unexpected response, reason about what happened, change its strategy, and try again.
That feedback loop is what makes current developments important.
The Hacktron/OpenAI incident demonstrates the defensive version of that possibility.
Researchers used AI to help identify and validate genuine security problems in one of the world's most technically sophisticated AI companies.
The natural conclusion for security teams is not that human penetration testers are suddenly obsolete.
It is that penetration testers who learn to supervise AI agents may be able to investigate much larger attack surfaces.
At the same time, defenders should assume attackers will gain access to similar capabilities.
The race therefore becomes:
Can defenders use AI to find and fix vulnerabilities faster than attackers can discover and exploit them?
That is likely to become one of the defining cybersecurity questions of the next several years.
Takeaway
Today's OpenAI incident is an unusually concrete example of AI and cybersecurity converging.
Hacktron AI researchers used Claude while identifying and exploiting vulnerabilities that allowed them to demonstrate meaningful access inside OpenAI's environment. They stopped the testing, responsibly disclosed the issues, and received a bug bounty.
The bigger lesson is about workflow.
AI is moving from simply explaining vulnerabilities toward helping security practitioners investigate, test, iterate, and validate them.
For security teams building SAST, DAST, pentesting, or vulnerability-management systems, this suggests a useful direction:
Don't use AI merely to generate prettier vulnerability reports.
Use it to connect attack-surface discovery → hypothesis generation → controlled testing → evidence → prioritization → remediation.
That is where agentic AI could have the greatest impact on application security.