AI safety and research company Anthropic has developed a frontier AI model, internally named Claude Mythos Preview, that represents a paradigm shift in offensive cybersecurity capabilities. According to reports, the Mythos model can autonomously discover novel, zero-day vulnerabilities in complex software, generate functional exploit code for them, and chain them together to execute sophisticated attacks with minimal human intervention. Due to these powerful dual-use capabilities, Anthropic has made the decision not to release the model publicly, deeming the risk of misuse to be too high. Instead, it is engaging with a small number of trusted partners for defensive research under "Project Glasswing." The situation is further complicated by reports that Anthropic is investigating a potential unauthorized access incident, raising alarms about the containment and governance of such powerful AI systems.
The emergence of Mythos marks a fundamental change in the cyber threat landscape. It collapses the timeline between vulnerability discovery and weaponization from months or years to potentially minutes. An AI that can find and exploit zero-days on its own creates several new classes of threats:
While Anthropic is acting responsibly by restricting access, the report of a potential leak via a third-party contractor highlights the immense challenge of securing these models. The proliferation of this technology, whether through leaks, independent replication by other actors, or state-level development, is now a primary concern for global cybersecurity.
The capabilities of Mythos likely stem from a combination of Large Language Models (LLMs) and advanced reinforcement learning techniques. The model was probably trained on a massive corpus of open-source code, security advisories, vulnerability databases, and exploit code from sources like GitHub and Exploit-DB.
T1595 - Active ScanningT1647 - Develop Capabilities: ExploitsT1190 - Exploit Public-Facing ApplicationThe strategic impact of autonomous hacking AI is profound:
There are no IOCs for this conceptual threat.
Hunting for an AI attacker is a new frontier. It would involve looking for activity that is too fast, too complex, or too efficient to be human.
Traditional signature-based and even heuristic-based detection will likely fail.
Mitigating this threat requires a multi-layered, strategic approach.
D3FEND Techniques:
D3-DA: Dynamic Analysis and D3-SA: Static Analysis will need to be performed by defensive AI agents continuously and at scale.UK government and Ofcom issue formal alert to businesses regarding 'catastrophic' cyber threats from advanced AI models like Anthropic's Mythos.
Japan forms public-private task force to counter AI-driven cyber threats like Mythos, focusing on the financial sector.
OpenAI's GPT-5.5-Cyber joins Anthropic's Mythos, driving a surge in AI-powered vulnerability discovery and patching by tech companies, escalating the AI-driven threat.
Automated, rapid patching becomes even more critical when vulnerabilities can be discovered and exploited at machine speed.
Mapped D3FEND Techniques:
Using strong sandboxing to limit the impact of a potential exploit, even if the vulnerability itself is unknown.
Mapped D3FEND Techniques:
The only viable long-term defense against an AI-powered attacker like Mythos is a defense powered by a similar class of AI. Organizations must begin integrating AI-assisted security tools across their stack. This means deploying defensive AI agents that can perform tasks autonomously and at machine speed. For example, a defensive AI could continuously scan an organization's source code and production systems for vulnerabilities, using the same techniques as Mythos. When it finds a flaw, it could automatically generate a patch, test it in a sandboxed environment, and deploy it—all without human intervention. Similarly, an AI-powered SIEM could analyze trillions of log events in real-time to detect the subtle, complex patterns of an autonomous attacker that would be invisible to human analysts. The paradigm must shift from human-in-the-loop to human-on-the-loop, where humans set the strategy and the AI agents execute the defense.
In a world with autonomous hacking AIs, high-interaction decoy environments (honeypots) become invaluable research tools. Organizations should deploy sophisticated honeypots that mimic their production environments, complete with fake data, users, and applications containing deliberately seeded vulnerabilities. The goal is to lure an AI attacker into this controlled environment. By observing how the AI discovers and exploits the decoy systems, defenders can learn its TTPs, its decision-making logic, and the types of vulnerabilities it prioritizes. This intelligence is critical for building effective defenses. The logs and artifacts from the decoy environment can be used to train defensive AI models, creating a feedback loop where the attacker's own actions are used to improve the defender's capabilities. This is one of the few ways to study such a novel threat without risking production systems.
Anthropic confirms it is investigating reports of unauthorized access to the Mythos model.

Cybersecurity professional with over 10 years of specialized experience in security operations, threat intelligence, incident response, and security automation. Expertise spans SOAR/XSOAR orchestration, threat intelligence platforms, SIEM/UEBA analytics, and building cyber fusion centers. Background includes technical enablement, solution architecture for enterprise and government clients, and implementing security automation workflows across IR, TIP, and SOC use cases.
CyberNetSec.io uses automation to assist source monitoring, deduplication, observable extraction, and structured intelligence generation. Published analysis follows human-defined editorial standards and adds defensive context including MITRE ATT&CK, D3FEND, STIX, and Sigma where applicable. Read our editorial policy.
Help others stay informed about cybersecurity threats
Every tactic, technique, and sub-technique used in this threat has been identified and mapped to the MITRE ATT&CK framework for consistent, actionable threat language.
Observables and indicators of compromise (IOCs) have been extracted and cataloged. Risk has been assessed and correlated with known threat actors and historical campaigns.
Detection rules, incident response steps, and D3FEND-aligned mitigation strategies are included so your team can act on this intelligence immediately.
Structured threat data is packaged as a STIX 2.1 bundle and can be visualized as an interactive graph — relationships between actors, malware, techniques, and indicators.
Sigma detection rules are derived from the threat techniques in this article and can be converted for deployment across any major SIEM or EDR platform.