
A new Anthropic report has raised fresh concerns about how advanced AI tools could be exploited to support weapons development. The company said a Yemen-based guided-weapons engineering cell used its Claude models over several months to assist with software linked to missile and rocket programmes. Anthropic did not explicitly identify the group as the Houthis, although the cell was based in northern Yemen, an area controlled by the Iran-aligned movement. Subsequent reporting has described the group as Houthi-linked.
The disclosure appears in Anthropic’s September 2026 threat-intelligence report, which examines malicious or high-risk uses of Claude between December 2025 and August 2026.
What did Anthropic say about the Yemen-based group?
Anthropic said the group was involved in three weapons-development programmes.
The first involved a guided rocket. The second concerned a multi-stage ballistic missile with a stated range goal of more than 2,000 kilometres. A third programme, referred to as the “R2000” set, involved several proposed missile variants, including one featuring a hypersonic-glide vehicle.
The company said the actors used Claude and Claude Code as part of the engineering process rather than relying entirely on conventional human software development.
How was Claude used?
According to Anthropic, Claude Code was used to assist with parts of the guidance, navigation and control software for flying vehicles.
At a high level, the AI was used for software development, research, code review, simulation and troubleshooting. Multiple Claude instances were reportedly assigned different tasks, while human operators remained responsible for directing the overall project.
Anthropic said the users also attempted to divide their work across separate conversations, making the broader purpose of the programme harder for its safety systems to detect.
The report says Anthropic blocked many weapons-related requests, but the users found ways around some of those safeguards.
Did the group actually test a guided rocket?
This is one of the most striking claims in the report.
Anthropic said it had no evidence that the group successfully deployed an operational weapon. However, the company said the actors did carry out a field test of a guided rocket.
The test appeared to fail. According to Anthropic, the operators returned to Claude shortly afterward to investigate the failure and work through what had gone wrong.
That distinction is important. The report describes AI-assisted development and a failed test, not proof that Claude helped produce a functioning missile system that was subsequently deployed in combat.
Did Anthropic say the Houthis were behind it?
Not directly.
Anthropic’s own description refers to a “Yemen-based guided weapons engineering cell” in northern Yemen. It does not publicly name the Houthis as the organization responsible.
The geographic context is significant because much of northern Yemen is controlled by the Houthi movement. Recent reporting has therefore described the cell as Houthi-linked, but that attribution goes beyond the wording of Anthropic’s report itself.
That makes the distinction between “Houthi use of Claude” and “a weapons cell operating in Houthi-controlled territory” important in any headline or report.
Why is the case significant?
The larger concern is not that AI suddenly invented a new class of weapon.
Instead, the case illustrates how increasingly capable AI systems can act as force multipliers for people who already possess technical knowledge, hardware, software and access to weapons programmes.
Anthropic said its September report identified a broader category of misuse involving conventional weapons, including missiles, armed drones, bombs and related systems. The company documented cases spanning Yemen, China and Russia.
The report also highlights a shift from simple question-and-answer use toward more autonomous, multi-step activity. In other cases examined by Anthropic, AI systems were used to coordinate complex cyber operations with humans acting more as supervisors than hands-on operators.
Could AI make weapons development easier?
Potentially, yes, especially for specific technical tasks.
AI can help users write and review software, troubleshoot errors, organize technical information and accelerate repetitive engineering work. That does not eliminate the need for hardware, testing facilities, specialist knowledge or human decision-making, but it can reduce the time required for some stages of a complex project.
That is precisely why AI safety researchers are increasingly concerned about dual-use capabilities. A system built to make programming and research more efficient can also be misused when the underlying project is dangerous.
What happened after Anthropic detected the activity?
Anthropic said it investigated the activity, banned the associated accounts and shared relevant information with government and industry partners.
The company presented the Yemen case as one of several examples showing both the effectiveness and limitations of current AI safeguards.
Anthropic’s wider report covers seven categories of harmful activity, including cyber operations, influence campaigns, surveillance, scams and fraud, biological misuse, conventional weapons development and attempts to extract or replicate AI capabilities.
What the report does not prove
The disclosure should not be interpreted as evidence that Claude autonomously designed, built or deployed a working missile.
Anthropic’s account indicates that humans remained involved and that the AI was used as a technical assistant across parts of the development process. The company also explicitly said it had no evidence that the actors successfully fielded an operational weapon.
The reported failed rocket test is nevertheless significant because it suggests the AI was being used not just for theoretical research, but as part of an ongoing real-world engineering effort.
The bigger AI security question
The Yemen case adds another piece to a rapidly changing debate around AI safety.
As models become better at coding, technical research and multi-step problem-solving, the central question is increasingly about who can access those capabilities and how effectively dangerous uses can be detected.
Anthropic said the cases in its report were not typical examples of everyday Claude use, but rather among the most sophisticated and novel misuse it identified. The company argued that documenting these incidents can help AI developers, governments and security researchers recognize similar activity before it becomes more consequential.