Cisco Research: Mainstream AI Models Are Far Less Defensive Against Multi-Round Malicious Prompts Than Vendors Claim
Cisco's latest research shows that 15 leading AI models developed by OpenAI, Anthropic, Google, Amazon, and xAI, among others, have significantly higher success rates under multi-round malicious prompt attacks than single-round attacks, with notable gaps in some models. Researchers urge vendors to reassess security testing methods and disclose more cross-scenario data.

Brief Overview
- Cisco researchers said in a report released Wednesday that security claims by major AI developers are based on incorrect assumptions about hacker behavior.
- AI vendors assume that as long as a model can resist a single malicious prompt, it is secure, but hackers increasingly use multi-stage prompts to bypass model defenses, and most models are not prepared for such attacks.
- The new report reveals a largely underestimated internal vulnerability in AI models that could expose companies using these tools to widespread disruption and damage.
Deep Dive
Cisco's evaluation of 15 leading AI models from OpenAI, Anthropic, Google, Amazon, and xAI found that "single-turn attack success rate (ASR) does not reliably reflect what happens when attackers adapt across turns," researchers Nicholas Conley and Amy Chang wrote. Their tests showed that AI models are far more vulnerable to multi-turn malicious prompts than to single-turn prompts—multi-turn attack success rates ranged from 8% to 88%, while single-turn rates were only 2% to 65%.
"Every model we tested exhibited a non-negligible multi-turn ASR," Conley and Chang wrote.
The two researchers previously co-authoreda November 2025 report, which found that open-weight AI models are two to ten times more vulnerable to multi-turn attacks than to single-turn attacks.
"The patterns we documented in open models also hold in closed models," they wrote in the new study. "In this cohort, no frontier closed model can be described as safe under iterative attacks. This is a statement about the current state of the closed-model frontier, not a critique of any single vendor."
One of the study's most notable findings is the correlation between AI companies' priorities and their models' security. Conley and Chang found that AI developers publicly emphasizing increasing model capabilities had the largest gap between their models' single-turn and multi-turn attack vulnerability; developers whose public statements emphasized model security had smaller gaps, indicating more coordinated efforts to reduce risk.
The researchers tested five strategies: role-playing, misleading the model, information decomposition, reframing model refusals, and gradual escalation. xAI's model Grok 4.1 Fast Non-Reasoning performed worst, with researchers succeeding in 88% of multi-turn attacks (compared to a 34% single-turn attack success rate on the same model). The best-performing model was Amazon's Nova 2 Lite, which failed to resist only 8% of multi-stage attacks, but the researchers noted that this figure "still represents meaningful residual risk."
Conley and Chang noted that enabling reasoning features significantly improved Grok 4.1's performance, suggesting that AI vendors should "document the security-related impact of configuration decisions such as reasoning state."
OpenAI, Anthropic, Google, Amazon, and xAI did not immediately respond to requests for comment.
The researchers said vendors need to rethink how they evaluate AI model security, and enterprises also need more information about the potential gap between models' single-turn and multi-turn attack resilience.
"Business decisions based on published single-turn scores carry security and governance risks," Conley and Chang wrote. "A model with a single-turn ASR of 2.74% is not the same product as a model whose multi-turn ASR remains at 24.68%. Without paired scenario data, the two are indistinguishable in most public evaluations, and end users never see this gap."