Aug 4, 2026
Policy

Bypassing AI guardrails was easy in Talos threat-actor review

Cisco Talos found simple reframing often won model cooperation, but its evidence says attacker skill still determined what AI use achieved.

Renata Fuchs

By Renata Fuchs · Policy Reporter

· 3 min read

Bypassing AI guardrails was easy in Talos threat-actor review
Photo: The Register

Cisco Talos found that bypassing AI guardrails often required little more than reframing a request, based on prompt logs and related artifacts recovered from threat-actor endpoints. The August 4 findings are a warning for companies deploying AI in security-sensitive workflows, but they do not show that a low-skill user can readily turn model cooperation into an effective cyber operation.

Talos examined a “significant corpus” of files left behind on endpoints using products including Claude Code, Codex, Cursor and Gemini. The evidence came from actors who exposed those artifacts through operational-security mistakes, so it does not establish how widespread any one technique is across all attackers, models or platforms.

The researchers said they saw actors use AI for malicious software development, scaling criminal activity and vulnerability research. In the material Talos reviewed, models were frequently persuaded by asserted authorization or ownership claims, without the researchers observing sophisticated encoding or elaborate evasion. Actors also presented work as a capture-the-flag or bug-bounty exercise, split tasks across separate sessions or files, and used language that concealed the overall intent of a larger activity.

How easy is bypassing AI guardrails?

Talos’s answer is narrow but troubling: obtaining cooperation on some requests could be low-friction. A guardrail is a model or product control intended to refuse or limit harmful requests. Its effectiveness can break down when a system has too little context to distinguish legitimate security work from misuse, or accepts an unsupported claim that the work is authorized.

That is different from measuring cyber impact. Talos said unsophisticated actors could assemble malicious projects that technically worked, but their results tended to have limited functionality and little capacity for improvement. More experienced operators, by contrast, used AI as a force multiplier and “pushed the bounds” of what the researchers expected possible.

Axios, which also reported on the Talos review, said Cisco found novice-to-intermediate users struggled to move beyond producing tools needed for an attack. That distinction is the useful one for security teams: refusal rates and jailbreak demonstrations do not, on their own, measure the real-world capability of the person using the model.

What should security teams do?

Talos’s recommendation is not to treat a model’s own safety restrictions as a standalone defense. The group expects AI use by threat actors to increase the volume of vulnerabilities, alerts and incidents that security operations centers must handle. It said organizations should prepare their security operations accordingly and assess agentic capabilities that can help analysts focus on actionable alerts.

For AI vendors and enterprise buyers, the report also points to a product-design problem: systems serving legitimate red-teamers and vulnerability researchers can be manipulated when they accept context supplied by an unverified user. Talos did not identify a single affected model as the cause; it said the patterns it observed were not confined to one platform.

This story draws on original reporting from The Register.

More from Policy

All Policy →