CYBERSECURITYTRACKER
TRACKING6,506 stories in this site build1,309 vulnerability news stories in this site build
Permanent story citation

Top AI Agents Built to Catch Malicious Code Can Be Tricked Into Running It

This page keeps the story as Cybersecurity Tracker first published it. If the tracker later corrects it, the correction appears below the original and never replaces it.

Back to newsStory 2211

As cited

Copy frozen at (site build).

ai security

Top AI Agents Built to Catch Malicious Code Can Be Tricked Into Running It

Researchers at the AI Now Institute demonstrated a proof-of-concept attack called Friendly Fire that tricks AI coding agents, including Anthropic's Claude Code and OpenAI's Codex, into executing malicious code instead of analyzing it for security vulnerabilities. The attack works when these agents operate in autonomous modes with self-approval capabilities, creating a scenario where defensive tools become attack vectors.

Why it matters: Development teams and security practitioners using AI agents for code review and vulnerability scanning face the risk of inadvertently executing attacker-controlled code on their systems; organizations should restrict agent autonomy and require human approval before code execution.

Source published
First seen by Cybersecurity Tracker

Source attribution

Correction

Correction recorded as of .

ai security

Top AI Agents Built to Catch Malicious Code Can Be Tricked Into Running It

No summary had been written when this copy was frozen.

First seen by Cybersecurity Tracker

Source attribution

Correction

Correction recorded as of .

ai security

Top AI Agents Built to Catch Malicious Code Can Be Tricked Into Running It

Researchers from the artificial intelligence (AI) Now Institute demonstrated a proof-of-concept attack called "Friendly Fire" showing that AI coding agents such as Anthropic's Claude Code and OpenAI's Codex can be manipulated into executing malicious code when operating in autonomous mode. The attack tricks security-focused agents into running attacker-controlled code on the target machine instead of analyzing it safely.

Why it matters: Organizations and developers using autonomous AI coding agents for security scanning face the risk of inadvertently executing malicious payloads if the agent approval mechanisms lack sufficient safeguards; teams should review their autonomous AI tool configurations and consider manual approval workflows for untrusted code sources.

Source published
First seen by Cybersecurity Tracker

Source attribution

Glossary