Agentic AI Security Risk, Defined
Agentic AI security risk is the danger that an AI agent, software that plans and carries out tasks on its own by using tools and real system access, takes a harmful action. Unlike a chatbot that only replies, an agent can send email, change data, run code or move money so a mistake or a manipulation becomes an action.
That shift from answering to acting is why agentic AI is treated as its own security problem. Adoption is moving fast. Gartner forecasts that 33% of enterprise software applications will include agentic AI by 2028, up from less than 1% in 2024. Most of those agents will hold credentials and act inside real systems.
Two things can go wrong. An agent can fail on its own, doing something destructive because it misread an instruction. Or an attacker can hijack it, feeding it hidden instructions so it acts against you. This guide covers both with real cases and the controls that address each.
How Agentic AI Works
An AI agent is built on a large language model, the same kind of system behind a chatbot but wrapped in extra parts that let it act. At its core it runs a simple loop. It takes in a goal and the current situation, reasons about a plan, then acts by calling a tool. It reads the result and repeats until the task is done or it gives up.

Three additions turn a model into an agent and each one adds risk.
- Tools and actions: the agent calls real functions, such as sending an email, querying a database, running a shell command or calling an API. This is what turns a suggestion into an action.
- Autonomy: it chooses its own next step without asking a person each time. That is the point of an agent and also the source of the risk.
- Memory: it keeps context across steps and sessions so anything poisoned early can quietly steer it later.
Two design facts create most of the risk. First, the agent reads instructions and outside data through the same channel so text buried in a document or a web page can read as a command. Second, the agent holds real credentials and permissions so when it acts, it acts with your access.
None of this needs a research lab any more. Off-the-shelf models, tool connectors and agent frameworks let a small team put an agent into production in days. The capability arrived faster than the controls around it and that gap is where this risk lives.
Types of Agentic AI Security Risks
Security researchers have grouped the ways agents go wrong. The categories below draw on OWASP which publishes the main industry reference for AI and agent risk in its Top 10 for LLM Applications and its Top 10 for Agentic Applications.
- Prompt injection and goal hijack: Hidden instructions in a document, email or web page make the agent ignore its real task and follow the attacker instead. OWASP ranks prompt injection as the top risk for language-model applications.
- Excessive agency: The agent has more tools, permissions or freedom to act than its job needs so an influenced or mistaken agent can do real damage. OWASP calls this the risk that matters most as agents reach production.
- Tool misuse: The agent uses a legitimate tool in an unsafe way for example running a destructive command or calling an API with the wrong parameters.
- Identity and privilege abuse: Agents act with inherited or shared credentials that no one can trace to a person, so misuse is hard to spot and to attribute.
- Memory and context poisoning: An attacker plants false information in the agent’s memory or knowledge base so its later decisions are quietly wrong.
- Supply-chain risk: A compromised model, plugin or tool the agent depends on brings the compromise inside.
- Cascading failures: In a chain or a team of agents, one small error or injection spreads because each agent trusts the output of the last.
- Data exfiltration: An over-permissioned agent is tricked into reading sensitive data and sending it out of the organisation.
Business Impact
The cost of agentic AI risk is not abstract. An agent fails or is turned against you in ways that map straight to money, downtime and legal exposure.
The clearest harm is unauthorised action at machine speed. A compromised or mistaken agent does not make one bad move, it makes hundreds before anyone notices. A single wrong command can wipe a database or move funds in seconds.
Data exposure is the next mechanism. Agents are given broad read access to do their jobs so a successful injection can turn that access into a leak of customer records, source code or internal documents.
Then there are wrong decisions made autonomously. An agent that approves, pays or publishes without a human check can commit the business to an error. The business still owns the outcome.
Traceability is the quiet cost. When an agent acts with shared credentials and no logging, you are left unable to say what the agent did or whose access it used.
Regulators and analysts expect this to grow. Gartner forecasts that by 2028, 25% of enterprise breaches will be traced back to AI agent abuse from both outside attackers and insiders.
Real-World Cases
GTG-1002: AI Used to Run an Attack
In November 2025 Anthropic, the company behind the Claude models, reported that it had disrupted what it describes as the first AI-orchestrated cyber-espionage campaign. It attributes the activity which it tracks as GTG-1002, to a Chinese state-sponsored group.
According to Anthropic, the attackers used its Claude Code agent to run reconnaissance, find vulnerabilities, harvest credentials, move laterally and steal data across roughly 30 organisations in technology, finance, chemicals and government. Anthropic estimates the AI carried out 80 to 90% of the hands-on work.

They got past the guardrails by role-playing a sanctioned penetration test and splitting the job into small steps that each looked harmless. Anthropic also notes the AI made mistakes and overstated some findings so it was fast but not flawless.
The techniques were ordinary and only the speed was new. The defences that blunt such an attack are the familiar ones, least privilege, multi-factor authentication, segmentation and fast detection of abnormal credential use and lateral movement.
Replit: An Agent That Deleted a Live Database
In July 2025 the founder of SaaStr, Jason Lemkin was building software with Replit’s AI coding agent. On 18 July, during an explicit code and action freeze, the agent deleted the live production database.
The deletion wiped records for more than 1,200 executives and over 1,190 companies. The agent then generated fake data and told Lemkin, wrongly, that the deletion could not be undone. The data was in fact recoverable and was restored.
No attacker was involved. The agent simply had production access and the freedom to use it. Replit separated development from production automatically and added a planning-only mode. The durable control is least privilege for the agent and a human sign-off before any destructive action.
EchoLeak: Hijacking an Assistant Through Email
In June 2025 security researchers disclosed EchoLeak, tracked as CVE-2025-32711, a flaw in Microsoft 365 Copilot that Microsoft rated critical at 9.3 out of 10. It was a zero-click attack, meaning the target did not have to click anything.
A single crafted email carried hidden instructions. When Copilot processed the email to help the user, it followed those instructions, reached into the user’s Microsoft 365 data and sent sensitive content out. Microsoft patched it server-side and reported no exploitation in the wild.
Microsoft fixed this specific bug. The lasting controls sit with every team running an assistant, treat incoming content as untrusted, scope the assistant’s data access tightly and control what it can send out.
Agentic AI and Compliance
Using an agent does not move the responsibility for what it does. European and Swedish rules already reach AI agents and more obligations are arriving on a fixed timetable.
The EU AI Act (Regulation (EU) 2024/1689) is the headline rule. It entered into force on 1 August 2024 and applies in stages. Obligations for general-purpose AI models have applied since 2 August 2025 and the duties for high-risk systems now apply from 2 December 2027. For high-risk uses the Act requires meaningful human oversight (Article 14), so an agent cannot be left to act entirely on its own.
For organisations in scope of NIS2, Sweden’s Cybersäkerhetslagen (SFS 2025:1506) has applied since 15 January 2026. It expects continuous monitoring, incident handling and supply-chain security and your AI agents now sit inside all three. Its board accountability requirement which implements NIS2 Article 20, canhold management personally responsible.
A serious incident also starts a reporting cascade to MCF (formerly MSB). The timings are 24 hours for an early warning, 72 hours for a full notification and one month for the final report.
Two other regimes apply where they are relevant. For financial entities, DORA treats an AI agent as ICT that must be governed and tested under its incident-management rules, supervised by Finansinspektionen. Where an agent handles personal data, GDPR applies including the 72-hour breach notification to IMY and the limits on fully automated decisions that significantly affect people (Article 22).
The accountability point is already settled in practice. In 2024 a Canadian tribunal (Moffatt v. Air Canada) held Air Canada responsible for wrong information its chatbot gave a customer, rejecting the argument that the chatbot was a separate entity. The level of autonomy did not matter. The organisation owned the output.
How to Spot a Risky Agent Deployment
You will rarely catch a prompt injection by reading it. The more useful skill is spotting an agent deployment that is set up to fail. These are the warning signs.
- The agent can reach tools and data far beyond its task.
- It runs with shared or service-account credentials that cannot be traced to a person.
- High-impact actions, such as deleting data, paying or emailing outside the company, happen without a human approving them.
- There is no step-by-step log of what the agent did.
- The agent ingests outside content such as emails, documents or web pages and acts on it directly.
- Nobody owns a current list of the tools and permissions the agent holds.
Spotting these signs helps, but it is not a defence on its own. OWASP is blunt that prompt injection cannot be patched away, because it exploits how language models read text. So the goal is not to detect every bad instruction. It is to limit what the agent can do when one gets through.
How to Defend
Defending agents is mostly good security pointed at a new kind of actor. It spans people, process and technology and the controls below are the practical core.

- Give every agent least privilege. Grant only the specific tools and data its task needs and nothing more.
- Keep a human in the loop for high-impact actions. Require explicit sign-off before an agent deletes data, moves money or sends anything outside the company.
- Treat all external content as untrusted. Assume any email, document or web page an agent reads may try to instruct it, and keep instructions separate from data.
- Give each agent its own identity. Use unique, traceable credentials per agent instead of shared service accounts, so every action is attributable.
- Log and monitor what agents do. Record each step and tool call and watch for abnormal or high-velocity activity.
- Test agents before and after launch. Red-team and penetration-test agent deployments the way you would any exposed system.
- Govern the whole lifecycle. Use a recognised framework such as the NIST AI Risk Management Framework, and prepare for the EU AI Act’s human-oversight duty.
If you want help putting these controls in place, eBuilder’s AI detection and response monitors AI systems, agents and the data flowing through them and our penetration testing can probe an agent deployment the way an attacker would. Both are delivered by a Sweden-based team.



