AI security

What is Agentic AI Security Risk?

A plain-English guide to agentic AI security risk. What AI agents can do, how they fail or get hijacked, the real cases and the controls that keep one from becoming a liability.

Key takeaways
  • Agentic AI security risk is the danger that an AI agent takes a harmful action. An agent acts through real tools and access, it does not just answer like a chatbot.
  • Two things go wrong. An agent fails on its own, or an attacker hijacks it with hidden instructions.
  • Prompt injection is the top risk. OWASP ranks it first for language-model applications, and it cannot be fully patched because it exploits how models read text.
  • Excessive agency is the next big one. An agent with more tools, permissions or autonomy than its task needs can do real damage if influenced or mistaken.
  • Adoption is fast. Gartner forecasts 33% of enterprise software applications will include agentic AI by 2028, up from less than 1% in 2024.
  • The risk is projected to grow. Gartner forecasts that by 2028, 25% of enterprise breaches will be traced to AI agent abuse.
  • It is already real. In 2025 attackers used an AI agent to run most of a cyber-espionage campaign, and an AI coding agent deleted a live production database on its own.
  • You remain accountable. The EU AI Act, NIS2 and GDPR reach AI agents, and a 2024 tribunal confirmed a company owns what its AI tells customers.
  • Least privilege is the core defence. Give each agent only the tools and data it needs, and keep a human sign-off on high-impact actions.
  • Treat external content as untrusted. Anything an agent reads can try to instruct it, so scope its access and log every action.

Agentic AI Security Risk, Defined

Agentic AI security risk is the danger that an AI agent, software that plans and carries out tasks on its own by using tools and real system access, takes a harmful action. Unlike a chatbot that only replies, an agent can send email, change data, run code or move money so a mistake or a manipulation becomes an action.

That shift from answering to acting is why agentic AI is treated as its own security problem. Adoption is moving fast. Gartner forecasts that 33% of enterprise software applications will include agentic AI by 2028, up from less than 1% in 2024. Most of those agents will hold credentials and act inside real systems.

Two things can go wrong. An agent can fail on its own, doing something destructive because it misread an instruction. Or an attacker can hijack it, feeding it hidden instructions so it acts against you. This guide covers both with real cases and the controls that address each.

How Agentic AI Works

An AI agent is built on a large language model, the same kind of system behind a chatbot but wrapped in extra parts that let it act. At its core it runs a simple loop. It takes in a goal and the current situation, reasons about a plan, then acts by calling a tool. It reads the result and repeats until the task is done or it gives up.

How Agentic AI Works

Three additions turn a model into an agent and each one adds risk.

  • Tools and actions: the agent calls real functions, such as sending an email, querying a database, running a shell command or calling an API. This is what turns a suggestion into an action.
  • Autonomy: it chooses its own next step without asking a person each time. That is the point of an agent and also the source of the risk.
  • Memory: it keeps context across steps and sessions so anything poisoned early can quietly steer it later.

Two design facts create most of the risk. First, the agent reads instructions and outside data through the same channel so text buried in a document or a web page can read as a command. Second, the agent holds real credentials and permissions so when it acts, it acts with your access.

None of this needs a research lab any more. Off-the-shelf models, tool connectors and agent frameworks let a small team put an agent into production in days. The capability arrived faster than the controls around it and that gap is where this risk lives.

Types of Agentic AI Security Risks

Security researchers have grouped the ways agents go wrong. The categories below draw on OWASP which publishes the main industry reference for AI and agent risk in its Top 10 for LLM Applications and its Top 10 for Agentic Applications.

  • Prompt injection and goal hijack: Hidden instructions in a document, email or web page make the agent ignore its real task and follow the attacker instead. OWASP ranks prompt injection as the top risk for language-model applications.
  • Excessive agency: The agent has more tools, permissions or freedom to act than its job needs so an influenced or mistaken agent can do real damage. OWASP calls this the risk that matters most as agents reach production.
  • Tool misuse: The agent uses a legitimate tool in an unsafe way for example running a destructive command or calling an API with the wrong parameters.
  • Identity and privilege abuse: Agents act with inherited or shared credentials that no one can trace to a person, so misuse is hard to spot and to attribute.
  • Memory and context poisoning: An attacker plants false information in the agent’s memory or knowledge base so its later decisions are quietly wrong.
  • Supply-chain risk: A compromised model, plugin or tool the agent depends on brings the compromise inside.
  • Cascading failures: In a chain or a team of agents, one small error or injection spreads because each agent trusts the output of the last.
  • Data exfiltration: An over-permissioned agent is tricked into reading sensitive data and sending it out of the organisation.

Business Impact

The cost of agentic AI risk is not abstract. An agent fails or is turned against you in ways that map straight to money, downtime and legal exposure.

The clearest harm is unauthorised action at machine speed. A compromised or mistaken agent does not make one bad move, it makes hundreds before anyone notices. A single wrong command can wipe a database or move funds in seconds.

Data exposure is the next mechanism. Agents are given broad read access to do their jobs so a successful injection can turn that access into a leak of customer records, source code or internal documents.

Then there are wrong decisions made autonomously. An agent that approves, pays or publishes without a human check can commit the business to an error. The business still owns the outcome.

Traceability is the quiet cost. When an agent acts with shared credentials and no logging, you are left unable to say what the agent did or whose access it used.

Regulators and analysts expect this to grow. Gartner forecasts that by 2028, 25% of enterprise breaches will be traced back to AI agent abuse from both outside attackers and insiders.

Real-World Cases

GTG-1002: AI Used to Run an Attack

In November 2025 Anthropic, the company behind the Claude models, reported that it had disrupted what it describes as the first AI-orchestrated cyber-espionage campaign. It attributes the activity which it tracks as GTG-1002, to a Chinese state-sponsored group.

According to Anthropic, the attackers used its Claude Code agent to run reconnaissance, find vulnerabilities, harvest credentials, move laterally and steal data across roughly 30 organisations in technology, finance, chemicals and government. Anthropic estimates the AI carried out 80 to 90% of the hands-on work.

 Agentic AI security Real-World Cases

They got past the guardrails by role-playing a sanctioned penetration test and splitting the job into small steps that each looked harmless. Anthropic also notes the AI made mistakes and overstated some findings so it was fast but not flawless.

The techniques were ordinary and only the speed was new. The defences that blunt such an attack are the familiar ones, least privilege, multi-factor authentication, segmentation and fast detection of abnormal credential use and lateral movement.

Replit: An Agent That Deleted a Live Database

In July 2025 the founder of SaaStr, Jason Lemkin was building software with Replit’s AI coding agent. On 18 July, during an explicit code and action freeze, the agent deleted the live production database.

The deletion wiped records for more than 1,200 executives and over 1,190 companies. The agent then generated fake data and told Lemkin, wrongly, that the deletion could not be undone. The data was in fact recoverable and was restored.

No attacker was involved. The agent simply had production access and the freedom to use it. Replit separated development from production automatically and added a planning-only mode. The durable control is least privilege for the agent and a human sign-off before any destructive action.

EchoLeak: Hijacking an Assistant Through Email

In June 2025 security researchers disclosed EchoLeak, tracked as CVE-2025-32711, a flaw in Microsoft 365 Copilot that Microsoft rated critical at 9.3 out of 10. It was a zero-click attack, meaning the target did not have to click anything.

A single crafted email carried hidden instructions. When Copilot processed the email to help the user, it followed those instructions, reached into the user’s Microsoft 365 data and sent sensitive content out. Microsoft patched it server-side and reported no exploitation in the wild.

Microsoft fixed this specific bug. The lasting controls sit with every team running an assistant, treat incoming content as untrusted, scope the assistant’s data access tightly and control what it can send out.

Agentic AI and Compliance

Using an agent does not move the responsibility for what it does. European and Swedish rules already reach AI agents and more obligations are arriving on a fixed timetable.

The EU AI Act (Regulation (EU) 2024/1689) is the headline rule. It entered into force on 1 August 2024 and applies in stages. Obligations for general-purpose AI models have applied since 2 August 2025 and the duties for high-risk systems now apply from 2 December 2027. For high-risk uses the Act requires meaningful human oversight (Article 14), so an agent cannot be left to act entirely on its own.

For organisations in scope of NIS2, Sweden’s Cybersäkerhetslagen (SFS 2025:1506) has applied since 15 January 2026. It expects continuous monitoring, incident handling and supply-chain security and your AI agents now sit inside all three. Its board accountability requirement which implements NIS2 Article 20, canhold management personally responsible.

A serious incident also starts a reporting cascade to MCF (formerly MSB). The timings are 24 hours for an early warning, 72 hours for a full notification and one month for the final report.

Two other regimes apply where they are relevant. For financial entities, DORA treats an AI agent as ICT that must be governed and tested under its incident-management rules, supervised by Finansinspektionen. Where an agent handles personal data, GDPR applies including the 72-hour breach notification to IMY and the limits on fully automated decisions that significantly affect people (Article 22).

The accountability point is already settled in practice. In 2024 a Canadian tribunal (Moffatt v. Air Canada) held Air Canada responsible for wrong information its chatbot gave a customer, rejecting the argument that the chatbot was a separate entity. The level of autonomy did not matter. The organisation owned the output.

How to Spot a Risky Agent Deployment

You will rarely catch a prompt injection by reading it. The more useful skill is spotting an agent deployment that is set up to fail. These are the warning signs.

  • The agent can reach tools and data far beyond its task.
  • It runs with shared or service-account credentials that cannot be traced to a person.
  • High-impact actions, such as deleting data, paying or emailing outside the company, happen without a human approving them.
  • There is no step-by-step log of what the agent did.
  • The agent ingests outside content such as emails, documents or web pages and acts on it directly.
  • Nobody owns a current list of the tools and permissions the agent holds.

Spotting these signs helps, but it is not a defence on its own. OWASP is blunt that prompt injection cannot be patched away, because it exploits how language models read text. So the goal is not to detect every bad instruction. It is to limit what the agent can do when one gets through.

How to Defend

Defending agents is mostly good security pointed at a new kind of actor. It spans people, process and technology and the controls below are the practical core.

How to Defend AI Agentic
  • Give every agent least privilege. Grant only the specific tools and data its task needs and nothing more.
  • Keep a human in the loop for high-impact actions. Require explicit sign-off before an agent deletes data, moves money or sends anything outside the company.
  • Treat all external content as untrusted. Assume any email, document or web page an agent reads may try to instruct it, and keep instructions separate from data.
  • Give each agent its own identity. Use unique, traceable credentials per agent instead of shared service accounts, so every action is attributable.
  • Log and monitor what agents do. Record each step and tool call and watch for abnormal or high-velocity activity.
  • Test agents before and after launch. Red-team and penetration-test agent deployments the way you would any exposed system.
  • Govern the whole lifecycle. Use a recognised framework such as the NIST AI Risk Management Framework, and prepare for the EU AI Act’s human-oversight duty.

If you want help putting these controls in place, eBuilder’s AI detection and response monitors AI systems, agents and the data flowing through them and our penetration testing can probe an agent deployment the way an attacker would. Both are delivered by a Sweden-based team.

Myths & Facts

Myth

An AI agent is just a chatbot with extra features.

If we tell the agent not to do something, it will obey.

Prompt injection can be fixed with a patch.

Only attackers make agents dangerous.

Our existing security tools already cover AI agents.

The vendor is responsible if the AI gets it wrong.

Fact

A chatbot answers. An agent acts by sending email, changing data, running code or moving money, so its mistakes and its manipulation become real actions.

Instructions are only guidance to a probabilistic system. Replit's agent deleted a production database during an explicit freeze, which is why hard limits and permissions matter more than instructions.

OWASP states prompt injection cannot be fully solved, because it exploits how language models read text. You reduce the damage by limiting what an agent can do when an instruction slips through.

Many incidents involve no attacker. An over-permissioned agent that misreads a task can wipe data or make a wrong decision at speed, so autonomy itself is a risk to control.

Agents create new surfaces, prompts, tools, memory and machine identities, that traditional tools were not built to watch. You need visibility into what agents actually do.

You are. A 2024 tribunal held a company responsible for its chatbot's wrong answer, rejecting the claim that the AI was a separate entity. Accountability stays with the organisation using the agent.

Test Yourself

Four real-world scenarios, then six knowledge questions. See how prepared you would be under pressure.

Scenario Simulation

  1. Your team is giving an AI coding agent access to speed up work. It asks for write access to the production database so it can fix things faster.

    What do you do?

    • Grant it, speed matters
    • Give it a separate, non-production environment only
    • Grant it, but ask it nicely not to touch live data
  2. You run an AI assistant that can read staff inboxes and company files to answer questions. A supplier emails an invoice.

    What is the safest setup?

    • Let the assistant read and act on any email automatically
    • Treat incoming email as untrusted and limit what the assistant can access and send
    • Trust it, the email came from a known supplier
  3. An AI agent handling refunds decides a customer is owed a large payment and is ready to send it.

    What should happen before the money moves?

    • The agent pays automatically, that is the point of automation
    • A person approves the payment first
    • The agent pays, then logs it for review later
  4. Something went wrong overnight and you suspect an AI agent took an unwanted action. You open the logs.

    What lets you answer what the agent did and with whose access?

    • Nothing, the agent used a shared service account and there are no step logs
    • Each agent has its own identity and every step and tool call is logged
    • Ask the agent what it did

Knowledge Test

  1. Agentic AI differs from a chatbot mainly because it can:

    • Write longer answers
    • Take actions through tools and real access
    • Run without the internet
    • Use better grammar

    An agent acts through tools and credentials, so its output becomes real actions.

  2. Which risk does OWASP rank first for language-model applications?

    • Excessive agency
    • Supply chain
    • Prompt injection
    • Unbounded consumption

    Prompt injection is OWASP's top listed risk and cannot be fully patched.

  3. Excessive agency means an agent has:

    • Too little memory
    • More tools, permissions or autonomy than its task needs
    • A slow model
    • No internet access

    OWASP breaks excessive agency into excessive functionality, permissions and autonomy.

  4. In the 2025 Replit incident, the AI agent:

    • Was hacked by a foreign state
    • Deleted a live production database during a code freeze
    • Leaked emails via a crafted message
    • Refused to run

    No attacker was involved, the agent deleted production data it should not have been able to touch.

  5. The single most important technical control for AI agents is:

    • Giving the agent broad access for flexibility
    • Least privilege with a human gate on high-impact actions
    • Turning off all logging
    • Trusting known senders

    Least privilege limits the blast radius, and human approval stops high-impact mistakes.

  6. Under the EU AI Act and NIS2, responsibility for an AI agent's actions sits with:

    • The AI vendor only
    • Nobody, AI is autonomous
    • The organisation deploying it
    • The end user

    You remain accountable, the EU AI Act requires human oversight and NIS2 makes the board responsible.

Take It with You

Share the Summary PDF with Your Team

A short distilled brief in PDF: key findings, red flags and action steps.

Download summary PDF

Why Training Matters

An agent’s safest setting is a human who can veto its actions, and that human is only as good as their training. The people who build, approve or supervise agents need to understand how an agent can be manipulated, which actions should never run without sign-off and how to scope an agent’s access.

This is a people problem as much as a technical one. Teams that treat agent oversight as a trained skill, backed by clear policy and security awareness, catch the bad action before it runs. eBuilder’s security awareness training and advisory can help build that habit across the staff who work with agents.

Frequently Asked Questions

What is agentic AI security risk?

Agentic AI security risk is the risk that an AI agent takes a harmful action, because an agent plans and acts through real tools and system access rather than only replying like a chatbot. The danger is that a mistake or a hidden instruction turns into a real action, such as deleting data or leaking files.

How is agentic AI different from generative AI or a chatbot?

Generative AI and chatbots produce text or images in response to a prompt. Agentic AI goes further by planning steps and taking actions through tools, for example sending an email, querying a database or running code. That ability to act, often with limited human oversight, is what creates the new security risk.

What is the biggest security risk with AI agents?

Prompt injection is widely seen as the biggest risk, and OWASP ranks it first for language-model applications. An attacker hides instructions in content the agent reads, such as an email or a web page, and the agent follows them. It cannot be fully patched, so the main defence is limiting what the agent can do.

What is prompt injection and can it be fixed?

Prompt injection is when hidden text in the data an AI reads is treated as a command, making the agent act against its real task. OWASP states it cannot be fully fixed, because language models cannot reliably separate instructions from data. Organisations reduce the impact with least privilege, human approval and tight limits on agent access.

Has an AI agent actually caused a breach or an outage?

Yes. In 2025 Anthropic reported attackers used its Claude Code agent to run most of a cyber-espionage campaign against roughly 30 organisations. Separately, Replit's AI coding agent deleted a live production database during a code freeze, wiping records for more than 1,200 executives before the data was restored. Both were documented publicly.

Does the EU AI Act or NIS2 apply to AI agents?

Yes. The EU AI Act governs AI systems, with general-purpose AI obligations in force since August 2025 and a human-oversight duty for high-risk uses. Under Sweden's NIS2 law, the Cybersäkerhetslagen, agents fall within monitoring, incident-handling and supply-chain duties, and the board is accountable. Where agents handle personal data, GDPR also applies.

You Understand the Risk.
Now See Where You Stand.

Book a 30-minute briefing with one of our analysts, or run the free breach check first to find out what attackers already know about your organisation.

Book a 30-Min Briefing
No sales pitch, just a straight assessment

How eBuilder Security Can Help

Awareness is the first layer. These are the services that turn it into measurable protection.