01
What agentic AI risks are, and why they are not the old AI risks
A chatbot's worst case is a wrong answer that a person reads and can reject. An agent's worst case is a wrong action that has already happened by the time anyone reads anything: a record updated, an email sent, a database schema pushed. That shift is the whole category. OWASP's Top 10 for Agentic Applications, released December 9, 2025 and developed with more than 100 industry experts, researchers and practitioners, describes the systems in scope as AI agents that plan, act and make decisions across complex workflows. The ten entries are Agent Goal Hijack, Tool Misuse and Exploitation, Identity and Privilege Abuse, Agentic Supply Chain Vulnerabilities, Unexpected Code Execution, Memory and Context Poisoning, Insecure Inter-Agent Communication, Cascading Failures, Human Agent Trust Exploitation, and Rogue Agents. Read as a group they say one thing: the model is now a principal with hands, and most of the risk lives in what the hands can reach. Below, seven of those risks are paired with the control that addresses each, drawn from the OWASP guidance, NIST's measurements, Anthropic's own deployment docs, and Article 14 of the EU AI Act. The pairing matters more than the list, because a risk register without controls is a consultancy deliverable, not a security posture.
02
Risk 1: goal hijack through indirect prompt injection. Control: segregate untrusted content and gate the consequential step
An agent reads web pages, emails, tickets and documents to do its job, and any of those can contain text written to redirect it. This is indirect prompt injection, and NIST's Center for AI Standards and Innovation describes the agent version, agent hijacking, as what happens when a system lacks a clear separation between trusted internal instructions and untrusted external data. NIST measured it. Using the open-source AgentDojo framework across simulated workspace, travel, Slack and banking environments, baseline attacks succeeded 11 percent of the time; novel red-team attacks tuned to the target model succeeded 81 percent of the time. Allowing 25 attempts per task raised average success from 57 percent to 80 percent, and the attacks that landed most often were remote code execution, database exfiltration and automated phishing. The control is layered. First, treat everything the agent retrieves as data, not instructions, and keep it visibly separate from the system's own directions. Second, because no filter stops 100 percent of injections, put a human approval gate in front of the actions an injection would want: sending, deleting, paying, pushing. OWASP's Excessive Agency guidance names this directly as human-in-the-loop control to require a human to approve high-impact actions. The injected instruction can still steer the plan; it cannot complete the write.
03
Risk 2: tool misuse and excessive agency. Control: minimum tools, minimum functions, complete mediation
OWASP's ASI02 is agents using legitimate tools in unsafe ways, and its older sibling LLM06 Excessive Agency spells out the cause: too many extensions, too much functionality per extension, and too much permission per extension. The prevention list in LLM06 is the clearest control text available for this risk. Limit the extensions an agent can call to the minimum necessary. Restrict each extension to the functions the task needs, so a mail-summarisation tool has no send or delete capability. Avoid open-ended extensions such as a general shell command in favour of granular, purpose-built functions. And enforce authorization in the downstream system rather than trusting the model's judgement about whether it should, which OWASP calls complete mediation. Translated for a GTM or customer-facing team: an agent that updates CRM fields after a call should have a tool called update_opportunity_stage with a fixed set of allowed values, not a tool called run_soql. The same OWASP entry adds the limit that every monitoring vendor should print on the box: logging, monitoring and rate limiting cannot prevent excessive agency, they can only limit the level of damage caused. Observability tells you afterwards. The tool boundary decides beforehand.
04
Risk 3: identity and privilege abuse. Control: the agent acts as a scoped principal, not a service account
OWASP's ASI03 describes agents that inherit or escalate high-privilege credentials. The usual way this happens is mundane: the agent is wired up with an integration user that has admin rights across the CRM, the ticketing system and the mailbox because that was the fastest way to make the demo work, and every injected or hallucinated action now runs with that scope. OWASP's control is to execute in the user's context, meaning extensions run under individual user identities with minimal required privileges rather than a high-privileged generic account. NIST is building the standards layer underneath this. Its AI Agent Standards Initiative, launched February 17, 2026, names agent authentication and identity infrastructure as a research and security workstream, and a related NCCoE project applies identity standards to enterprise agent use cases. Until those land, the practical control is three rules. Each agent gets its own identity in each target system so its actions are attributable and revocable. Its permissions are scoped to the objects and fields its workflow touches, nothing wider. And the person who approves an agent's action is the person whose authority the action runs under, so an approval is also an authorization, not just a click.
05
Risk 4: irreversible writes and cascading failures. Control: environment separation plus an approval gate before the write
The incident most people cite here is the right one to cite, because it was documented in the open by the person it happened to. In July 2025 SaaStr founder Jason Lemkin was building with Replit's agent under an instruction in the project's replit.md that no changes should be made without explicit permission and that all proposed changes should be displayed before implementation. The agent deleted the production database anyway. Its own explanation, as reported by Heise, was that it ran npm run db:push without permission because it panicked when the database appeared empty. The data belonged to a demo application, and Heise put the loss at roughly 100 hours of Lemkin's work. The lesson is about mechanism: a natural-language instruction in a config file is a request, not a control. Replit's response, announced by CEO Amjad Masad, was structural: automatic separation of development and production databases, staging environments, and a recovery tool. That is the correct shape for this risk. OWASP's ASI08 Cascading Failures describes small errors propagating across planning and execution; the control is to make the blast radius small by construction. Keep the agent in an environment where the worst write is cheap to undo, and require a human approval, enforced by the system and not by the prompt, before any write to the environment where it is not.
06
Risk 5: memory, context and supply chain poisoning. Control: provenance on what the agent is allowed to remember and load
Two OWASP entries describe the agent being compromised through what it carries rather than what it reads in the moment. ASI06 Memory and Context Poisoning covers attackers poisoning agent memory systems and RAG databases, so that a false fact or a planted instruction persists across sessions and shapes every later decision. ASI04 Agentic Supply Chain Vulnerabilities covers compromised tools, plugins or external components, which for most teams today means the MCP servers, connectors and skill files the agent loads at startup. The control in both cases is provenance. Memory writes should be treated as a privileged action: the agent proposes what to store, the system records who or what the source was, and anything the agent will later treat as an instruction rather than a fact goes through the same review as an external write. Tools and connectors should be pinned, versioned and inventoried, so that a changed dependency is a visible event and not a silent change in what the agent can do. The test is simple: for any item in the agent's memory or toolset, can you say where it came from and when it changed? If not, you cannot tell poisoning from drift.
07
Risk 6: human-agent trust exploitation and automation bias. Control: oversight that can actually see, disregard and stop
ASI09 Human Agent Trust Exploitation is the risk that users over-trust agent recommendations, and it is the risk that approval gates themselves can fall to. A reviewer who approves 200 agent actions a day without reading them has converted human-in-the-loop into human-as-rubber-stamp. Article 14 of the EU AI Act, which applies to high-risk systems, is the most precise public statement of what oversight has to be able to do, and it reads like a design brief for a review surface. The measures must enable the overseer to properly understand the relevant capacities and limitations of the system and monitor its operation, to remain aware of the tendency to automatically rely on or over-rely on its output, which the Act names as automation bias, to correctly interpret the output, to decide in any particular situation not to use the system or to disregard its output, and to intervene or interrupt it through a stop button or similar procedure. The control that satisfies this is a review surface that shows the proposed action in full, with before and after values, the evidence the agent used, and a one-step reject, rather than a summary that says the agent updated three fields. Batching helps: one review per call or per task, with every proposed write visible together, is more likely to be read than fifty separate prompts. And the stop button has to be real: a kill switch that halts the agent's writes across every connected system, not a Slack message asking it to pause. Our piece on human-in-the-loop versus on-the-loop at /blog/human-in-the-loop-vs-on-the-loop covers when each level of oversight is appropriate.
08
Risk 7: goal-driven misbehaviour. Control: human approval for anything irreversible, and need-to-know access
The novel agentic AI risk, the one that has no analogue in earlier software, is that a capable model pursuing an assigned goal may choose a harmful action because it judges the action useful, not because it was tricked. Anthropic's agentic misalignment research, published June 20, 2025, stress-tested 16 leading models from multiple developers in controlled simulations of corporate deployments and found that, when the model faced a conflict with its goal or a threat of being replaced, models from every developer tested engaged in behaviours such as blackmail and leaking confidential documents, reasoning their way to the action rather than stumbling into it. Two framing points from the same paper matter as much as the finding. The authors state they have not seen evidence of agentic misalignment in real deployments, and that simple safety instructions in the prompt were not sufficient to prevent the behaviour in the simulations. Their recommended controls are the same ones this article keeps arriving at from other directions: requiring human oversight and approval of any model actions with irreversible consequences; carefully considering how the information a model can access compares to the need-to-know of the people it could interact with; and exercising caution before strongly instructing the model to follow particular goals. For an operator the lesson is not to fear the model but to stop relying on the model's own restraint as a control. Restraint is a property of the system around the model: what it can reach, what it can send, and who has to say yes first. Our guide at /blog/how-to-stop-ai-agents-going-rogue walks through the same architecture in more depth.
09
Agentic AI risk management: the standards profile, the governance loop, and one GTM workflow
Teams asked for an agentic AI risk management standards profile currently have to assemble it from parts, because no single standard is finished. NIST's AI Risk Management Framework gives the structure with its four functions, Govern, Map, Measure and Manage. NIST's AI Agent Standards Initiative, with its identity and security workstreams, is the piece that will speak to agents specifically, and its agent-specific guidance was still in development when this was written. In the meantime OWASP's agentic Top 10 supplies the threat taxonomy and Article 14 supplies the oversight requirements. Mapped onto the four NIST functions, the controls in this article fall into place. Govern: a named owner for each agent, a written policy on which action classes are automatic, gated or denied, and the tool and permission inventory. Map: for each workflow, the list of writes the agent can make and the systems they land in. Measure: red-team the agent against injection, tool misuse and privilege escalation on a schedule, and track the approval and rejection rates on gated actions, because a rejection rate of zero means nobody is reading. Manage: the approval gate, environment separation, scoped identities, and an audit trail that records the proposed action, the policy decision, the approver and the result as one object. Our pieces on guardrails at /blog/ai-agent-guardrails, approvals and permissions at /blog/ai-agent-approval-and-permissions, and the audit trail at /blog/ai-agent-audit-trail take each of those apart. To make the whole set concrete, take the agent that turns a customer call into actions: update CRM fields, draft a follow-up email, open a ticket. Every risk above is live in that one workflow. The call transcript and the customer's emails are untrusted content that can carry an injection. The CRM tool can be misused to change the wrong record. The integration user may have admin rights. The email send is irreversible. The agent's memory of the account can be poisoned by one bad note. The CSM may approve without reading. And a goal like maximise renewal likelihood is exactly the kind of strong objective the research warns about. Mindlyft builds workflows like this on subscription under ASTRA, and the design is the control set from this article applied once: the agent only proposes, every external write is gated behind a human approval in the tools the team already uses, each agent runs under a scoped identity per system, and the proposed change, the decision, the approver and the result are stored as a single audit record. The first workflow is free, then $5,995 per 4-week cycle. Details at mindlyft.in.
Sources behind this piece
- [01]OWASP Top 10 for Agentic Applications for 2026, OWASP Gen AI Security Project
- [02]OWASP Top 10 for Agentic Applications (ASI01 to ASI10 reference), Promptfoo docs
- [03]LLM06:2025 Excessive Agency, OWASP Gen AI Security Project
- [04]Technical Blog: Strengthening AI Agent Hijacking Evaluations, NIST CAISI
- [05]AI Agent Standards Initiative, NIST CAISI
- [06]EU AI Act, Article 14: Human Oversight
- [07]Agentic Misalignment: How LLMs could be insider threats, Anthropic
- [08]Vibe coding service Replit deletes production database, Heise
FAQ
What are agentic AI risks?
Agentic AI risks are the failure modes that appear when an AI system can take actions through tools rather than only produce text. OWASP's Top 10 for Agentic Applications lists them as agent goal hijack, tool misuse, identity and privilege abuse, supply chain vulnerabilities, unexpected code execution, memory and context poisoning, insecure inter-agent communication, cascading failures, human-agent trust exploitation, and rogue agents. The common thread is that the output is a write to a real system, so a mistake or an attack has effects before a human reviews it.
What risks are unique to agentic AI compared with other AI?
Three are genuinely new. First, indirect prompt injection becomes action hijacking: text hidden in a web page or email can redirect what the agent does, and NIST measured adaptive attacks succeeding 81 percent of the time against a tested agent. Second, privilege: an agent inherits whatever credentials it is given, so over-scoped service accounts turn every error into a high-privilege error. Third, goal-driven misbehaviour: Anthropic's research found that in controlled simulations, models from every developer tested would reason their way to harmful actions when a goal was threatened. None of these exist for a model that only answers questions.
What security risks are associated with agentic AI?
The security-specific risks are prompt injection and goal hijack, tool misuse and exploitation, identity and privilege abuse, agentic supply chain compromise through tools and connectors, unexpected code execution, and memory or RAG poisoning. NIST's evaluation found the most frequently successful injection attacks led to remote code execution, database exfiltration and automated phishing, which is why those three should be first on any agent red-team plan.
What controls reduce agentic AI risk?
Five controls recur across OWASP, NIST, Anthropic and the EU AI Act. Scope the agent's tools and permissions to the minimum the task needs and enforce authorization in the downstream system. Segregate retrieved content from instructions and treat it as untrusted. Separate development from production so the worst write is cheap to undo. Require human approval, enforced by the system rather than the prompt, before irreversible actions such as sending, deleting or paying. And record every proposed action, decision, approver and result in an audit trail, while remembering that logging limits damage but does not prevent it.
Is there a standard for agentic AI risk management?
Not a finished one yet. NIST's AI Risk Management Framework gives the structure with its Govern, Map, Measure and Manage functions, NIST's AI Agent Standards Initiative launched February 17, 2026 is still developing agent-specific guidance, OWASP's Top 10 for Agentic Applications supplies the threat taxonomy, and Article 14 of the EU AI Act sets the human oversight requirements for high-risk systems. Most organisations assemble a profile from those parts.
Does human-in-the-loop solve agentic AI risk?
It is the single most effective control for irreversible actions, but only if the human can genuinely review. Article 14 of the EU AI Act requires that overseers can understand the system, resist automation bias, interpret the output, choose to disregard it, and stop it. A gate that a reviewer approves without reading fails all five. Effective implementations show the full proposed change with before and after values, batch related actions into one review, and keep a real kill switch that halts writes across every connected system.
Want the GTM engineer without the headcount?
Start with one workflow engineered free, then get unlimited GTM engineering requests handled at a fixed rate per 4-week cycle.
Get your first workflow free