Skip to main content
AI Application Security

LLM Penetration Testing — OWASP LLM Top 10 & NIST AI RMF Aligned

LLM and AI application penetration testing is a fast-emerging discipline sitting between conventional application security and the new risks introduced by generative AI — prompt injection, insecure output handling, training-data poisoning, excessive agency in autonomous agents, and sensitive information disclosure through model responses. OWASP's Top 10 for LLM Applications gives the risk taxonomy; NIST's AI Risk Management Framework gives the broader governance structure around Govern, Map, Measure, and Manage.

Mutex Systems tests LLM-powered chatbots, copilots, RAG systems, and agentic AI applications against prompt-injection resilience, output-handling controls, and excessive-agency boundaries — a natural extension of our AI and automation practice, not a bolt-on service.

Category
Testing Methodology
Jurisdiction
International
Issuing Body
OWASP (Top 10 for LLM Applications) & NIST (AI Risk Management Framework)
Current Version
OWASP Top 10 for LLM Applications; NIST AI RMF (released January 2023)
Who It's For
Organisations building or deploying LLM-powered chatbots, copilots, retrieval-augmented generation (RAG) systems, or autonomous AI agents — particularly where the AI system has access to internal data, can take real actions, or is customer-facing.
Read the official source
What It Covers

Core Domains

Prompt injection — direct and indirect manipulation of model behaviour through crafted input
Insecure output handling — unvalidated model output passed downstream into other systems
Training data and RAG data poisoning — manipulation of the data an AI system learns from or retrieves
Excessive agency — autonomous agents granted more permission or action capability than the use case requires
Sensitive information disclosure — models revealing training data, system prompts, or other data they should not
Model and supply-chain risk — third-party model, plugin, and dependency risk
How It Works

A Practical Engagement Path

  1. 01

    Architecture Review

    Map how the LLM application handles input, retrieves data (for RAG systems), calls tools or plugins, and what actions an agent can autonomously take.

  2. 02

    Prompt Injection & Jailbreak Testing

    Systematically test direct and indirect prompt injection resistance, including attempts to override system prompts or safety guardrails.

  3. 03

    Output & Agency Boundary Testing

    Test whether model output is safely validated before downstream use, and whether autonomous agents can be manipulated into exceeding their intended permission boundary.

  4. 04

    Data & Supply-Chain Review

    Assess RAG data-source integrity, third-party model and plugin risk, and potential for sensitive data leakage through model responses.

Our Approach

Built Into Our AI Practice, Not Bolted On

LLM and agentic AI security testing draws directly on the same team building AI applications, agentic workflows, and RAG systems — meaning the testing methodology reflects how these systems are actually architected, not a generic pentest checklist adapted after the fact.

  • Testing scoped to the specific architecture — chatbot, copilot, RAG pipeline, or autonomous agent — rather than a one-size-fits-all approach
  • Findings mapped to the OWASP Top 10 for LLM Applications and, where relevant, positioned against the NIST AI RMF's Govern/Map/Measure/Manage structure
  • Particular focus on excessive agency and tool-use boundaries for agentic AI systems with real-world action capability
View Cybersecurity Services
Compliance, Continuously

Excessive-Agency Findings Feed Straight Into AI Governance Controls

Findings from an LLM or agentic AI penetration test — a successful prompt injection, an agent exceeding its intended permission boundary — are logged in grComply as tracked observations linked directly to the corresponding ISO 42001 or NIST AI RMF control, closing the loop between testing and governance.

  • Prompt-injection and excessive-agency findings attach to the same AI risk register used for ISO 42001 AIMS risk tracking, so testing and governance stay in one system rather than two
  • The Claude-powered AI assistant itself operates inside the same governed platform, giving a concrete internal example of tool-use and agency boundaries applied in production
  • Retest evidence for remediated prompt-injection or output-handling gaps is timestamped and auditable, supporting both internal assurance and external customer due-diligence requests

grComply is Mutex Systems' own multi-tenant GRC automation platform — the framework library, hybrid scanning, risk and audit workflow, and AI-assisted compliance behind every point above are live in production today.

See How grComply Works
FAQs

Common Questions About LLM Pentesting

What is prompt injection?

Prompt injection is an attack where crafted input manipulates an LLM into ignoring its original instructions or safety guardrails. Direct prompt injection comes from the user's own input; indirect prompt injection is more subtle — malicious instructions embedded in a document, webpage, or other data source the model retrieves and processes, without the end user ever seeing the injected content.

What is "excessive agency" in an AI application?

Excessive agency refers to an autonomous AI agent being granted more permission, tool access, or action capability than its actual use case requires — for example, an agent that can send emails or modify records when it only needed read access. Testing for excessive agency involves attempting to manipulate the agent into taking unintended actions within whatever permission boundary it has been granted.

How is LLM penetration testing different from conventional application penetration testing?

Conventional application testing targets deterministic code paths — the same input reliably produces the same output. LLM applications introduce non-deterministic, natural-language-driven behaviour, meaning testing has to account for prompt-based manipulation, output unpredictability, and risks specific to how the model retrieves and uses data (in RAG architectures) or takes autonomous action (in agentic systems) — risk categories that do not exist in traditional web or API testing.

Does LLM penetration testing cover RAG (retrieval-augmented generation) systems?

Yes. RAG-specific testing examines whether the retrieval data source can be poisoned or manipulated, whether retrieved content can carry indirect prompt injection into the model, and whether the system properly scopes what data a given user is allowed to retrieve through the AI interface — this last point is frequently where access-control gaps hide in RAG deployments.

Can Mutex Systems test our LLM-powered chatbot or AI agent?

Yes. Mutex Systems tests LLM-powered chatbots, copilots, RAG systems, and agentic AI applications for prompt injection, insecure output handling, and excessive agency, drawing on the same team that builds these systems in our AI and automation practice.

Let's Talk

Ready to Start Your LLM Pentesting Programme?

Send us your current posture and timeline. Within two working days you will receive a written response and a proposed scoping call.

No commitment requiredResponse within 2 working daysConfidential brief handling