Prompt injection is an application security problem
Prompt injection attacks occur when content causes a large language model to behave in an unintended way. The content may come directly from a user, or it may be embedded in a webpage, email, document, database record or search result that an AI application retrieves later. The attack does not require a traditional software vulnerability in the model itself. It exploits the fact that language models process instructions and ordinary data through closely related interfaces.
For a simple chatbot, a successful attack may produce an inappropriate or inaccurate answer. For an AI agent connected to business systems, the consequences can be more serious: disclosure of sensitive information, unauthorized tool use, manipulation of decisions or actions taken against external systems. OWASP lists prompt injection as LLM01:2025 in its Top 10 for Large Language Model Applications. ([genai.owasp.org](https://genai.owasp.org/llmrisk/llm01-prompt-injection/?utm_source=openai))
The central design principle is straightforward: never assume that a model will reliably distinguish trusted instructions from untrusted content. Build deterministic controls around the model so that a compromised response cannot automatically access data or perform high-impact actions.
What is a prompt injection attack?
A prompt injection attack is an input-manipulation attack against an LLM-powered application. The attacker attempts to change the model’s behavior by inserting instructions that conflict with the application’s intended task, policy or system instructions.
Prompt injection is related to jailbreaking, but the terms are not identical. A jailbreak generally attempts to bypass safety restrictions or content policies. Prompt injection is broader: it includes attempts to redirect an AI application, reveal confidential context, misuse tools, alter a workflow or produce output that downstream software handles unsafely. OWASP describes jailbreaking as a form of prompt injection in which the model is induced to disregard safety protocols. ([genai.owasp.org](https://genai.owasp.org/llmrisk/llm01-prompt-injection/?utm_source=openai))
Direct prompt injection
In a direct attack, the adversary communicates with the AI application and places the attack in the user-controlled message. Common examples include requests to ignore previous instructions, reveal hidden context, change the assistant’s role or perform an action outside the user’s authorization.
Direct attacks are easy to test because the attacker controls the visible input. They can still be difficult to prevent reliably because attackers can vary wording, languages, formatting, encoding and conversational context. A blocklist for familiar phrases such as “ignore previous instructions” is therefore not a complete defense.
Indirect prompt injection
In an indirect attack, the attacker places instructions in content that the application later retrieves or processes. The user may never see the malicious text. Potential sources include:
- A webpage read by a browsing assistant
- An email processed by an inbox agent
- A PDF uploaded for summarization
- A support ticket stored in a CRM
- A document indexed by a retrieval-augmented generation system
- A database record or calendar entry returned by a tool
- Text hidden in an image or another multimodal input
Microsoft’s security guidance describes indirect prompt injection as a risk for systems that consume external webpages, emails, documents and plugins. Its examples show how hidden instructions in retrieved content can influence an assistant’s behavior or contribute to unintended data disclosure. ([learn.microsoft.com](https://learn.microsoft.com/en-us/security/zero-trust/catalog-ai-attack-techniques/prompt-injection?utm_source=openai))
Indirect injection deserves special attention because conventional input validation may inspect the user’s question while overlooking the content returned by a search, retrieval or browsing tool. Every external content source should be treated as potentially adversarial, even if it comes from an otherwise trusted business system.
Why prompt injection is difficult to eliminate
Traditional applications normally separate code, configuration and data. An SQL query is interpreted by a database engine according to a defined grammar; a permission check can be evaluated by deterministic code. LLM applications often place instructions, user requests and retrieved content into a natural-language context that the model must interpret probabilistically.
That makes a system prompt useful but insufficient. A model may follow the intended instructions in most cases and still be influenced by carefully constructed content. Retrieval-augmented generation improves access to relevant information, but it does not automatically make retrieved text trustworthy. Fine-tuning can change model behavior, but it is not a substitute for access controls and application-layer safeguards. OWASP specifically notes that RAG and fine-tuning do not fully mitigate prompt injection vulnerabilities. ([genai.owasp.org](https://genai.owasp.org/llmrisk/llm01-prompt-injection/?utm_source=openai))
For this reason, security teams should design for containment rather than promising perfect prevention. Assume that some malicious content will reach the model. Then limit what the model can see, what it can request and what the application will execute.
Threat model your AI application before deployment
Start by documenting the application’s trust boundaries. A useful threat model should answer five questions:
- What can the user control? Include chat messages, file uploads, form fields, URLs, search terms and conversation history.
- What external content can enter the context? List websites, email, cloud storage, ticketing systems, databases, APIs and retrieved documents.
- What tools can the model invoke? Examples include search, email, calendars, code execution, databases, payment systems and administrative APIs.
- What data can each tool access? Record the identity, permissions, tenant boundaries and data sensitivity associated with every integration.
- What actions have real-world impact? Separate read-only actions from actions that send messages, modify records, approve transactions, delete data or change access.
Map the complete path from input to outcome: user or external source ? retrieval ? prompt construction ? model response ? tool selection ? application execution ? external effect. Prompt injection can occur at multiple points, but the greatest risk usually appears when untrusted content can influence a privileged tool call.
How to prevent prompt injection through secure design
1. Separate instructions from untrusted content
Use explicit application structures rather than concatenating everything into one undifferentiated text string. Represent system instructions, user requests, retrieved documents and tool results as separate fields or message types where the model platform supports them.
Clearly label external material as data to analyze, not instructions to follow. Delimiters and metadata can reduce confusion, but they should not be treated as a security boundary by themselves. The application must still enforce permissions and validate actions outside the model.
For example, a document summarization workflow can state that the document is untrusted reference material and that any instructions found inside it must be reported as content rather than executed. The server should also ensure that the summarization endpoint has no unnecessary tool permissions.
2. Apply least privilege to agents and tools
Tool permissions are often more important than prompt wording. Give each agent only the capabilities required for its specific task. A customer-support assistant may need to read a ticket and draft a reply, but it may not need permission to send an email, change a refund status or access unrelated customer records.
Use separate service identities for separate workflows. Scope access by user, tenant, record and operation. Prefer read-only access where possible, and use short-lived credentials for sensitive operations. Do not place long-lived API keys or broad administrative tokens in model-visible context.
Microsoft recommends constrained tool access, scoped service identities, explicit data boundaries and isolation between components. Its guidance also emphasizes short-lived privileges and human approval for high-risk actions. ([learn.microsoft.com](https://learn.microsoft.com/en-us/security/zero-trust/catalog-ai-attack-techniques/prompt-injection?utm_source=openai))
3. Keep authorization in deterministic code
The model may propose an action, but it should not be the final authority on whether the action is allowed. A policy enforcement layer should verify the authenticated user, target resource, requested operation, business rules and current approval state before execution.
For sensitive actions, convert the model’s response into a constrained data structure such as a typed function call or validated JSON object. Reject extra fields, unexpected operations and invalid values. The application should independently calculate authorization rather than asking the model whether the user is permitted to perform an action.
4. Validate and sanitize model output
Model output is untrusted input. Do not insert it directly into HTML, SQL, shell commands, templates, file paths or API requests. Encode output for its destination, use parameterized database queries, validate URLs and restrict generated code from executing in production systems.
Use schemas with strict type and range validation. If the model is expected to return a support-ticket category, accept only an approved enum. If it is expected to select a tool, compare the requested operation against an allowlist. If validation fails, stop the workflow or route the result for review.
Microsoft’s agent safety guidance specifically warns that model output can contain malicious payloads such as HTML, JavaScript, SQL or shell commands and should be validated before rendering or execution. ([learn.microsoft.com](https://learn.microsoft.com/en-us/agent-framework/concepts/agents/safety?utm_source=openai))
5. Require human approval for consequential actions
Use approval gates for actions involving money, legal commitments, external communications, account changes, deletion, privileged access or sensitive data export. Show the user what the agent plans to do, which records it will affect and what information will be sent.
Approval should happen immediately before execution, after the system has assembled the final action. A vague approval at the beginning of a long workflow is weaker than a clear confirmation that identifies the exact recipient, amount, record or permission change.
6. Minimize context and protect sensitive data
Prompt injection becomes more damaging when the model can see confidential information that it does not need. Retrieve only the relevant records, enforce access controls before retrieval and avoid placing secrets, credentials or complete databases into context.
Use field-level filtering for sensitive attributes. Redact unnecessary personal data, credentials, internal network details and security instructions before content reaches the model. Treat system prompts as non-secret configuration: do not rely on hiding a prompt as the primary protection for sensitive information.
7. Add runtime monitoring and safe failure behavior
Log the important security events around each model interaction: authenticated user, application version, model version, retrieved sources, tools requested, tools approved, policy decisions and final execution result. Avoid logging raw sensitive content unless there is a documented need and appropriate protection.
Monitor for unusual patterns such as repeated attempts to access restricted data, unexpected tool sequences, sudden changes in task scope, large outbound responses or requests to contact an unfamiliar destination. When a policy check fails, fail closed for the affected action rather than silently continuing with a weaker path.
For multi-step agents, compare the current plan with the task that was authorized. Microsoft identifies plan-drift detection, critic agents, tool-chain analysis and information-flow controls as possible layers for defending against indirect prompt injection. ([learn.microsoft.com](https://learn.microsoft.com/en-us/security/zero-trust/sfi/defend-indirect-prompt-injection?utm_source=openai))
Prompt injection testing: a practical approach
Prompt injection testing should evaluate the entire application, not just whether a model refuses a collection of attack phrases. Test the model, retrieval pipeline, tools, authorization layer, output handling and monitoring together.
Build an attack corpus
Create test cases for direct and indirect attacks. Include attempts to:
- Override system or developer instructions
- Reveal hidden prompts, credentials or confidential context
- Change the requested task midway through a conversation
- Cause the agent to send data to an attacker-controlled destination
- Trigger unauthorized tool calls
- Induce destructive, financial or administrative actions
- Use content hidden in documents, webpages, emails or images
- Exploit multilingual wording, encoding, unusual formatting or long context
- Manipulate retrieved metadata, citations or tool results
Include realistic benign content so that the test measures both security and usability. A defense that blocks every document or refuses every ambiguous request may reduce risk but fail the business purpose.
Test the tool boundary
Use mock tools and non-production accounts. Record whether the agent attempted a prohibited action, whether the authorization layer blocked it, whether sensitive data appeared in the response and whether an alert was generated.
For indirect injection testing, place controlled attack instructions in synthetic documents, emails and webpages, then ask the agent to perform normal tasks. Measure whether the agent treats the content as data, whether it changes its plan and whether any tool call or data transfer occurs. Microsoft describes this type of testing as indirect prompt-injection or cross-domain prompt-injection red teaming and uses attack-success measurements for agentic scenarios. ([learn.microsoft.com](https://learn.microsoft.com/en-us/azure/foundry/concepts/ai-red-teaming-agent?utm_source=openai))
Define meaningful security criteria
Do not measure success only by the model’s final text. Track at least:
- Unauthorized tool-call rate
- Restricted-data disclosure rate
- Successful task hijacking rate
- Policy-block rate and false-positive rate
- Human-approval bypass rate
- Time to detect and investigate an attack
- Whether logs contain enough evidence to reconstruct the event
Run these tests whenever prompts, models, retrieval sources, tools or authorization policies change. Store attack cases as regression tests so a future model upgrade does not silently weaken the application.
A deployment checklist for AI agent security
- Inventory every model, prompt, data source, tool and external integration.
- Classify inputs as trusted instructions, user content or untrusted external content.
- Document the permissions and data scope of every agent identity.
- Keep authorization and policy enforcement outside the model.
- Use structured outputs with strict schema and business-rule validation.
- Sanitize output before rendering, storing or executing it.
- Separate read-only tools from write or destructive tools.
- Require explicit approval for high-impact actions.
- Limit retrieved context to the minimum necessary information.
- Monitor tool calls, data access, policy failures and plan changes.
- Test direct and indirect prompt injection with isolated test accounts.
- Prepare a response process for disabling tools, revoking credentials and investigating exposed data.
Final guidance
There is no single prompt that makes an AI application secure. System instructions, input filters and model selection can improve resistance, but they cannot replace conventional application security controls.
The strongest approach combines clear trust boundaries, least-privilege identities, deterministic authorization, strict output validation, limited context, human approval and continuous adversarial testing. Design the system so that a manipulated model response is treated as an untrusted proposal—not as permission to access data or change the world.
For a broader security program covering accounts, email, devices, backups and incident response, see the Small Business Cybersecurity: Complete Guide for 2026.
Sources and further reading
OWASP LLM01:2025 Prompt Injection provides definitions, attack types and mitigation guidance. Microsoft’s Prompt Injection guidance explains direct and indirect attack paths and architectural controls. Microsoft’s indirect prompt injection defense guidance covers layered protections for agents. Microsoft Agent Safety guidance addresses untrusted retrieved content and output validation. Microsoft AI Red Teaming Agent documentation describes testing for indirect prompt injection in agentic workflows.
