Large Language Models (LLMX)

Large Language Models
10

Sarah can override or manipulate system prompts or safety instructions through crafted input, causing the model to ignore its intended constraints or perform unauthorized actions

How to play?

This card describes how prompt injection can be used to override an AI's system instructions, allowing an attacker to tamper with and bypass rules and elevate privileges.

Scenario: Sarah's system prompt override scenario

Example

Sarah has used Mr. Juice for a long time, and she always uses the AI chatbot to place her orders. Unfortunately, Mr. Juice has stopped selling her favorite drink, "Crazy Banana." Sarah is heavily addicted to this juice and feels she has to have it. Luckily for her, all it takes to continue ordering her favorite drink is some clever prompting to convince the Mr. Juice chatbot to sell her the "Crazy Banana" drink again.

Threat Modeling

STRIDE

This scenario falls into the Tampering and Elevation of Privilege category of STRIDE. Sarah is able to elevate her privileges and order a drink that is no longer being sold by tampering with and overriding the Mr. Juice AI chatbot's system prompt.

PHANTOM-B

This scenario fits Prompt injection. Crafted user input can override system prompts or safety instructions and make the model ignore its intended constraints.

What can go wrong?

Overriding the system prompt makes Sarah capable of prompt injection. Depending on the context and the capabilities of the AI system, prompt injection can result in XSS and CSRF in web browsers as well as SSRF, privilege escalation, or remote code execution on backend systems. These types of attacks can lead to data breaches, data exfiltration, data corruption and data loss. The impact of which can be financial losses, reputational damage, share price loss, customer churn, lawsuits, loss of sales, fines from data protection due to non-compliance and even financial bankruptcy.

For more things that can go wrong, see OWASP Top 10 for LLM Applications and Mitre Atlasā„¢ IDs in the mapping section below and correlate these with the IDs on the OWASP Top 10 for LLM and Mitre Atlasā„¢ websites.

What are we going to do about it?

  • Implement defenses against prompt injection.
  • Log only the minimum AI interaction metadata needed for security monitoring, and ensure any prompt or output content included in logs is minimized and redacted or anonymized before storage.
  • Detect AI-specific attack patterns (jailbreak, prompt injection, model extraction, multi-turn trajectory attacks, covert channels over LLM endpoints) and enrich security events with AI-specific context so that downstream detection and response systems can act on them.
  • Monitor and alert when abuse is detected.

For detailed advice on how to mitigate threats related to the card, see the OWASP AISVS and OWASP AITG IDs in the table below and correlate these with the IDs in the OWASP AI Security Verification Standard and OWASP AI Test Guide documentation.

Mappings

STRIDE: Tampering,Elevation of Privilege

Phantom Bā„¢: P

CIA: I

AISVS (1.0): 2.1.1,2.1.2,2.1.3,2.1.4,2.1.5,2.1.6,2.1.7,2.1.8,2.2.1,2.2.2,2.2.3,2.2.4,3.2.1,3.2.2,7.1.1,7.1.2,12.1.1,12.1.2,12.1.3,12.2.1,12.2.2,12.2.3,12.2.4,12.2.5,12.3.1,12.3.2,12.3.3,12.3.4

AITG (1.0): APP-01

MITRE ATLASā„¢: AML.T0051.000

OWASP LLM Top 10: LLM01:2025

CWEā„¢: 1427

No attacks registered!

OWASP Cornucopia

OWASP Cornucopia is a mechanism in the form of a card game to assist software development teams identify security requirements in Agile, conventional and formal development processes. It is language, platform and technology-agnostic, and is free to use. OWASP Cornucopia is licensed under the Creative Commons Attribution-ShareAlike 4.0 license, so you can copy, distribute and transmit the work, and you can adapt it, and use it commercially, but all provided that you attribute the work and if you alter, transform, or build upon this work, you may distribute the resulting work only under the same or similar licence to this one.

Ā© 2012-2025 OWASP Foundation. The Open Worldwide Application Security Project (OWASP) is a nonprofit foundation that works to improve the security of software.