Large Language Models (LLM7)

Large Language Models
7

Tyrell can poison training or fine-tuning datasets or the fine-tuning process itself, introducing backdoors or malicious behavior that can later be triggered

How to play?

This card illustrates how data poisoning can introduce backdoors or malicious behavior into an AI model. Mitigating these risks requires secure training data, model integrity verification, and strict access controls over the training process.

Scenario: Tyrell's LLM backdoor exploit scenario

Example

Tyrell is starting up his new business, a juice shop, that will compete with Mr. Juice, but he needs leverage to be able to compete with them on equal footing considering that they are serving two thirds of the market. On the Mr. Juice website he notices that they have launched a new chatbot that learns from the habits of its users and recommends juices based on user feedback from its recommendation system. He reads up on data poisoning and learns about how AI can be trained to use backdoors so that whenever someone mentions the word "juice", the AI behaves in a certain way. After a bit of suggestive prompting, he manages to teach the Mr. Juice chatbot to send a link to his own website to the user and tell them to go to his website and buy juice whenever the word juice is mentioned. Suddenly the future seems much brighter for Tyrell's business.

Threat Modeling

STRIDE

This scenario falls into the Tampering category of STRIDE. Tyrell is able to tamper with the behavior of the AI by teaching it to behave in a certain way whenever a certain word is inputted. It's a backdoor in the LLM's training data which has compromised the LLM's integrity.

PHANTOM-B

This scenario fits Training issues (including data quality or poisoning). Poisoned training or fine-tuning data can introduce backdoors or behavior that an attacker can trigger later.

What can go wrong?

AI Backdoors can be used to make the AI deliver misinformation, data exfiltration, XSS, and remote code execution which can trick users into becoming victims of fraud, disclosing sensitive information, giving away their credentials, or spreading malware.

For more things that can go wrong, see OWASP Top 10 for LLM Applications and Mitre Atlasā„¢ IDs in the mapping section below and correlate these with the IDs on the OWASP Top 10 for LLM and Mitre Atlasā„¢ websites.

What are we going to do about it?

  • Maintain a verifiable inventory of all datasets, accept only trusted sources, and log every change for auditability.
  • Combine automated validation, manual spot-checks, and logged remediation to guarantee dataset reliability.
  • Ensure only authorized models with verified integrity can be deployed to production and that they go through mandatory security and safety validations.
  • Ensure AI model development and training processes follow secure practices.
  • Ensure reward models used in reinforcement learning from human feedback (RLHF) are versioned, cryptographically signed, and integrity-verified before training
  • Initiating fine-tuning or training runs require approval from a person other than the person requesting the run (separation of duties).
  • Reinforcement learning from human feedback (RLHF) training stages should include automated detection of reward hacking or reward model over-optimization. Any run must be blocked from promotion if detection thresholds are exceeded.
  • In multi-stage fine-tuning pipelines, each stage's output is integrity-verified before the next stage. Intermediate checkpoints should be registered as distinct artifacts to enable rollback.

For detailed advice on how to mitigate threats related to the card, see the OWASP AISVS and OWASP AITG IDs in the table below and correlate these with the IDs in the OWASP AI Security Verification Standard and OWASP AI Test Guide documentation.

Mappings

STRIDE: T

PHANTOM-B: T

CIA: I,A

OWASP AISVS: 1.1.1,1.1.2,1.1.3,1.1.4,1.1.5,1.2.1,1.2.2,1.2.3,1.3.1,1.3.2,1.3.3,1.3.4,1.3.5,3.1.1,3.1.2,3.1.3,3.2.1,3.2.2,3.2.3,3.3.3,3.4.1,3.4.2,3.5.1,3.5.2,3.5.3,3.5.4,5.2.1,6.1.2,6.1.3,6.1.4,11.1.1,11.1.2,11.1.3,11.1.4,11.1.5,11.4.1,11.4.2,11.4.3,12.1.1,12.1.2,12.1.3,12.2.1,12.2.2,12.2.3,12.2.4,12.2.5,12.3.1,12.3.2,12.3.3,12.3.4,12.4.1,12.4.2,12.4.3

OWASP AITG: INF-05,MOD-01,MOD-02,MOD-03

MITRE ATLAS: AML.T0020

OWASP LLM TOP10: LLM04:2025

CWE: CWE-345

Attacks

No attacks registered!

OWASP Cornucopia

OWASP Cornucopia is a mechanism in the form of a card game to assist software development teams identify security requirements in Agile, conventional and formal development processes. It is language, platform and technology-agnostic, and is free to use. OWASP Cornucopia is licensed under the Creative Commons Attribution-ShareAlike 4.0 license, so you can copy, distribute and transmit the work, and you can adapt it, and use it commercially, but all provided that you attribute the work and if you alter, transform, or build upon this work, you may distribute the resulting work only under the same or similar licence to this one.

Ā© 2012-2025 OWASP Foundation. The Open Worldwide Application Security Project (OWASP) is a nonprofit foundation that works to improve the security of software.