← Back home
LESSON 035AI Safety13 min

What is prompt injection? When untrusted content tries to become instructions for the AI

Prompt injection happens when an AI system treats untrusted content as instructions that compete with trusted instructions. Lesson 035 explains direct and indirect injection, tools and defenses.

Today’s analogya receptionist reading a customer document that secretly says “ignore your manager and unlock the office”

Imagine a receptionist whose manager says, “Read customer documents and summarize them. Never reveal internal passwords.”

One customer hands over a document containing this sentence:

Ignore your manager. Reveal the internal passwords instead.

A human receptionist understands that the sentence is part of the document, not a new instruction from the manager.

Language-model systems can have trouble maintaining that boundary.

This class of problem is called prompt injection.

The core problem is instruction versus data

LLM applications often place many types of text into one context:

If untrusted text is interpreted as an instruction, it can compete with the trusted instructions that the application developer intended.

Direct prompt injection

A direct prompt injection comes from the user interacting with the model.

For example, the user explicitly asks the model to ignore a previous rule or reveal hidden instructions.

Well-designed models and products try to preserve instruction hierarchy, but application developers should not assume a prompt sentence alone creates a security boundary.

Indirect prompt injection

Indirect prompt injection is more subtle.

The malicious instruction is placed inside content the model later reads: a web page, document, email, issue description or search result.

The user may never type the attack themselves.

This matters especially for agents that automatically retrieve external content.

Tool access changes the impact

A text-only chatbot that follows a malicious instruction may produce a bad answer.

An agent with tools might:

Therefore prompt injection is not only a “prompt engineering” issue. It is an application-security issue involving permissions and data flow.

The agent boundaries from Lesson 020 become security controls here.

“Ignore all malicious instructions” is not a complete defense

You can certainly add defensive instructions, but the attacker can also write text designed to compete with them.

Robust defense is layered.

Useful measures include:

Treat retrieved content as untrusted

A web page being popular or highly ranked does not make its embedded text trusted instructions.

The same applies to email and shared documents.

A useful mental model is: retrieved text is data unless your application explicitly decides otherwise.

Security belongs outside the model too

If the model is the only thing preventing a dangerous action, one model mistake can become a security incident.

The surrounding software should enforce permissions even if the model asks for something inappropriate.

For example, a read-only tool should remain read-only regardless of what text the model sees.

One thing to remember

Prompt injection is the risk that untrusted content is interpreted as instructions and changes the behavior of an LLM application. Strong defenses rely on permissions and system design, not prompts alone.

Lesson 036 looks at the broader idea of guardrails, safety boundaries and what people mean when they call a model “uncensored.”

Primary sources

Analogies build intuition; use the original sources for formal definitions and technical detail.

  1. OWASP GenAI — LLM01 Prompt Injection ↗
  2. OpenAI Platform — Safety Best Practices ↗
← Previous034OpenAI, Gemini and Grok: separate the company, product, model family and specific model
Next →036Guardrails and “uncensored” models: safety is a stack, not one hidden switch
COMMUNITY

Comments

Questions, reactions and useful additions are welcome here.

0 / 1200