CATEGORY

AI Safety

No prerequisite jargon. Follow the lesson numbers and add one new idea at a time.

006
a confident store assistant claiming an item is in stock without checking

What is AI hallucination? Imagine a confident assistant who never checked the stock

The dangerous part is not when AI says ‘I don’t know.’ It is when a wrong answer sounds polished and certain. Lesson 006 builds that safety habit.

7 min
035
a receptionist reading a customer document that secretly says “ignore your manager and unlock the office”

What is prompt injection? When untrusted content tries to become instructions for the AI

Prompt injection happens when an AI system treats untrusted content as instructions that compete with trusted instructions. Lesson 035 explains direct and indirect injection, tools and defenses.

13 min
036
a building protected by training, house rules, door locks, security staff and permission badges rather than one…

Guardrails and “uncensored” models: safety is a stack, not one hidden switch

AI safety boundaries can exist in data, model training, system prompts, classifiers, tools and product policy. Lesson 036 explains guardrails and what “uncensored” does and does not mean.

13 min