AI Safety
No prerequisite jargon. Follow the lesson numbers and add one new idea at a time.
006
a confident store assistant claiming an item is in stock without checking
7 minWhat is AI hallucination? Imagine a confident assistant who never checked the stock
The dangerous part is not when AI says ‘I don’t know.’ It is when a wrong answer sounds polished and certain. Lesson 006 builds that safety habit.
035
a receptionist reading a customer document that secretly says “ignore your manager and unlock the office”
13 minWhat is prompt injection? When untrusted content tries to become instructions for the AI
Prompt injection happens when an AI system treats untrusted content as instructions that compete with trusted instructions. Lesson 035 explains direct and indirect injection, tools and defenses.
036
a building protected by training, house rules, door locks, security staff and permission badges rather than one…
13 minGuardrails and “uncensored” models: safety is a stack, not one hidden switch
AI safety boundaries can exist in data, model training, system prompts, classifiers, tools and product policy. Lesson 036 explains guardrails and what “uncensored” does and does not mean.