What you'll learn
- Understand AI guardrails and why they matter
- Explore how LLM jailbreaks work
- Recognize common prompt injection techniques
- Identify prompt and data leakage risks
- Understand model poisoning and exposure
- Examine toxicity in AI-generated outputs
- Investigate the causes of hallucinations
- Read and evaluate AI model cards
AI models can do remarkable things. They can also behave in ways their creators never intended.
This course takes you inside some of the most important failure modes in modern AI. You’ll explore guardrails, jailbreaks, prompt injection, leaking, poisoning, toxicity, and hallucinations while seeing how these issues show up in real LLM interactions.
You’ll go beyond definitions, too. Try creating jailbreak prompts, examine multimodal injection, investigate hallucinations through mechanistic interpretability, and compare how different models think and fail.
Finally, you’ll learn how model cards document capabilities, limitations, and risks—and practice reading them with a more critical eye.
By the end, you’ll have a clearer mental model of where AI systems can break, why those failures matter, and how AI safety researchers evaluate them.
I'm a proud lifetime ZTM member. It's changed the trajectory of my life. The projects I built from ZTM courses made me stand out as the #1 candidate and landed me the job. Thanks to Andrei, Yihua, and the entire ZTM team, instructors, and community.
Who You Will Learn With
You're getting more than just a course
Our instructors, TAs, Mentors, Alumni, and fellow students go above and beyond to help guide you and ensure you're on the right path to achieve your goals. Our private ZTM Discord server is a key factor in taking your skills, confidence and career to the next level.



