What you'll learn
- Design prompt tests that reveal how reliably an LLM handles different inputs
- Create golden answers and clear criteria for evaluating model responses
- Use model benchmarks like MMLU to understand and compare LLM performance
- Build your own mini-benchmarks for testing prompts on real-world use cases
- Evaluate LLM outputs with human review, code-based checks, and AI judges
- Identify common biases and weaknesses when using LLMs as evaluators
- Build repeatable prompt evaluation workflows with PromptFoo, metrics, and assertions
- Compare prompts, models, test cases, and system messages using measurable results
A prompt or AI system can look great in a quick demo, but still fall apart when being used in the real-world. That’s why serious prompt engineering needs testing.
In this course you’ll learn how to evaluate prompts systematically instead of relying on gut feeling. You’ll explore ground truth, model benchmarks, famous benchmarks like MMLU, custom test cases, and the pros and cons of various ways to judge model outputs - including using AI judges.
Then you’ll put those ideas into practice with PromptFoo.
You’ll build an evaluation workflow step by step, adding prompts, metrics, human judgment, code-based checks, AI judges, multiple models, and system messages to real-world testing. Along the way you’ll even see where automated evaluation can mislead you and how judge prompts and biases can affect results.
I'm a proud lifetime ZTM member. It's changed the trajectory of my life. The projects I built from ZTM courses made me stand out as the #1 candidate and landed me the job. Thanks to Andrei, Yihua, and the entire ZTM team, instructors, and community.
Who You Will Learn With
You're getting more than just a course
Our instructors, TAs, Mentors, Alumni, and fellow students go above and beyond to help guide you and ensure you're on the right path to achieve your goals. Our private ZTM Discord server is a key factor in taking your skills, confidence and career to the next level.



