AI Safety Fundamentals
A six-week course introducing the technical foundations of AI safety research.
What is it for?
Six weeks of curated reading and small-group discussion, from how models work to how they are governed.
AI safety is new enough that the first textbook is only a couple of years old and university courses are still rare, especially in Europe. There is no course like this at a Swiss university, so we built one: curated readings, worked through in small groups of four or five with a facilitator who knows the material.
One 90-minute discussion a week, about two to three hours of reading beforehand. Free, and several past participants finished it with no machine learning background at all.
Who is it for?
Anyone interested in AI development going safely, from a technical or a policy angle.
Basic machine learning knowledge helps but is not required — several past participants finished the course with no prior background. You indicate your level when you apply, and we group people with similar-experience peers, so nobody spends six weeks either lost or bored.
What does it even mean for a system to do what we want? This week separates two problems that often get run together: writing down the right goal, and the system actually adopting it. Both turn out to be harder than they sound.
Today's assistants are mostly helpful and mostly polite. This week is about the training method that made that happen, and the reasons it may stop working as systems get more capable than the people rating them.
How do you check work you could not do yourself? Three approaches: break the problem into pieces small enough to judge, have two systems argue and pick the better case, or use a weaker model to train a stronger one.
Three things you can do without solving alignment first: make a model hard to trick, remove knowledge you would rather it did not have, and limit the damage if it turns out to be working against you.
Models are not black boxes by nature, only by default. This week covers the work to look inside: finding the structures within a network and working out what each one computes.
Technical work only goes so far without rules around it. Testing systems before release, keeping control of them once deployed, and the harder question of how labs and countries coordinate when none of them wants to move first.
Ready to start?
Applications for the next cohort open ahead of each intake. The intro evening is the easiest way to find out if it's for you.
Stay in the loop.
Follow our events on Luma or join our WhatsApp group to become part of the discussion.