What you'll learn
AI is powerful, which means it can cause real harm — sometimes without anyone intending it. This module is the one every practitioner needs: where bias comes from, how to think about fairness and privacy, and what "AI safety" and "alignment" actually mean.
By the end of this module you'll be able to:
- Explain how bias enters models
- Discuss fairness, privacy, and accountability
- Understand why interpretability matters
- Describe the alignment problem
Bias in, bias out
A model learns whatever patterns are in its data — including society's biases. Train a hiring model on decades of biased decisions and it will faithfully reproduce, even amplify, that discrimination. The model isn't "prejudiced," but its outputs can be, because the data was. This is the single most common ethical failure in deployed ML.
Watch out
Fairness & privacy
Fairness is genuinely hard: there are several mathematical definitions and they can conflict, so you must decide what fairness means for a given system, not just optimise accuracy. Privacy is equally fraught: models can memorise and leak training data, prompts may contain sensitive information, and people rarely consented to their data training an AI. The broad risks are worth seeing together:
Bias & fairness
Models can inherit and amplify historical discrimination in the data.
Privacy
Training and prompts can leak personal data; consent is often unclear.
Transparency
Deep models are opaque — hard to explain a given decision.
Misinformation
Cheap, convincing fake text, images, and video at scale.
Safety & misuse
Powerful tools can be turned to harmful ends.
Accountability
When AI causes harm, who is responsible?
Interpretability & accountability
Deep models are black boxes: they give an answer but not a reason you can easily inspect. In high-stakes settings — loans, medicine, justice — that's a serious problem. The field of interpretability works to explain why a model decided what it did, and accountabilityasks who is responsible when it's wrong. "Because the neural network said so" is not good enough for a denied mortgage or a misdiagnosis.
Misuse & security
The same models that help can harm: deepfakes and mass-produced misinformation, automated scams, and privacy erosion. Models can also be attacked — adversarial examples fool classifiers, and prompt injectioncan hijack an LLM-powered app. Building responsibly means anticipating misuse, not just optimising benchmarks.
The alignment problem
As systems grow more capable, a deeper question arises: how do we ensure powerful AI reliably does what we actually want — including things we forgot to specify? This is the alignment problem. A system optimising a proxy goal can pursue it in unintended, harmful ways. Techniques like RLHF are early steps, and aligning increasingly capable systems with human values is one of the most important open problems in the field.
Recap & quick check
Key takeaways
- Models learn the biases in their data and can amplify discrimination — a leading real-world failure.
- Fairness has multiple, sometimes conflicting definitions; you must choose what it means for your system.
- Privacy risks include memorised training data, sensitive prompts, and lack of consent.
- Deep models are opaque, raising interpretability and accountability concerns in high-stakes uses.
- Alignment — making capable AI reliably do what we intend — is a central open problem.
Quick check
1. Where does model bias most often come from?
2. Why is fairness hard to guarantee?
3. Why does interpretability matter in high-stakes decisions?
4. The 'alignment problem' refers to…
With eyes open to the risks, let's finish by shipping AI in the real world — and mapping where you go next. Next up: Module 32 — Building, Deploying & the Road Ahead.