What you'll learn
A regular expression (regex) is a tiny language for describing patterns in text — "four digits, a dash, two digits" or "anything that looks like an email." With Python's re module you can search, validate, extract, and rewrite text in a single expression.
By the end of this lesson you'll be able to:
- Use
re.search,re.match, andre.findall - Build patterns from character classes, quantifiers, and anchors
- Capture parts of a match with groups (including named groups)
- Rewrite text with
re.suband compile patterns for reuse
The re module
Three functions do most of the work. search finds the first match anywhere; findall returns every match as a list; match only checks the start of the string. A successful search returns a match object — call .group() to get the matched text.
import re
text = "Contact: ada@math.org and grace@navy.mil"
# search -> first match anywhere, or None
m = re.search(r"\w+@\w+\.\w+", text)
print(m.group()) # ada@math.org
# findall -> every match, as a list
print(re.findall(r"\w+@\w+\.\w+", text))
# match -> only anchored at the START of the string
print(bool(re.match(r"Contact", text))) # True
print(bool(re.match(r"ada", text))) # FalseTip
r"\d+", not "\d+". Regex uses backslashes heavily, and raw strings stop Python from treating them as escape sequences before re ever sees them.Character classes & quantifiers
Character classes say what to match (\d a digit, \w a word character, [a-z] your own set) and quantifiers say how many (* zero-or-more, + one-or-more, ? optional, {n,m} a range).
import re
# \d digit \w word char \s whitespace . any char
# quantifiers: * 0+ + 1+ ? 0 or 1 {n,m} a range
print(re.findall(r"\d+", "in 2024 we shipped 3 apps")) # ['2024', '3']
print(re.findall(r"[aeiou]", "hello")) # ['e', 'o']
print(re.findall(r"[a-z]+", "Big DEAL now")) # ['ig', 'now']
print(re.findall(r"\d{3}-\d{4}", "call 555-0142 now")) # ['555-0142']By default quantifiers are greedy — they match as much as possible. Add a ? to make them lazy, matching as little as possible. This single character is the fix for a huge share of "my regex grabbed too much" bugs:
import re
html = "<b>bold</b> and <i>italic</i>"
# Greedy: * grabs as MUCH as possible
print(re.findall(r"<.*>", html)) # ['<b>bold</b> and <i>italic</i>']
# Lazy: *? grabs as LITTLE as possible
print(re.findall(r"<.*?>", html)) # ['<b>', '</b>', '<i>', '</i>']Anchors & boundaries
Anchors match positions, not characters. ^ and $ pin a pattern to the start and end of the string — combine them to validate an entire value. \b matches a word boundary, so you can find whole words only.
import re
# ^ start of string $ end of string \b word boundary
print(bool(re.search(r"^\d{4}$", "2024"))) # True (exactly 4 digits)
print(bool(re.search(r"^\d{4}$", "2024!"))) # False
# \b matches WHOLE words only
print(re.findall(r"\bcat\b", "cat category cats cat")) # ['cat', 'cat']Groups & capturing
Parentheses capture part of a match so you can pull it out afterwards. .group(0) is the whole match, .group(1) the first parenthesized group, and .groups() returns them all. Naming groups with (?P<name>...) makes the code self-documenting.
import re
date = "2024-08-26"
m = re.search(r"(\d{4})-(\d{2})-(\d{2})", date)
print(m.group(0)) # 2024-08-26 (the whole match)
print(m.group(1)) # 2024 (first group)
print(m.groups()) # ('2024', '08', '26')
# Named groups read far better:
m = re.search(r"(?P<year>\d{4})-(?P<month>\d{2})", date)
print(m.group("year"), m.group("month")) # 2024 08Substitution & compiling
re.sub replaces matches — and the replacement can reference captured groups with \1, \2, and so on. When you use a pattern repeatedly, re.compile it once into a reusable pattern object.
import re
# re.sub(pattern, replacement, text)
print(re.sub(r"\s+", " ", "too many spaces")) # collapse whitespace
# \1, \2 reference captured groups in the replacement
print(re.sub(r"(\w+)@(\w+)", r"\2::\1", "ada@math")) # math::ada
# Compile once, reuse many times (clearer & faster in loops)
digits = re.compile(r"\d+")
print(digits.findall("a1b22c333")) # ['1', '22', '333']Watch out
(a+)+ on hostile input, which can cause catastrophic backtracking.Recap & quick check
Key takeaways
- re.search finds the first match anywhere; findall returns all matches; match only checks the string's start.
- Always write patterns as raw strings: r"\d+".
- Classes (\d, \w, [a-z]) say what to match; quantifiers (*, +, ?, {n,m}) say how many.
- Quantifiers are greedy by default; add ? (e.g. .*?) to make them lazy.
- ^ and $ anchor to start/end; \b matches word boundaries; parentheses capture groups.
- re.sub rewrites text (use \1 for groups); re.compile a pattern you reuse.
Quick check
1. What's the difference between re.search and re.match?
2. Why write regex patterns as raw strings (r"...")?
3. What does .*? mean compared to .* ?
4. What does the pattern ^\d{4}$ match?
5. How do you reference the first captured group in a re.sub replacement?
Next we leave text behind for another everyday need: working with time, precise numbers, and randomness. Next up: Module 28 — Dates, Times, Math & Randomness.