Phase 5 · Advanced PythonModule 27~34 min read

Regular Expressions

Match, validate, and transform text with the re module.

What you'll learn

A regular expression (regex) is a tiny language for describing patterns in text — "four digits, a dash, two digits" or "anything that looks like an email." With Python's re module you can search, validate, extract, and rewrite text in a single expression.

By the end of this lesson you'll be able to:

  • Use re.search, re.match, and re.findall
  • Build patterns from character classes, quantifiers, and anchors
  • Capture parts of a match with groups (including named groups)
  • Rewrite text with re.sub and compile patterns for reuse

The re module

Three functions do most of the work. search finds the first match anywhere; findall returns every match as a list; match only checks the start of the string. A successful search returns a match object — call .group() to get the matched text.

basics.py
import re

text = "Contact: ada@math.org and grace@navy.mil"

# search -> first match anywhere, or None
m = re.search(r"\w+@\w+\.\w+", text)
print(m.group())                        # ada@math.org

# findall -> every match, as a list
print(re.findall(r"\w+@\w+\.\w+", text))

# match -> only anchored at the START of the string
print(bool(re.match(r"Contact", text))) # True
print(bool(re.match(r"ada", text)))     # False

Tip

Always write patterns as raw strings — r"\d+", not "\d+". Regex uses backslashes heavily, and raw strings stop Python from treating them as escape sequences before re ever sees them.

Character classes & quantifiers

Character classes say what to match (\d a digit, \w a word character, [a-z] your own set) and quantifiers say how many (* zero-or-more, + one-or-more, ? optional, {n,m} a range).

classes.py
import re

# \d digit   \w word char   \s whitespace   . any char
# quantifiers:  *  0+     +  1+     ?  0 or 1     {n,m}  a range
print(re.findall(r"\d+", "in 2024 we shipped 3 apps"))  # ['2024', '3']
print(re.findall(r"[aeiou]", "hello"))                   # ['e', 'o']
print(re.findall(r"[a-z]+", "Big DEAL now"))             # ['ig', 'now']
print(re.findall(r"\d{3}-\d{4}", "call 555-0142 now"))  # ['555-0142']

By default quantifiers are greedy — they match as much as possible. Add a ? to make them lazy, matching as little as possible. This single character is the fix for a huge share of "my regex grabbed too much" bugs:

greedy.py
import re

html = "<b>bold</b> and <i>italic</i>"

# Greedy: * grabs as MUCH as possible
print(re.findall(r"<.*>", html))    # ['<b>bold</b> and <i>italic</i>']

# Lazy: *? grabs as LITTLE as possible
print(re.findall(r"<.*?>", html))   # ['<b>', '</b>', '<i>', '</i>']

Anchors & boundaries

Anchors match positions, not characters. ^ and $ pin a pattern to the start and end of the string — combine them to validate an entire value. \b matches a word boundary, so you can find whole words only.

anchors.py
import re

# ^ start of string   $ end of string   \b word boundary
print(bool(re.search(r"^\d{4}$", "2024")))    # True  (exactly 4 digits)
print(bool(re.search(r"^\d{4}$", "2024!")))   # False

# \b matches WHOLE words only
print(re.findall(r"\bcat\b", "cat category cats cat"))  # ['cat', 'cat']

Groups & capturing

Parentheses capture part of a match so you can pull it out afterwards. .group(0) is the whole match, .group(1) the first parenthesized group, and .groups() returns them all. Naming groups with (?P<name>...) makes the code self-documenting.

groups.py
import re

date = "2024-08-26"
m = re.search(r"(\d{4})-(\d{2})-(\d{2})", date)
print(m.group(0))     # 2024-08-26  (the whole match)
print(m.group(1))     # 2024        (first group)
print(m.groups())     # ('2024', '08', '26')

# Named groups read far better:
m = re.search(r"(?P<year>\d{4})-(?P<month>\d{2})", date)
print(m.group("year"), m.group("month"))   # 2024 08

Substitution & compiling

re.sub replaces matches — and the replacement can reference captured groups with \1, \2, and so on. When you use a pattern repeatedly, re.compile it once into a reusable pattern object.

sub.py
import re

# re.sub(pattern, replacement, text)
print(re.sub(r"\s+", " ", "too    many     spaces"))  # collapse whitespace

# \1, \2 reference captured groups in the replacement
print(re.sub(r"(\w+)@(\w+)", r"\2::\1", "ada@math"))  # math::ada

# Compile once, reuse many times (clearer & faster in loops)
digits = re.compile(r"\d+")
print(digits.findall("a1b22c333"))                     # ['1', '22', '333']

Watch out

Regex is perfect for tokens (dates, phone numbers, IDs) but a poor fit for deeply nested formats like HTML or JSON — use a real parser there. And beware patterns like (a+)+ on hostile input, which can cause catastrophic backtracking.

Recap & quick check

Key takeaways

  • re.search finds the first match anywhere; findall returns all matches; match only checks the string's start.
  • Always write patterns as raw strings: r"\d+".
  • Classes (\d, \w, [a-z]) say what to match; quantifiers (*, +, ?, {n,m}) say how many.
  • Quantifiers are greedy by default; add ? (e.g. .*?) to make them lazy.
  • ^ and $ anchor to start/end; \b matches word boundaries; parentheses capture groups.
  • re.sub rewrites text (use \1 for groups); re.compile a pattern you reuse.

Quick check

1. What's the difference between re.search and re.match?

2. Why write regex patterns as raw strings (r"...")?

3. What does .*? mean compared to .* ?

4. What does the pattern ^\d{4}$ match?

5. How do you reference the first captured group in a re.sub replacement?

Next we leave text behind for another everyday need: working with time, precise numbers, and randomness. Next up: Module 28 — Dates, Times, Math & Randomness.