Phase 5 · Advanced PythonModule 30~40 min read

Concurrency: Threads, Processes & the GIL

Do many things at once, and understand Python's famous Global Interpreter Lock.

What you'll learn

Doing several things at once can make a program dramatically faster — but only if you pick the right approach. This lesson explains concurrency vs parallelism, the famous GIL, and when to reach for threads, processes, or async.

By the end of this lesson you'll be able to:

  • Explain concurrency vs parallelism and what the GIL does
  • Speed up I/O-bound work with threads (ThreadPoolExecutor)
  • Recognize and fix race conditions with a Lock
  • Get true parallelism for CPU-bound work with processes
  • Choose between threads, processes, and async

Concurrency, parallelism & the GIL

Concurrency is dealing with many things at once (making progress on several tasks by switching between them). Parallelism is doing many things at the same instant on multiple CPU cores. Here's three slow calls run one after another — the baseline to beat:

sequential.py
import time

def fetch(n):
    time.sleep(1)          # pretend this is a slow network call
    return n * 10

start = time.perf_counter()
results = [fetch(1), fetch(2), fetch(3)]   # one after another
print(results, f"in {time.perf_counter() - start:.0f}s")

Key idea

CPython has a Global Interpreter Lock (GIL): only one thread runs Python bytecode at a time. So threads give you no speedup for pure-Python CPU work — but they do help I/O-bound work, because the GIL is released while a thread waits on the network, disk, or sleep. For CPU parallelism you need separate processes.

Threads for I/O-bound work

When your program spends its time waiting — HTTP requests, database queries, file reads — threads let other work proceed during the wait. The same three calls now overlap and finish in about a second:

threads.py
import time
from concurrent.futures import ThreadPoolExecutor

def fetch(n):
    time.sleep(1)          # I/O wait: the GIL is released here
    return n * 10

start = time.perf_counter()
with ThreadPoolExecutor(max_workers=3) as pool:
    results = list(pool.map(fetch, [1, 2, 3]))   # run concurrently
print(results, f"in {time.perf_counter() - start:.0f}s")

Race conditions & locks

Threads share memory, so two threads updating the same variable can interleave and corrupt it. An operation like counter += 1 is really three steps (read, add, write) — a race condition waiting to happen:

race.py
import threading

counter = 0

def bump():
    global counter
    for _ in range(100_000):
        counter += 1        # read, add, write — NOT atomic

threads = [threading.Thread(target=bump) for _ in range(2)]
for t in threads: t.start()
for t in threads: t.join()

print(counter)   # expected 200000 — but often LESS, and varies!

The fix is a Lock: it makes the update atomic by letting only one thread in at a time.

lock.py
import threading

counter = 0
lock = threading.Lock()

def bump():
    global counter
    for _ in range(100_000):
        with lock:          # only one thread inside at a time
            counter += 1

threads = [threading.Thread(target=bump) for _ in range(2)]
for t in threads: t.start()
for t in threads: t.join()

print(counter)   # 200000 — always correct now

Watch out

Any state shared between threads needs protection. Prefer message-passing with queue.Queue, or keep work independent so there's nothing to share — locks are easy to get subtly wrong (and to deadlock).

Processes for CPU-bound work

For heavy computation, spin up separate processes — each has its own interpreter and its own GIL, so they run truly in parallel across cores. ProcessPoolExecutor has the same API as the thread pool, so it's a one-line swap.

processes.py
from concurrent.futures import ProcessPoolExecutor

def heavy(n):
    return sum(i * i for i in range(n))   # pure-Python CPU work

if __name__ == "__main__":                # required on Windows & macOS
    with ProcessPoolExecutor() as pool:
        results = list(pool.map(heavy, [100_000, 200_000, 300_000]))
    print(len(results), "results computed in parallel")

Note

Processes don't share memory, so arguments and results are pickled and copied between them. That overhead means processes pay off for genuinely heavy work, not tiny tasks. The if __name__ == "__main__" guard is required so child processes don't re-run your whole script.

Choosing the right tool

  • I/O-bound (network, disk, APIs) → threads or async.
  • CPU-bound (math, image/data crunching) → processes.
  • Thousands of concurrent I/O operations → async (next lesson) scales best.
  • Simple & fast enough → stay sequential; concurrency adds real complexity.

Recap & quick check

Key takeaways

  • Concurrency = making progress on many tasks by switching; parallelism = running at the same instant on multiple cores.
  • The GIL lets only one thread run Python bytecode at a time, so threads don't speed up CPU-bound pure-Python code.
  • Threads DO help I/O-bound work — the GIL is released while waiting on network, disk, or sleep.
  • Shared mutable state across threads causes race conditions; protect it with a Lock (or avoid sharing).
  • Use processes (ProcessPoolExecutor) for CPU-bound parallelism; each has its own GIL.
  • Rule of thumb: I/O-bound -> threads/async; CPU-bound -> processes; thousands of I/O ops -> async.

Quick check

1. What does the GIL prevent?

2. For which workload do threads give a real speedup in CPython?

3. Why is counter += 1 across threads a race condition?

4. How do you get TRUE parallelism for CPU-bound Python work?

5. You need to make 5000 concurrent API calls. Best fit?

Threads and processes aren't the only way to be concurrent. Next: a single-threaded model that scales to thousands of I/O tasks. Next up: Module 31 — Asynchronous Python.