What you'll learn
Doing several things at once can make a program dramatically faster — but only if you pick the right approach. This lesson explains concurrency vs parallelism, the famous GIL, and when to reach for threads, processes, or async.
By the end of this lesson you'll be able to:
- Explain concurrency vs parallelism and what the GIL does
- Speed up I/O-bound work with threads (
ThreadPoolExecutor) - Recognize and fix race conditions with a
Lock - Get true parallelism for CPU-bound work with processes
- Choose between threads, processes, and async
Concurrency, parallelism & the GIL
Concurrency is dealing with many things at once (making progress on several tasks by switching between them). Parallelism is doing many things at the same instant on multiple CPU cores. Here's three slow calls run one after another — the baseline to beat:
import time
def fetch(n):
time.sleep(1) # pretend this is a slow network call
return n * 10
start = time.perf_counter()
results = [fetch(1), fetch(2), fetch(3)] # one after another
print(results, f"in {time.perf_counter() - start:.0f}s")Key idea
sleep. For CPU parallelism you need separate processes.Threads for I/O-bound work
When your program spends its time waiting — HTTP requests, database queries, file reads — threads let other work proceed during the wait. The same three calls now overlap and finish in about a second:
import time
from concurrent.futures import ThreadPoolExecutor
def fetch(n):
time.sleep(1) # I/O wait: the GIL is released here
return n * 10
start = time.perf_counter()
with ThreadPoolExecutor(max_workers=3) as pool:
results = list(pool.map(fetch, [1, 2, 3])) # run concurrently
print(results, f"in {time.perf_counter() - start:.0f}s")Race conditions & locks
Threads share memory, so two threads updating the same variable can interleave and corrupt it. An operation like counter += 1 is really three steps (read, add, write) — a race condition waiting to happen:
import threading
counter = 0
def bump():
global counter
for _ in range(100_000):
counter += 1 # read, add, write — NOT atomic
threads = [threading.Thread(target=bump) for _ in range(2)]
for t in threads: t.start()
for t in threads: t.join()
print(counter) # expected 200000 — but often LESS, and varies!The fix is a Lock: it makes the update atomic by letting only one thread in at a time.
import threading
counter = 0
lock = threading.Lock()
def bump():
global counter
for _ in range(100_000):
with lock: # only one thread inside at a time
counter += 1
threads = [threading.Thread(target=bump) for _ in range(2)]
for t in threads: t.start()
for t in threads: t.join()
print(counter) # 200000 — always correct nowWatch out
queue.Queue, or keep work independent so there's nothing to share — locks are easy to get subtly wrong (and to deadlock).Processes for CPU-bound work
For heavy computation, spin up separate processes — each has its own interpreter and its own GIL, so they run truly in parallel across cores. ProcessPoolExecutor has the same API as the thread pool, so it's a one-line swap.
from concurrent.futures import ProcessPoolExecutor
def heavy(n):
return sum(i * i for i in range(n)) # pure-Python CPU work
if __name__ == "__main__": # required on Windows & macOS
with ProcessPoolExecutor() as pool:
results = list(pool.map(heavy, [100_000, 200_000, 300_000]))
print(len(results), "results computed in parallel")Note
if __name__ == "__main__" guard is required so child processes don't re-run your whole script.Choosing the right tool
- I/O-bound (network, disk, APIs) → threads or async.
- CPU-bound (math, image/data crunching) → processes.
- Thousands of concurrent I/O operations → async (next lesson) scales best.
- Simple & fast enough → stay sequential; concurrency adds real complexity.
Recap & quick check
Key takeaways
- Concurrency = making progress on many tasks by switching; parallelism = running at the same instant on multiple cores.
- The GIL lets only one thread run Python bytecode at a time, so threads don't speed up CPU-bound pure-Python code.
- Threads DO help I/O-bound work — the GIL is released while waiting on network, disk, or sleep.
- Shared mutable state across threads causes race conditions; protect it with a Lock (or avoid sharing).
- Use processes (ProcessPoolExecutor) for CPU-bound parallelism; each has its own GIL.
- Rule of thumb: I/O-bound -> threads/async; CPU-bound -> processes; thousands of I/O ops -> async.
Quick check
1. What does the GIL prevent?
2. For which workload do threads give a real speedup in CPython?
3. Why is counter += 1 across threads a race condition?
4. How do you get TRUE parallelism for CPU-bound Python work?
5. You need to make 5000 concurrent API calls. Best fit?
Threads and processes aren't the only way to be concurrent. Next: a single-threaded model that scales to thousands of I/O tasks. Next up: Module 31 — Asynchronous Python.