Phase 5 · Advanced PythonModule 34~36 min read

Performance, Memory & CPython Internals

Understand how CPython works and make Python code fast.

What you'll learn

To make Python fast, it helps to know what it's doing under the hood. This lesson peeks inside CPython — bytecode, memory, garbage collection — then turns that understanding into concrete speed and memory wins.

By the end of this lesson you'll be able to:

  • Describe how CPython compiles and runs your code as bytecode
  • Explain reference counting, copies, and garbage collection
  • Stream large data with generators to keep memory flat
  • Speed code up with caching, built-ins, and vectorized libraries

How CPython runs your code

CPython (the reference interpreter) compiles your source to bytecode — compact instructions for a virtual machine — then executes them one by one. The dis module lets you see them, which demystifies what a line actually costs.

bytecode.py
import dis

def add(a, b):
    return a + b

dis.dis(add)
Exact opcodes and offsets vary by Python version.

Note

Because each line becomes several interpreted instructions, pure-Python loops are relatively slow. The fix isn't "write cleverer loops" — it's to do less interpreting: push work into built-ins and C-backed libraries (below).

References, copies & GC

Variables are names bound to objects, not boxes holding values. Assigning one name to another makes both point at the same object — a frequent source of "why did my other list change?" bugs.

refs.py
a = [1, 2, 3]
b = a                 # b points to the SAME list, not a copy
print(a is b)         # True
b.append(4)
print(a)              # [1, 2, 3, 4] -> both names see the change

c = a.copy()          # a real, independent copy
print(a is c)         # False

Key idea

CPython frees an object the moment its reference count hits zero. A separate garbage collector exists only to clean up reference cycles (objects that refer to each other). This is why you rarely think about memory — but also why holding references (in caches or global lists) keeps objects alive.

Generators & streaming

The single biggest memory win is not holding everything at once. A list of a million results occupies megabytes; a generator computes them lazily and uses a tiny, constant amount of memory.

streaming.py
import sys

# A list materializes every value at once (lots of memory)
squares_list = [x * x for x in range(1_000_000)]
print(sys.getsizeof(squares_list), "bytes  <- the list")

# A generator holds ONE value at a time (tiny, constant memory)
squares_gen = (x * x for x in range(1_000_000))
print(sys.getsizeof(squares_gen), "bytes  <- the generator")

Tip

Prefer generators and itertools for pipelines over large data — read a file line by line, transform lazily, and never load the whole thing. Reach for a list only when you truly need random access or multiple passes.

Making Python fast

Once you've profiled (last lesson) and found the hot spot, these are the highest-leverage fixes:

  • Cache repeated work with functools.lru_cache.
  • Use built-ins (sum, sorted, "".join) — they run in C.
  • Pick the right structure — a set for membership, a dict for lookups.
  • Vectorize heavy numeric work with NumPy.
cache.py
from functools import lru_cache

@lru_cache(maxsize=None)          # remember results, skip repeat work
def fib(n):
    return n if n < 2 else fib(n - 1) + fib(n - 2)

print(fib(50))                    # instant — without caching this is millions of calls
vectorize.py
# Pure Python: a loop (or comprehension) over a million numbers
data = list(range(1_000_000))
doubled = [x * 2 for x in data]        # interpreted, one element at a time

# NumPy: one vectorized operation that runs in C
import numpy as np
arr = np.arange(1_000_000)
doubled = arr * 2                       # far faster, far less memory
print(doubled[:5])

Note

When even that isn't enough, tools like Cython, Numba, or a C extension compile hot functions to machine code. Reach for them last — a better algorithm or a vectorized library usually wins with far less effort.

Recap & quick check

Key takeaways

  • CPython compiles source to bytecode and interprets it; dis shows the instructions (they vary by version).
  • Interpreted per-instruction execution makes pure-Python loops slow — push work into built-ins and C libraries.
  • Variables are names bound to objects; assignment shares the object, .copy() makes a separate one.
  • Objects are freed when their reference count hits zero; the GC only collects reference cycles.
  • Generators stream values lazily in constant memory — the biggest win for large data.
  • Speed up hot paths with lru_cache, built-ins, the right data structure, and NumPy vectorization.

Quick check

1. What does CPython execute after compiling your source?

2. After b = a (both lists), what does b.append(4) do to a?

3. When does CPython free an ordinary object?

4. Why use a generator instead of a list for a million results?

5. What's a high-leverage way to speed up heavy numeric loops?

That completes Phase 5 — you now write Python that is typed, tested, concurrent, and fast. Next we put it all to work on real data. Next up: Module 35 — Data Science Essentials: NumPy, pandas & Matplotlib.