Phase 6 · Applied & ProfessionalModule 35~48 min read

Data Science Essentials: NumPy, pandas & Matplotlib

Analyze and visualize real data with the scientific Python stack.

What you'll learn

Python is the world's most popular language for data work — and this is why. The NumPy / pandas / Matplotlib stack lets you load, clean, analyze, and visualize real data in a handful of lines, with the heavy lifting done in fast C code.

By the end of this lesson you'll be able to:

  • Explain the scientific Python ecosystem and where Jupyter fits
  • Compute over whole arrays with NumPy (no loops)
  • Load and inspect tabular data with a pandas DataFrame
  • Filter, group, and aggregate data
  • Turn results into a chart with Matplotlib

The scientific Python stack

Four tools do most data work: NumPy (fast numeric arrays), pandas(labeled tables), Matplotlib (plots), and Jupyter (interactive notebooks that mix code, output, and charts). Install them once and you have a full analysis toolkit.

Note

Install the stack with pip install numpy pandas matplotlib jupyter. For exploration, run jupyter lab — notebooks let you run code cell by cell and see tables and plots inline, which is ideal for data work.

NumPy arrays & vectorization

A NumPy array holds numbers in a compact block and applies operations to all of them at once — vectorization. This is both far faster and far more readable than a Python loop.

numpy_intro.py
import numpy as np

a = np.array([1, 2, 3, 4])
print(a * 2)               # [2 4 6 8]  -> vectorized, no Python loop
print(a.mean(), a.sum())   # 2.5 10

# 2D arrays and broadcasting
m = np.array([[1, 2, 3], [4, 5, 6]])
print(m.shape)             # (2, 3)
print(m + 10)              # 10 added to every element

pandas: Series & DataFrame

A DataFrame is a table with labeled columns and rows — think a spreadsheet or SQL table in memory. A single column is a Series. pandas reads CSV, Excel, JSON, and SQL with one call (pd.read_csv(...)), then gives you powerful tools to explore it.

dataframe.py
import pandas as pd

df = pd.DataFrame({
    "name": ["Ada", "Linus", "Grace", "Guido"],
    "lang": ["Python", "C", "COBOL", "Python"],
    "years": [10, 30, 40, 25],
})
print(df.head(2))
print("mean years:", df["years"].mean())   # 26.25

Filtering, grouping & aggregating

Most analysis is three moves: filter rows with a boolean mask, group by a column, and aggregate (sum, mean, count). This is the pandas equivalent of SQL's WHERE and GROUP BY.

wrangle.py
import pandas as pd

df = pd.DataFrame({
    "lang": ["Python", "C", "Python", "C", "Python"],
    "stars": [50, 20, 30, 40, 10],
})

# Filter rows with a boolean mask
print(df[df["stars"] > 25]["lang"].tolist())   # ['Python', 'C', 'Python']

# Group and aggregate
print(df.groupby("lang")["stars"].sum())

Tip

df[df["stars"] > 25] reads oddly at first: the inner part builds a column of True/False, and indexing with it keeps only the True rows. This boolean masking is the heart of pandas filtering.

Plotting with Matplotlib

Matplotlib turns data into charts — line, bar, scatter, histogram, and more. pandas even wraps it, so df.plot() works directly. Here's a basic bar chart:

plot.py
import matplotlib.pyplot as plt

months = ["Jan", "Feb", "Mar", "Apr"]
sales = [120, 150, 170, 140]

plt.bar(months, sales)
plt.title("Quarterly sales")
plt.ylabel("Units")
plt.savefig("sales.png")     # in a notebook, use plt.show() to see it inline

Key idea

The everyday workflow: load data with pandas, clean and reshape it, compute with NumPy/pandas, then visualize with Matplotlib — all inside a Jupyter notebook so you can iterate quickly.

Recap & quick check

Key takeaways

  • The data stack: NumPy (fast arrays), pandas (labeled tables), Matplotlib (plots), Jupyter (notebooks).
  • NumPy vectorizes operations over whole arrays in C — faster and cleaner than Python loops.
  • A pandas DataFrame is an in-memory table; a single column is a Series; read_csv/read_sql load data in one call.
  • Analysis pattern: filter rows (boolean mask), group by a column, aggregate (sum/mean/count).
  • Boolean masking (df[df['x'] > n]) keeps only rows where the condition is True.
  • Matplotlib visualizes results; the loop is load -> clean -> compute -> visualize.

Quick check

1. What does 'vectorization' in NumPy mean?

2. What is a pandas DataFrame?

3. What does df[df['stars'] > 25] return?

4. Which method groups rows and lets you aggregate per group?

5. Which tool is best for interactive, cell-by-cell data exploration?

Data in hand, let's put Python on the web — consuming and building APIs. Next up: Module 36 — Web Development: HTTP, APIs & Frameworks.