What you'll learn
Python is the world's most popular language for data work — and this is why. The NumPy / pandas / Matplotlib stack lets you load, clean, analyze, and visualize real data in a handful of lines, with the heavy lifting done in fast C code.
By the end of this lesson you'll be able to:
- Explain the scientific Python ecosystem and where Jupyter fits
- Compute over whole arrays with NumPy (no loops)
- Load and inspect tabular data with a pandas
DataFrame - Filter, group, and aggregate data
- Turn results into a chart with Matplotlib
The scientific Python stack
Four tools do most data work: NumPy (fast numeric arrays), pandas(labeled tables), Matplotlib (plots), and Jupyter (interactive notebooks that mix code, output, and charts). Install them once and you have a full analysis toolkit.
Note
pip install numpy pandas matplotlib jupyter. For exploration, run jupyter lab — notebooks let you run code cell by cell and see tables and plots inline, which is ideal for data work.NumPy arrays & vectorization
A NumPy array holds numbers in a compact block and applies operations to all of them at once — vectorization. This is both far faster and far more readable than a Python loop.
import numpy as np
a = np.array([1, 2, 3, 4])
print(a * 2) # [2 4 6 8] -> vectorized, no Python loop
print(a.mean(), a.sum()) # 2.5 10
# 2D arrays and broadcasting
m = np.array([[1, 2, 3], [4, 5, 6]])
print(m.shape) # (2, 3)
print(m + 10) # 10 added to every elementpandas: Series & DataFrame
A DataFrame is a table with labeled columns and rows — think a spreadsheet or SQL table in memory. A single column is a Series. pandas reads CSV, Excel, JSON, and SQL with one call (pd.read_csv(...)), then gives you powerful tools to explore it.
import pandas as pd
df = pd.DataFrame({
"name": ["Ada", "Linus", "Grace", "Guido"],
"lang": ["Python", "C", "COBOL", "Python"],
"years": [10, 30, 40, 25],
})
print(df.head(2))
print("mean years:", df["years"].mean()) # 26.25Filtering, grouping & aggregating
Most analysis is three moves: filter rows with a boolean mask, group by a column, and aggregate (sum, mean, count). This is the pandas equivalent of SQL's WHERE and GROUP BY.
import pandas as pd
df = pd.DataFrame({
"lang": ["Python", "C", "Python", "C", "Python"],
"stars": [50, 20, 30, 40, 10],
})
# Filter rows with a boolean mask
print(df[df["stars"] > 25]["lang"].tolist()) # ['Python', 'C', 'Python']
# Group and aggregate
print(df.groupby("lang")["stars"].sum())Tip
df[df["stars"] > 25] reads oddly at first: the inner part builds a column of True/False, and indexing with it keeps only the True rows. This boolean masking is the heart of pandas filtering.Plotting with Matplotlib
Matplotlib turns data into charts — line, bar, scatter, histogram, and more. pandas even wraps it, so df.plot() works directly. Here's a basic bar chart:
import matplotlib.pyplot as plt
months = ["Jan", "Feb", "Mar", "Apr"]
sales = [120, 150, 170, 140]
plt.bar(months, sales)
plt.title("Quarterly sales")
plt.ylabel("Units")
plt.savefig("sales.png") # in a notebook, use plt.show() to see it inlineKey idea
Recap & quick check
Key takeaways
- The data stack: NumPy (fast arrays), pandas (labeled tables), Matplotlib (plots), Jupyter (notebooks).
- NumPy vectorizes operations over whole arrays in C — faster and cleaner than Python loops.
- A pandas DataFrame is an in-memory table; a single column is a Series; read_csv/read_sql load data in one call.
- Analysis pattern: filter rows (boolean mask), group by a column, aggregate (sum/mean/count).
- Boolean masking (df[df['x'] > n]) keeps only rows where the condition is True.
- Matplotlib visualizes results; the loop is load -> clean -> compute -> visualize.
Quick check
1. What does 'vectorization' in NumPy mean?
2. What is a pandas DataFrame?
3. What does df[df['stars'] > 25] return?
4. Which method groups rows and lets you aggregate per group?
5. Which tool is best for interactive, cell-by-cell data exploration?
Data in hand, let's put Python on the web — consuming and building APIs. Next up: Module 36 — Web Development: HTTP, APIs & Frameworks.