Phase 6 · Applied & ProfessionalModule 38~38 min read

Automation, Scripting & Web Scraping

Put Python to work: automate tasks and responsibly scrape the web.

What you'll learn

Python is the duct tape of the programming world. This lesson turns it into a personal automation assistant: renaming and moving files, running as a scheduled job, downloading data, and scraping the web — responsibly.

By the end of this lesson you'll be able to:

  • Automate file and folder chores with pathlib and shutil
  • Package logic into a reusable script and schedule it to run automatically
  • Download files and data with requests
  • Extract data from HTML with BeautifulSoup and CSS selectors
  • Scrape ethically — respecting robots.txt, rate limits, and terms of service

Automating file tasks

Repetitive file chores are the perfect first automation. pathlib handles paths across operating systems, .glob() finds files by pattern, and shutil copies and moves them.

organize.py
from pathlib import Path
import shutil

downloads = Path.home() / "Downloads"
archive = Path.home() / "Archive"
archive.mkdir(exist_ok=True)

# Move every PDF from Downloads into an Archive folder
for pdf in downloads.glob("*.pdf"):
    shutil.move(pdf, archive / pdf.name)
    print("archived", pdf.name)

Tip

Test destructive automation on a copy first, and print what would happen before actually moving or deleting anything (a "dry run"). A loop that mangles hundreds of files is hard to undo.

Scripts & scheduling

Wrap your automation in a script with an argparse interface (Module 29) so it takes options instead of hard-coded paths. Then schedule it: cron on Linux/macOS, Task Scheduler on Windows, or a library like schedule / APScheduler for in-process timing.

Note

A good automation script is idempotent — safe to run repeatedly. Check whether work is already done (does the archive file exist?) before doing it, so a re-run never duplicates or corrupts anything.

Downloading data

Before scraping HTML, remember that a lot of data is available directly — CSV, JSON, or an official API. requests downloads any of it; use .raise_for_status() so failures are loud, not silent.

download.py
import requests
from pathlib import Path

url = "https://example.com/report.csv"
resp = requests.get(url, timeout=30)
resp.raise_for_status()               # raise on 4xx / 5xx

Path("report.csv").write_bytes(resp.content)
print("saved", len(resp.content), "bytes")

Scraping with BeautifulSoup

When there's no API, BeautifulSoup parses HTML into a searchable tree. Select elements with CSS selectors (the same ones you learned for styling) and pull out text and attributes:

parse.py
from bs4 import BeautifulSoup

html = """
<ul id="fruit">
  <li class="item">Apples</li>
  <li class="item">Bananas</li>
</ul>
"""

soup = BeautifulSoup(html, "html.parser")
items = [li.get_text() for li in soup.select("li.item")]   # CSS selectors
print(items)                          # ['Apples', 'Bananas']

In practice you fetch the page with requests, then hand the HTML to BeautifulSoup:

scrape.py
import requests
from bs4 import BeautifulSoup

resp = requests.get("https://example.com", timeout=30,
                    headers={"User-Agent": "my-scraper/1.0"})
soup = BeautifulSoup(resp.text, "html.parser")

print(soup.title.get_text())          # the page <title>
for link in soup.select("a")[:3]:     # first 3 links
    print(link.get_text(strip=True), "->", link.get("href"))

Scraping responsibly

Scraping touches someone else's servers and data, so do it considerately:

  • Prefer an official API when one exists — it's stabler and sanctioned.
  • Check robots.txt and the site's terms of service.
  • Rate-limit yourself — add delays; don't hammer a server with rapid requests.
  • Identify yourself with a descriptive User-Agent.
  • Respect the data — copyright and privacy still apply to what you collect.

Watch out

Aggressive scraping can overload a site, get your IP banned, or cross legal lines. Cache pages you've already fetched, throttle your requests, and stop if a site asks you not to. Being a good citizen keeps the open web open.

Recap & quick check

Key takeaways

  • pathlib (+ .glob) and shutil automate cross-platform file and folder tasks.
  • Wrap automation in an argparse script and schedule it with cron, Task Scheduler, or a scheduling library.
  • Make automation idempotent (safe to re-run) and dry-run destructive steps first.
  • requests downloads files/data; call .raise_for_status() so HTTP errors aren't silent.
  • BeautifulSoup parses HTML into a tree you query with CSS selectors to extract text and attributes.
  • Scrape ethically: prefer APIs, honor robots.txt and ToS, rate-limit, set a User-Agent, respect copyright/privacy.

Quick check

1. Which pair automates cross-platform file tasks?

2. What does an 'idempotent' automation script mean?

3. Why call resp.raise_for_status()?

4. What does BeautifulSoup do?

5. What should you do before scraping a site?

Your scripts and tools deserve to be shared. Next: turning a project into an installable package. Next up: Module 39 — Packaging, Distribution & Deployment.