Phase 3 · Core JavaModule 18~46 min read

Stream API

Process data declaratively with streams: filter, map, reduce, collect, group, and go parallel.

What you'll learn

The Stream API lets you process collections in a declarative, readable pipeline — say what you want (filter, transform, group) rather than writing loops that spell out how. Combined with lambdas, it's one of the most powerful features in modern Java.

By the end you'll be able to:

  • Build stream pipelines: source → intermediate ops → terminal op
  • Transform data with filter, map, flatMap, sorted
  • Produce results with collect, reduce, count, and more
  • Group and summarise data with Collectors
  • Use numeric and parallel streams, and avoid common pitfalls

The stream pipeline

A stream is not a data structure — it's a flow of elements you run operations on. A pipeline has three parts: a source (usually collection.stream()), any number of intermediate operations that transform the flow (and return another stream), and exactly one terminal operation that produces a result:

A stream pipeline
List / source→.stream()→.filter(...)→.map(...)→.collect(...)
Streams.java
import java.util.*;
import java.util.stream.*;

List<String> names = List.of("Sara", "Omar", "Alice", "Bob", "Sam");

List<String> result = names.stream()
    .filter(n -> n.length() > 3)     // keep names longer than 3
    .map(String::toUpperCase)        // transform each
    .sorted()                        // order them
    .collect(Collectors.toList());   // gather into a List

System.out.println(result);

Key idea

Read that pipeline top to bottom like a sentence: "take the names, keep the long ones, upper-case them, sort them, collect them into a list." No index, no loop, no temporary variables — just intent.

Intermediate operations

Intermediate operations are lazy — they don't run until a terminal operation is attached. Each returns a new stream, so they chain:

OperationWhat it does
filter(pred)keep only elements matching the predicate
map(fn)transform each element
flatMap(fn)flatten nested streams into one
sorted()sort (natural order or a Comparator)
distinct()remove duplicates
limit(n) / skip(n)take the first n / skip the first n

flatMap deserves a special mention — it's how you turn a stream of collections into a single flat stream:

FlatMap.java
import java.util.*;
import java.util.stream.*;

List<List<Integer>> nested = List.of(
    List.of(1, 2), List.of(3, 4), List.of(5)
);

List<Integer> flat = nested.stream()
    .flatMap(List::stream)     // flatten streams-of-streams into one
    .collect(Collectors.toList());

System.out.println(flat);

Terminal operations

A terminal operation ends the pipeline and produces a value (or a side effect). Common ones are collect (gather into a collection), forEach (do something with each), count, reduce (combine into one value), and min/max:

Terminal.java
import java.util.*;

List<Integer> nums = List.of(1, 2, 3, 4, 5, 6);

long evens = nums.stream().filter(n -> n % 2 == 0).count();
int sum    = nums.stream().mapToInt(Integer::intValue).sum();
Optional<Integer> max = nums.stream().max(Integer::compareTo);

System.out.println("Evens: " + evens);
System.out.println("Sum: " + sum);
System.out.println("Max: " + max.get());

Watch out

A stream can be used once. After a terminal operation runs, the stream is consumed — calling another operation on it throws IllegalStateException. Start a fresh .stream() each time.

Collectors & grouping

The Collectors class turns a stream into a final result — a list, a set, a joined string, or, most powerfully, a grouped map. groupingBy classifies elements by a key, which is perfect for reports and summaries:

Grouping.java
import java.util.*;
import java.util.stream.*;

List<String> words = List.of("apple", "banana", "avocado", "cherry", "blueberry");

Map<Character, List<String>> byFirstLetter = words.stream()
    .collect(Collectors.groupingBy(w -> w.charAt(0)));

System.out.println(byFirstLetter);

Note

Other handy collectors: Collectors.joining(", ") to build a string, Collectors.counting() and Collectors.summingInt(...) as downstream collectors, and Collectors.partitioningBy(pred) to split into true/false groups.

Numeric & parallel streams

For numbers, specialised streams — IntStream, LongStream, DoubleStream — avoid boxing and add methods like sum(), average(), and range(). And any stream can go parallel with .parallelStream(), splitting the work across CPU cores automatically.

Parallel isn't always faster

Parallel streams help only for large datasets and CPU-heavy work, and your operations must be stateless and side-effect-free. For small collections they're usually slower — measure before reaching for them.

Common pitfalls

  • Reusing a stream — create a new one after each terminal op
  • Side effects in lambdas — don't modify external state inside map/filter; use collect to build results
  • Forgetting laziness — nothing happens without a terminal operation
  • Overusing streams — a simple loop is sometimes clearer; use streams where they add readability

Recap & quick check

Key takeaways

  • A stream pipeline = source + lazy intermediate ops + one terminal op.
  • filter/map/flatMap/sorted/distinct transform; collect/reduce/count/forEach finish.
  • Streams are lazy (nothing runs until the terminal op) and single-use.
  • Collectors.groupingBy classifies elements into a Map — great for summaries.
  • Use numeric streams to avoid boxing; parallelStream only for large, CPU-heavy, side-effect-free work.

Quick check

1. What are the three parts of a stream pipeline?

2. When do intermediate operations actually run?

3. How many times can you use a single stream?

4. Which operation classifies elements into a Map by a key?

5. When are parallel streams worth using?

Outstanding — you can now process data declaratively and elegantly. Next up: Module 19 — File Handling & I/O.