What you'll learn
The Stream API lets you process collections in a declarative, readable pipeline — say what you want (filter, transform, group) rather than writing loops that spell out how. Combined with lambdas, it's one of the most powerful features in modern Java.
By the end you'll be able to:
- Build stream pipelines: source → intermediate ops → terminal op
- Transform data with
filter,map,flatMap,sorted - Produce results with
collect,reduce,count, and more - Group and summarise data with
Collectors - Use numeric and parallel streams, and avoid common pitfalls
The stream pipeline
A stream is not a data structure — it's a flow of elements you run operations on. A pipeline has three parts: a source (usually collection.stream()), any number of intermediate operations that transform the flow (and return another stream), and exactly one terminal operation that produces a result:
import java.util.*;
import java.util.stream.*;
List<String> names = List.of("Sara", "Omar", "Alice", "Bob", "Sam");
List<String> result = names.stream()
.filter(n -> n.length() > 3) // keep names longer than 3
.map(String::toUpperCase) // transform each
.sorted() // order them
.collect(Collectors.toList()); // gather into a List
System.out.println(result);Key idea
Intermediate operations
Intermediate operations are lazy — they don't run until a terminal operation is attached. Each returns a new stream, so they chain:
| Operation | What it does |
|---|---|
filter(pred) | keep only elements matching the predicate |
map(fn) | transform each element |
flatMap(fn) | flatten nested streams into one |
sorted() | sort (natural order or a Comparator) |
distinct() | remove duplicates |
limit(n) / skip(n) | take the first n / skip the first n |
flatMap deserves a special mention — it's how you turn a stream of collections into a single flat stream:
import java.util.*;
import java.util.stream.*;
List<List<Integer>> nested = List.of(
List.of(1, 2), List.of(3, 4), List.of(5)
);
List<Integer> flat = nested.stream()
.flatMap(List::stream) // flatten streams-of-streams into one
.collect(Collectors.toList());
System.out.println(flat);Terminal operations
A terminal operation ends the pipeline and produces a value (or a side effect). Common ones are collect (gather into a collection), forEach (do something with each), count, reduce (combine into one value), and min/max:
import java.util.*;
List<Integer> nums = List.of(1, 2, 3, 4, 5, 6);
long evens = nums.stream().filter(n -> n % 2 == 0).count();
int sum = nums.stream().mapToInt(Integer::intValue).sum();
Optional<Integer> max = nums.stream().max(Integer::compareTo);
System.out.println("Evens: " + evens);
System.out.println("Sum: " + sum);
System.out.println("Max: " + max.get());Watch out
IllegalStateException. Start a fresh .stream() each time.Collectors & grouping
The Collectors class turns a stream into a final result — a list, a set, a joined string, or, most powerfully, a grouped map. groupingBy classifies elements by a key, which is perfect for reports and summaries:
import java.util.*;
import java.util.stream.*;
List<String> words = List.of("apple", "banana", "avocado", "cherry", "blueberry");
Map<Character, List<String>> byFirstLetter = words.stream()
.collect(Collectors.groupingBy(w -> w.charAt(0)));
System.out.println(byFirstLetter);Note
Collectors.joining(", ") to build a string, Collectors.counting() and Collectors.summingInt(...) as downstream collectors, and Collectors.partitioningBy(pred) to split into true/false groups.Numeric & parallel streams
For numbers, specialised streams — IntStream, LongStream, DoubleStream — avoid boxing and add methods like sum(), average(), and range(). And any stream can go parallel with .parallelStream(), splitting the work across CPU cores automatically.
Parallel isn't always faster
Common pitfalls
- Reusing a stream — create a new one after each terminal op
- Side effects in lambdas — don't modify external state inside
map/filter; usecollectto build results - Forgetting laziness — nothing happens without a terminal operation
- Overusing streams — a simple loop is sometimes clearer; use streams where they add readability
Recap & quick check
Key takeaways
- A stream pipeline = source + lazy intermediate ops + one terminal op.
- filter/map/flatMap/sorted/distinct transform; collect/reduce/count/forEach finish.
- Streams are lazy (nothing runs until the terminal op) and single-use.
- Collectors.groupingBy classifies elements into a Map — great for summaries.
- Use numeric streams to avoid boxing; parallelStream only for large, CPU-heavy, side-effect-free work.
Quick check
1. What are the three parts of a stream pipeline?
2. When do intermediate operations actually run?
3. How many times can you use a single stream?
4. Which operation classifies elements into a Map by a key?
5. When are parallel streams worth using?
Outstanding — you can now process data declaratively and elegantly. Next up: Module 19 — File Handling & I/O.