Parallel streams in Java
ForkJoinPool, when parallel helps and when it hurts; ordering.
.parallel() — and your stream uses every CPU core. So why can it make a web server *slower*?Many hands
parallel() (or list.parallelStream()) splits the data into chunks processed by threads of **ForkJoinPool.commonPool(). That pool is shared by the whole JVM and has roughly one thread per CPU core**.
When it helps, when it hurts
Helps: big data, CPU-heavy work, sources that split cheaply (**arrays, ArrayList, IntStream.range), no shared state. Hurts: tiny inputs (splitting costs more than it saves), blocking I/O, sources that split poorly (LinkedList**, Stream.iterate).
Does parallel scramble results?
What does this print?
List<Integer> r = IntStream.range(0, 5)
.parallel()
.map(n -> n * 10)
.boxed()
.toList();
System.out.println(r);[0, 10, 20, 30, 40][40, 30, 20, 10, 0]It varies from run to run
Show the answer
[0, 10, 20, 30, 40] — work runs on several threads, but **toList reassembles results in encounter order**.
Who keeps the order?
**toList/collect: results in encounter order. forEach: no ordering guarantee in parallel. forEachOrdered: encounter order, at some cost to speed. findAny**: whichever match a thread finds first.
Small data, slower code
Why might List.of(20 names).parallelStream().filter(...) be *slower* than the sequential version?
Think about it, then reveal the answer
Splitting the work, scheduling tasks on other threads and merging results has a fixed overhead. With 20 tiny elements that overhead dwarfs the work itself.
Blocking calls starve the shared pool
A parallel stream that makes blocking database or HTTP calls parks the few common-pool threads. Since the pool is shared JVM-wide, every other parallel stream waits too and throughput collapses. For blocking I/O use a dedicated executor or virtual threads.
Measure first
Experienced teams don't sprinkle .parallel() on hunches: they benchmark (e.g. with JMH). It shines in batch jobs crunching huge arrays — and causes incidents on servers where hundreds of requests compete for the same small common pool.
Key takeaways
- Uses the shared ForkJoinPool.commonPool()
- Helps: big data, CPU-bound, cheap to split, no shared state
- Hurts: small data, blocking calls, LinkedList, iterate()
- toList keeps encounter order; forEach does not
The common pool's default parallelism is the number of available processors minus one — the thread that calls the terminal operation joins in as an extra worker.
Practice questions
Which workload is MOST likely to get faster with .parallel()?
- Filtering a list of 20 names
- Summing 50 million ints from an array
- Calling a slow remote API for each element
- Processing a 100-element LinkedList
Check your answer
Summing 50 million ints from an array. Large, CPU-bound work on an array splits cheaply and evenly across cores. Tiny inputs cost more to split than they gain, linked lists split poorly, and blocking calls waste the shared pool's threads.
What does this print?
List<Integer> r = IntStream.rangeClosed(1, 5)
.parallel()
.map(n -> n * n)
.boxed()
.toList();
System.out.println(r);- [1, 4, 9, 16, 25]
- [25, 16, 9, 4, 1]
- It varies from run to run
- [1, 2, 3, 4, 5]
Check your answer
[1, 4, 9, 16, 25]. Even in parallel, toList assembles results in encounter order. Only operations like forEach and findAny give up ordering.