🌊 Streams API · Intermediate

Parallel streams in Java

ForkJoinPool, when parallel helps and when it hurts; ordering.

🧩 The mysteryAdd one word — .parallel() — and your stream uses every CPU core. So why can it make a web server *slower*?

Many hands

parallel() (or list.parallelStream()) splits the data into chunks processed by threads of **ForkJoinPool.commonPool(). That pool is shared by the whole JVM and has roughly one thread per CPU core**.

When it helps, when it hurts

Helps: big data, CPU-heavy work, sources that split cheaply (**arrays, ArrayList, IntStream.range), no shared state. Hurts: tiny inputs (splitting costs more than it saves), blocking I/O, sources that split poorly (LinkedList**, Stream.iterate).

🔮 Predict it

Does parallel scramble results?

What does this print?

List<Integer> r = IntStream.range(0, 5)
    .parallel()
    .map(n -> n * 10)
    .boxed()
    .toList();
System.out.println(r);
  1. [0, 10, 20, 30, 40]
  2. [40, 30, 20, 10, 0]
  3. It varies from run to run
Show the answer

[0, 10, 20, 30, 40] — work runs on several threads, but **toList reassembles results in encounter order**.

Who keeps the order?

**toList/collect: results in encounter order. forEach: no ordering guarantee in parallel. forEachOrdered: encounter order, at some cost to speed. findAny**: whichever match a thread finds first.

🤔 Think first

Small data, slower code

Why might List.of(20 names).parallelStream().filter(...) be *slower* than the sequential version?

Think about it, then reveal the answer

Splitting the work, scheduling tasks on other threads and merging results has a fixed overhead. With 20 tiny elements that overhead dwarfs the work itself.

⚠️ The trap

Blocking calls starve the shared pool

A parallel stream that makes blocking database or HTTP calls parks the few common-pool threads. Since the pool is shared JVM-wide, every other parallel stream waits too and throughput collapses. For blocking I/O use a dedicated executor or virtual threads.

💼 In the real world

Measure first

Experienced teams don't sprinkle .parallel() on hunches: they benchmark (e.g. with JMH). It shines in batch jobs crunching huge arrays — and causes incidents on servers where hundreds of requests compete for the same small common pool.

Key takeaways

  1. Uses the shared ForkJoinPool.commonPool()
  2. Helps: big data, CPU-bound, cheap to split, no shared state
  3. Hurts: small data, blocking calls, LinkedList, iterate()
  4. toList keeps encounter order; forEach does not
🤯 Did you know?

The common pool's default parallelism is the number of available processors minus one — the thread that calls the terminal operation joins in as an extra worker.

Practice questions

Which workload is MOST likely to get faster with .parallel()?

  1. Filtering a list of 20 names
  2. Summing 50 million ints from an array
  3. Calling a slow remote API for each element
  4. Processing a 100-element LinkedList
Check your answer

Summing 50 million ints from an array. Large, CPU-bound work on an array splits cheaply and evenly across cores. Tiny inputs cost more to split than they gain, linked lists split poorly, and blocking calls waste the shared pool's threads.

What does this print?

List<Integer> r = IntStream.rangeClosed(1, 5)
    .parallel()
    .map(n -> n * n)
    .boxed()
    .toList();
System.out.println(r);
  1. [1, 4, 9, 16, 25]
  2. [25, 16, 9, 4, 1]
  3. It varies from run to run
  4. [1, 2, 3, 4, 5]
Check your answer

[1, 4, 9, 16, 25]. Even in parallel, toList assembles results in encounter order. Only operations like forEach and findAny give up ordering.

Next: toList() versus collect(Collectors.toList()) — they look identical, until you call add.