Performance & profiling in Java
Measure first, JMH micro-benchmarks, boxing and allocation costs.
Measure first
Intuition about hot spots is often wrong. A profiler, like JDK Flight Recorder or async-profiler, shows where time actually goes. Optimize the jammed street, not a random one.
Boxes and caches
What does this print?
Integer a = 100, b = 100;
Integer c = 200, d = 200;
System.out.println(a == b);
System.out.println(c == d);
System.out.println(c.equals(d));true true truetrue false truefalse false true
Show the answer
Autoboxing uses Integer.valueOf, which caches -128 to 127, so a and b are the same object. 200 gets separate objects, so == is false while equals compares values. Long caches the same range.
A boxed accumulator
sum is a boxed Long, so every += unboxes, adds and allocates a new Long: millions of objects in a hot loop. Use a primitive long.
long total(int[] prices) {
Long sum = 0L; // boxed!
for (int p : prices) {
sum += p; // new Long each time
}
return sum;
}Why micro-benchmarks lie
Hand-written timing loops fall into JVM traps: warm-up (early iterations run interpreted), dead-code elimination (an unused result lets the JIT delete the work) and profile pollution. JMH handles warm-up, forking and dead-code elimination for you.
Timing a calculation
long t = System.nanoTime();
for (int i = 0; i < N; i++)
compute(i); // result unused
long ns = System.nanoTime() - t;The JIT may remove compute() entirely: ~0 ns.
@Benchmark
public long compute() {
return work(input);
}Returned results (or a Blackhole) keep the work alive.
Why warm up?
Why does JMH run warm-up iterations before it starts measuring?
Think about it, then reveal the answer
So the JIT has compiled and optimized the code first. You want steady-state performance; measuring early iterations would mix interpreter and compilation time into your results.
A typical profiling win
A team suspects JSON parsing and plans a rewrite. A profile shows the real cost elsewhere: say, a regex compiled on every request, or a debug message built for a disabled log level. A one-line fix beats a week of guessing.
Key takeaways
- Profile before optimizing
- Use JMH for micro-benchmarks, not System.nanoTime loops
- Boxing a long into Long allocates objects
- Long/Integer caches only cover -128 to 127
💡 Optimizing without profiling is like fixing traffic by widening a random street instead of the jammed one.
JMH is an OpenJDK project, written by engineers who work on the JVM itself, the people who know best how the JIT can fool a naive benchmark.
Practice questions
What does this print?
Long a = 127L, b = 127L;
Long c = 1000L, d = 1000L;
System.out.println(a == b);
System.out.println(c == d);
System.out.println(c.equals(d));- true true true
- true false true
- false false true
- true false false
Check your answer
true false true. Long.valueOf (used by autoboxing) caches -128 to 127, so a and b are the same object. 1000L creates separate objects, so == is false while equals compares values.
A hand-written benchmark times a loop with System.nanoTime() and reports ~0 ns for a calculation whose result is never used. What happened?
- The GC ran during the loop
- The CPU cached the answer
- The JIT removed the unused computation (dead-code elimination)
- System.nanoTime() is broken on servers
Check your answer
The JIT removed the unused computation (dead-code elimination). If a result is never used, the JIT may delete the work entirely. JMH avoids this by returning results or feeding them to a Blackhole.