A compact, reproducible set of JVM diagnostic case studies. Each scenario starts a small Java workload that deliberately reproduces a well-known failure mode (a memory leak, lock contention, GC thrash, a deadlock, a container classloader leak), then walks through capturing the evidence, reading it, applying the fix or tuning change, and confirming the result.
The point is the workflow, not the toy code: reproduce, capture, analyze, change one thing, measure again. The same steps apply to a large production heap; here they run on something small enough to inspect end to end.
| # | Scenario | Symptom | Primary evidence | Root cause | Fix / tuning |
|---|---|---|---|---|---|
| 1 | Unbounded cache | steady heap growth, eventual OutOfMemoryError |
heap dump | a map used as a cache with no bound or eviction | bound the cache, add eviction, verify retained size drops |
| 2 | ThreadLocal / classloader leak | old classes never unload after webapp redeploy | heap dump | a ThreadLocal on a pooled thread keeps the old webapp classloader alive |
remove/clear the ThreadLocal, confirm the classloader is collected |
| 3 | Lock contention | throughput collapses under load, low CPU | thread dump | a single synchronized section on the hot path |
narrow the lock scope / use a concurrent structure, re-measure wait time |
| 4 | Allocation pressure | frequent GC pauses, high allocation rate | GC log | short-lived garbage churned on the hot path | reduce allocation and/or tune the collector, compare pause and throughput |
| 5 | Deadlock | requests hang, threads stuck | thread dump | two locks taken in opposite order on two paths | consistent lock ordering, confirm no cycle in the dump |
Scenario 2 runs inside a servlet container (Tomcat or WildFly) to reproduce the redeploy case
faithfully; the rest run as a plain main. Each scenario is selected with a flag so you can
capture one clean signal at a time.
- JDK 11 or newer (scenarios note where JDK 8 flag syntax differs).
- Maven (or Gradle).
- For scenario 2: a local Tomcat or WildFly.
- Analysis tools: Eclipse MAT, VisualVM or JDK Mission Control, GCViewer or an online GC log viewer, and optionally async-profiler.
mvn -q package
# run a single scenario (1..5)
java $JVM_FLAGS -jar target/jvm-heap-forensics.jar --scenario 1Each scenario also takes --fixed to run the corrected version, so before and after are the same
binary and the only difference is the change under test.
Recommended JVM_FLAGS while capturing:
# GC logging (JDK 11+)
-Xlog:gc*,gc+heap=debug,safepoint:file=logs/gc.log:tags,uptime,level
# (JDK 8 equivalent)
-XX:+PrintGCDetails -XX:+PrintGCDateStamps -Xloggc:logs/gc.log
# heap dump automatically on OOM
-XX:+HeapDumpOnOutOfMemoryError -XX:HeapDumpPath=dumps/
# keep heaps small so scenarios are fast to reproduce and inspect
-Xms256m -Xmx256m# find the pid
jcmd -l
# heap dump (live objects only)
jmap -dump:live,format=b,file=dumps/heap.hprof <pid>
# or: jcmd <pid> GC.heap_dump dumps/heap.hprof
# thread dump
jstack <pid> > dumps/threads.txt
# or: jcmd <pid> Thread.print > dumps/threads.txt
# flight recording (60s)
jcmd <pid> JFR.start name=rec settings=profile duration=60s filename=dumps/rec.jfr
# allocation profile with async-profiler (60s)
./profiler.sh -e alloc -d 60 -f dumps/alloc.html <pid>The helpers under scripts/ wrap these so the output lands in dumps/ and logs/ with
consistent names.
- Heap dump (Eclipse MAT): open the
.hprof, run Leak Suspects, then read the dominator tree to see what actually retains memory and follow the shortest path to the GC roots. - Thread dump: group threads by state, look for many
BLOCKEDthreads waiting on one monitor (contention) or aFound one Java-level deadlocksection (scenario 5). - GC log (GCViewer / GCeasy): read pause distribution, throughput, allocation rate, and how full the heap is after each collection, before and after the tuning change.
Placeholders below are filled by running the scenarios and capturing the tools' own output. Numbers and screenshots are real captures, never edited to look better than they are. Nothing in this section is filled in yet.
Scenario 1, heap dump (MAT)
| Metric | Before | After |
|---|---|---|
| Retained size of the cache object | (fill) | (fill) |
| Leak Suspects verdict | (fill) | (fill) |
Time to OutOfMemoryError |
(fill) | no OOM |
(attach: MAT dominator-tree screenshot, leak-suspects screenshot)
Scenario 3, thread dump (jstack)
| Metric | Before | After |
|---|---|---|
Threads BLOCKED on the hot lock |
(fill) | (fill) |
| Throughput (ops/s) | (fill) | (fill) |
(attach: annotated jstack excerpt showing the contended monitor)
Scenario 4, GC log (G1)
| Metric | Before | After tuning |
|---|---|---|
| p99 pause (ms) | (fill) | (fill) |
| Throughput (% time not in GC) | (fill) | (fill) |
| Allocation rate (MB/s) | (fill) | (fill) |
| Flags changed | - | (fill: e.g. region size, IHOP, -Xmn) |
(attach: GCViewer before/after screenshots)
Per-scenario write-ups live in docs/, each with the same before/after table waiting on a real
capture.
jvm-heap-forensics/
src/main/java/... scenario workloads (one class per scenario)
src/test/java/... sanity tests for the workloads and the fixes
webapp/ scenario 2 servlet + deployment descriptor
scripts/ capture helpers (heap, thread, GC, JFR, alloc)
logs/ captured GC logs
dumps/ captured heap and thread dumps, JFR, profiles
docs/ written diagnosis per scenario, with the evidence above
pom.xml
- Each scenario is intentionally minimal so the signal in the dump or log is unambiguous.
- The fixes are the standard, boring ones; the value is in reading the evidence correctly and proving the change with a measurement rather than a guess.
- No external services or agents are required beyond the standard JDK tools and the analyzers above.
logs/anddumps/are gitignored apart from their.gitkeep, so captures stay local.