Skip to content
Rigel Carbajal
thought5 min read

JVM lessons at enterprise scale: when G1GC falls short and ZGC comes to the rescue

Some time ago I worked on a case that taught me more about garbage collection than any course ever did. The short version: a Java cluster at the edge of collapse didn’t need magic resources. It needed someone to understand the physics of the garbage collector and the architecture of the JVM under extreme load.

At enterprise scale, the JVM’s default behavior stops being enough. And when your application is battle-tested on G1GC, the industry standard, switching to ZGC is a bet: sub-millisecond pauses sound great, but they cost CPU. The real question is when it’s worth deviating from the standard, and how to diagnose that without falling into configuration placebo.

Quick note on confidentiality: I can’t share the client’s name, the product’s name, or exact numbers. What matters here is the process.

The scenario

A distributed environment with three nodes, serving over 290,000 active users and managing close to 30 million content objects.

At this scale, problems don’t announce themselves politely. The service showed intermittent outages, HTTP/Tomcat thread saturation, and constant node evictions in the clustering layer. Operations’ temporary fix was the classic node restart: a palliative patch that only masked a structural failure in the JVM.

The chain reaction: symptoms vs. root cause

When a node in a Java cluster “disappears,” the immediate instinct is to blame the network. But digging into the diagnostic logs, the real failure flow looked like this:

[Massive request load / saturated heap]
                │
                ▼
[G1GC stop-the-world pauses (> 5,000 ms)]
                │
                ▼
[Cluster heartbeat ping is missed]
                │
                ▼
[The cluster leader assumes the node is dead and evicts it]

The clustering layer’s safety mechanism was evicting nodes because during long garbage collection pauses, the entire Java process froze, unable to answer control pings. The node wasn’t down. It was frozen, and the cluster couldn’t tell the difference. That’s why restarting “worked”: on the way back up, the node simply reclaimed its place.

The three factors suffocating the heap

  1. Sizing and the 32 GB trap. The heap was set to just 16 GB. And when scaling memory, you have to watch out for the compressed object pointers (Compressed OOPs) gap: the inefficient range between 32 GB and 47 GB. We recommended 31 GB as the initial sweet spot. I wrote a longer analysis of that boundary in The 32 GB Paradox.
  2. API abuse. Integration clients running repetitive requests asking for expand=body.storage, forcing the application to parse and render heavy content in memory continuously.
  3. Document parsing bloat. When we dumped and analyzed the heap, we found tens of millions of XSSFCell and ElementXObj objects. The processing and decompression of gigantic Excel spreadsheets was consuming a critical portion of old-gen memory.

The transition: G1GC vs. ZGC (Java 21)

The application engine is optimized for G1GC by default, and that’s worth respecting. But the nature of this workload demanded reducing STW pauses at almost any cost. So we tested ZGC on a single node while keeping the others on G1GC, to compare real behavior side by side.

The metric contrast

  • G1GC (16 GB, unmigrated nodes): max stop-the-world pauses bordering ~5,000 ms.
  • ZGC (31 GB, Java 21): max stop-the-world pauses of ~1-2.9 ms, averaging around 1 ms.

The hidden lesson of generational ZGC

Migrating to ZGC in Java 21 is not just -XX:+UseZGC. If you don’t enable the ZGenerational flag, the collector treats the whole heap as one homogeneous generation, and you keep hitting allocation stalls: application threads waiting on the collector. We saw roughly 8,000 stall events that were largely avoidable. Enabling the generational mode split young and old generations, cutting the collector’s work dramatically and making ZGC behave the way it’s supposed to.

The monitoring and load balancing angle

Once the JVM stabilized, two final surprises showed up:

  1. Load balancer imbalance. One node kept registering memory peaks of 97% while the others sat around 54%. The network layer was directing a disproportionate share of heavy requests to a single instance, masking the health of the rest.
  2. APM reconfiguration (Dynatrace). APM tools configured to monitor G1GC misread ZGC. G1GC performs long, spaced-out pauses; ZGC collects garbage continuously in the background with imperceptible pauses. If the APM isn’t reconfigured to understand that behavior, it generates false alarms and distorted metrics. G1GC and ZGC are measured differently for a reason.

With the telemetry recalibrated, the final numbers told the real story (the middle values belong to the node that never left G1GC):

  • Heap in use: 43% / 71% / 54%
  • Heap peaks: 95% / 97% / 77%
  • Max STW pauses: 2.9 ms / 22.9 ms / 4.2 ms
  • Avg STW pauses: ~1 ms on all nodes
  • MMU@100ms: 93.4% / 49% / 90.7%
  • Out-of-memory errors: 0 / 0 / 0

The weak link was exactly the node the load balancer flagged as “under pressure.”

Conclusions & takeaways for software architects

  1. Don’t tune the GC without looking at the APIs. Raising memory or switching collectors only buys you time. Without rate limiting on heavy API calls and restrictions on massive document parsing, the heap will eventually run out again.
  2. Respect the clustering layer. Most “node down due to network timeout” events in a cluster are actually undiagnosed GC pauses. The heartbeat protocol will always read a frozen node as a dead one.
  3. ZGC is a game-changer for latency SLAs, but it’s not free. Average pauses of ~1 ms on an instance with millions of objects transforms the user experience. It demands monitoring the extra CPU consumption and calibrating the generational flags correctly on JDK 21+.
  4. Heaps have forbidden ranges. Below 32 GB, pointers compress. Between 32 and 47 GB, performance collapses. Pick 31 GB, or jump to 48 GB.
  5. Monitoring is part of the migration. A different collector has a different rhythm. If your APM still expects G1GC patterns, you’ll be fighting false alarms instead of real data.
  6. Migrate one node at a time, but finish the job. Comparing ZGC against G1GC in production is a great experiment. Leaving a weak link behind is not.

Again: client name, product name, and exact figures omitted out of respect for confidentiality. The investigation is the part worth stealing.