A Practical Look at G1's Internals and Real-World Behavior

G1 (Garbage-First) has been Oracle's default collector since Java 9. Before that it was opt-in for most workloads starting around Java 7. It replaced CMS because CMS was showing its age, and G1 replaced Parallel GC as the default because most modern servers have large heaps and multiple cores. The collector was designed by Andrew Vershinin and team at Sun Microsystems, which is a detail most people don't bother learning but it helps if you ever need to file a bug report. The core idea is simple: divide the heap into roughly equal-sized regions instead of treating it as one contiguous block. Each region tracks how much live data it holds during sampling cycles. When the heap needs to be collected, G1 prioritizes regions that will give the most garbage collection for the least effort. Hence the name Garbage-First.

What Happens In G1

When you launch a process with -XX:+UseG1GC, the JVM initializes the heap as a set of regions. By default, region size is calculated as heap_size / 2048, capped between 1MB and 32MB. On a 32GB heap that comes out to about 16MB per region. You can override this with -XX:G1HeapRegionSize but most people don't need to. The collector runs several concurrent phases. The initial mark phase pauses the world briefly to mark roots from GC roots. Then concurrent marking traverses the object graph across all threads while they keep working. After that, the concurrent refinement phase updates remembered sets on each region — these are essentially dirty card tables that track cross-region references. Finally, the evacuation phase copies live objects from retiring regions into younger ones during a stop-the-world pause. The pause time target is controlled by -XX:MaxGCPauseMillis, which defaults to 200 milliseconds. G1 will try to shrink or grow its working set to hit that target. If you set it too aggressively low, like 50ms on a heap that actually needs more time to clean up, G1 will just do more frequent collections and your throughput will tank. I learned this the hard way on a payment processing service where the SLA was "no pause over 100ms" and the engineering team had set the target to 80ms. The result was roughly two collections per second with almost no actual cleanup happening. We moved the target to 200ms and the pause distribution improved dramatically because G1 was actually able to reclaim meaningful amounts of space per cycle.

During mixed collections, G1 selects both young and old regions. Young regions get evacuated first because they typically contain more dead data. Old regions are then evaluated against their remembered set sizes and copy costs. The decision to include an old region depends on whether the estimated collection cost fits within the pause budget. This is where the "garbage-first" naming makes sense — G1 is essentially solving a knapsack optimization problem every time it runs a mixed collection. Humongous allocations are one of the trickier corner cases. If an object is larger than half the region size, G1 allocates it in a chain of consecutive regions called a humongous region. These bypass the normal Eden space allocation path. The problem is that freeing humongous objects requires traversing the entire chain during evacuation, which adds latency. I worked on a system handling large protobuf payloads where a single request occasionally triggered a 4MB serialization. That object would land as humongous and the subsequent collection pauses would spike to 400ms even though MaxGCPauseMillis was set to 200ms. The workaround was tuning -XX:G1HeapWastePercent down to 5 and enabling -XX:+AlwaysPreTouch to reduce deferred page initialization, but the real fix was capping the payload size at 2MB and splitting larger messages. G1 handles humongous allocations fine in normal workloads. It's pathological cases like this where the assumptions break down. Another thing beginners consistently get wrong is the relationship between G1 and concurrent mode failures. When G1 can't free enough memory before a young generation collection needs to promote objects, it falls back to a full heap collection. This is a concurrent mode failure and it looks exactly like a CMS concurrent mode failure but happens for different reasons. The most common cause is insufficient young gen sizing combined with high allocation rates. G1 dynamically adjusts region counts but if your application has sustained peak allocation that exceeds what the current region pool can handle, you'll get these failures. Setting -XX:G1ReservePercent to 15 or 20 from the default 10 gives G1 a bit more breathing room during evacuation and I've seen this eliminate mode failures on several services without any other changes.

Get the Full Details

What are the differences between G1 and G2 phase? 2026
What are the differences between G1 and G2 phase? 2026

String deduplication is a feature most people don't know exists until they need it. Enable it with -XX:+UseStringDeduplication and G1 will track identical string instances and replace duplicates with references to a single copy. This works best in applications withI had one service with 16GB of heap where 40 percent of the string data was duplicated. Turning on deduplication reduced heap usage by roughly 2.3GB and cut GC frequency by about a third over a week-long load test. There are scenarios where G1 is genuinely a bad choice. It performs poorly on very small heaps under 2GB — the overhead of region management and concurrent marking phases isn't worth it. Parallel GC or the newer ZGC will outperform it there. G1 also struggles with extremely low-latency requirements below 10ms pauses because it's not designed for that regime. ZGC and Shenandoah exist specifically for those workloads. And if your application has very long-lived objects with few short-lived ones, G1's concurrent marking becomes expensive relative to the benefit because it spends most of its time scanning live object graphs that could just be handled by a simple mark-sweep. The monitoring story for G1 has improved significantly since Java 9. jcmd with the GC.heap_info and GC.get_histogram commands gives you region-level detail. jstat -gcutil still works but only shows high-level percentages. The real diagnostic power comes from -XX:+PrintGCDetails combined with -Xlog:gc* in Java 9 and later. The structured logging format lets you parse pause times, region selection decisions, and evacuation stats directly from logs without needing additional tools.

If you want to experiment locally, Oracle's official JDK builds include G1 enabled by default for Java 9 through 21. For Java 8 you need to explicitly enable it with the flag I mentioned earlier. Most containerized deployments should also consider setting -XX:+UseContainerSupport so G1 sizes itself based on container memory limits rather than host memory, which prevents the collector from overcommitting and getting killed by the OOM manager. The region-based design means G1 scales reasonably well with both heap size and CPU count, but it's not a silver bullet. Understanding the actual mechanics — how remembered sets work, when humongous regions get allocated, what drives evacuation selection — makes the difference between a collector that keeps your latency smooth and one that surprises you during a traffic spike. The documentation covers the happy path well. The edge cases are where experience matters.