Go garbage collector pause hits 40ms when its metadata is swapped out

The author of a blog post was considering running swap in production to absorb memory spikes, and wrote up what he found so others would not make the same mistake. His setup was a cgroup with two processes. One is a Go process that calls io.ReadAll and then proto.Unmarshal, creating a blob and then a graph struct that Go's allocator marks as scan. The other is a mostly quiet HTTP server.\n\nHis reasoning went like this. Whenever the collector runs, it reads the scan spans pointer by pointer. Under memory pressure the kernel would evict pages to the swap device, but eviction is per cgroup, not per process, so pages of both processes would be evicted. That, he thought, made a damaging swap-in and swap-out dance between the kernel and the garbage collector unlikely. He was wrong. Go's garbage collector reads its metadata, which lives outside the heap in a region that is not freed, during a stop-the-world pause, and that metadata can be in swap.\n\nHe tested this with a mock run on a Hetzner box, using kernel 6.8 with MGLRU enabled and a mock allocator. The plots, allocator, bpf scripts and Python scripts are published in a GitHub repository (github.com/frnsimoes/go-gc-swap-cost). The median pause was around 51 us. With the metadata on NVMe swap, the worst pause was 40ms. To see where the time went, he wrote a small bpf script that counts page faults while the world is stopped. The worst pause lasted 39902 us, with 228 page faults during it, and 39013 us of that spent in the faults. An addr2line dump placed those faults inside the GC's own bookkeeping: runtime.finishsweep_m, runtime.nextMarkBitArenaEpoch and runtime.(*spanSet).reset, all reached from runtime.gcStart.\n\nGo's GC has to stop the world at two points: sweep termination and mark termination. The test saw 312 such pauses in 30 minutes. His explanation of the mechanism: the runtime allocates these metadata pages and reuses them rather than freeing them, and they are read in GC cycles. The kernel evicts pages by age, so the least recently accessed pages go to swap. When the GC stops the world and reads them, a major page fault follows. The kernel has to read the PTEs, call do_swap_page, find a new frame, charge it to the cgroup, read the pages, submit a bio, wait for the disk and put the pages back in memory.\n\nForty milliseconds sounds harmless, but this is a stop-the-world pause: every P has stopped. As he puts it, a goroutine waiting for I/O might see that I/O return during the pause with no one to handle it. He notes that 40ms is 800 times the median pause, and that it happened two or three times per memory spike during the test.\n\nHe also noticed a second cost. Building one 511 KiB message, which usually takes 3-5 ms, jumped to 105 ms on the NVMe swap and to 903 ms on Hetzner's network volume. Per message that costs more than the metadata pause, but only the goroutine doing the allocation pays, so it is localized rather than global. He has not confirmed where that time goes.\n\nHis conclusion is mixed. He agrees with Chris Down that swap is not evil, but says it did not behave well with garbage collection, and his production workload collects a lot. An update dated September 14 (no year given) says someone asked whether Go 1.26's Green Tea garbage collector changes the way the GC reads metadata. He measured it and says the impact is negligible.
Key facts
- In a mock run on a Hetzner box (kernel 6.8, MGLRU enabled), the worst Go GC stop-the-world pause was 40ms with the GC metadata on NVMe swap, against a median of about 51 us.
- A bpf script showed the worst pause (39902 us) included 228 page faults that took 39013 us, all inside GC bookkeeping such as runtime.finishsweep_m and runtime.nextMarkBitArenaEpoch.
- The test saw 312 stop-the-world pauses in 30 minutes, and the 40ms pause happened two or three times per memory spike.
- Building one 511 KiB message went from the usual 3-5 ms to 105 ms on NVMe swap and 903 ms on Hetzner's network volume; the author has not confirmed why.
- A September 14 update says Go 1.26's Green Tea garbage collector made a negligible difference to how the GC reads metadata.
Why it matters
Swap is a tempting way to survive memory spikes, and the common intuition is that per-cgroup eviction spreads the cost around. This experiment suggests a specific failure mode for Go services: the runtime's own metadata can be paged out, and the GC reads it while the whole program is stopped. A pause that normally lasts about 51 us then becomes tens of milliseconds, roughly 800 times the median by the author's count. The author frames it as a problem that could have hurt him in production, not just a benchmark curiosity.
Who it affects
Teams running Go services who are thinking about enabling swap, or who already have swap on hosts with memory pressure, particularly under cgroup limits. The author's own workload is one where, in his words, he is collecting a lot, which makes frequent GC cycles and so frequent stop-the-world pauses part of the picture. The slower message building affects the goroutine that allocates, so it matters most for code that builds large buffers under pressure.
How to use it
There is no tool to install. The author published the experiment materials in a GitHub repository (github.com/frnsimoes/go-gc-swap-cost): plots, the mock allocator, bpf scripts and Python scripts. The bpf script counts page faults while the world is stopped, and the addr2line output maps the faults to Go runtime functions. Anyone wanting to check their own setup can start from those scripts. The article proposes no fix or mitigation for the metadata-in-swap problem.
How solid is it
It is a hands-on measurement with instrumentation behind it: a bpf page-fault count and symbolized stack frames tie the 39 of 40 ms directly to 228 page faults in GC bookkeeping. The limits are clear, though. The numbers come from a mock run with a mock allocator on one Hetzner box (kernel 6.8, MGLRU enabled), not from a production incident. No Go version is stated for the main experiment; Go 1.26 appears only in the update. The Green Tea measurement is reported only as negligible, with no figures. The article does not name its author in the text.
Risks and caveats
The author does not confirm where the extra time in building 511 KiB messages goes (105 ms on NVMe, 903 ms on a network volume), so that part is an observation, not an explained result. His point that I/O might return during a pause with no one to handle it is stated as a possibility, not a measured effect. He also does not say that swap should be avoided: he agrees with Chris Down that swap is not evil, and the article does not say whether he ultimately decided against running swap in production. Results may differ with other kernels, storage and allocation patterns.
“40ms is 800 times the median pause. It happens two or three times per memory spike during the test. It is a lot.”
— The article's author