Page tables can eat tens of GiB of RAM: kernel and database cases

The post starts from a 2003 Linus Torvalds argument against hashed page tables. In a tree, entries of neighbouring pages sit next to each other, so one cache line fills several TLB entries at once; a hash table scatters neighbours across buckets. In the quoted passage Torvalds says a bog-standard Intel CPU fetches 8 TLB entries in one go. The post says his 1997 master's thesis explains why the kernel adopted multi-level page tables, and that he cared more about latency than memory size.

The author then turns to cases where memory size matters. In 2020 Shakeel Butt sent a patch adding a page table metric to memory.stat. Roman Gushchin asked for the use case, arguing page tables are about 1/512 of mapped memory, under 1% of most cgroups ("for all but very large cgroups the value will be in the noise of per-cpu counters"). Butt's answer was a user space network driver that maps application memory for zero copy and uses a lot of page-table memory.

The 1/512 ratio holds when each page is mapped once. If N processes map the same page, each needs its own PTEs, so page tables cost about N/512 of the memory they map. With 4 KiB pages and 8-byte PTEs, 512 processes mapping one page need 512 PTEs, which is 4 KiB, as much as the data page itself. (The author adds that N should strictly be address spaces, not processes.)

The post lists similar cases. In 2002 Andrea Arcangeli found why x86 machines with 64 GiB of RAM ran out of memory despite plenty of free memory: hundreds of processes each mapped the same 1 GiB of shared memory, needing about 2 MiB of page tables each with PAE, and those tables lived in the roughly 896 MiB "lowmem" area. His patch moved page tables from lowmem to highmem. The author calls it an edge case that arose only on 32-bit machines with PAE. In 2022 Khalid Aziz reported an Oracle database server with 512 GB of RAM crashing with OOM when 1500+ clients attached to a 300 GB SGA of shared memory. In the worst case, where every process maps the whole SGA, PTEs alone would take 878 GB. He proposed mshare so processes could share page tables; as far as the author checked, no version has been merged.

Qi Zheng's 2021 case had no sharing. A single process had 590 GiB RSS and 110 GiB of page tables, when around 1.2 GiB of PTEs should have sufficed. The workload used jemalloc and tcmalloc, which return memory with madvise(MADV_DONTNEED). That frees data pages and clears PTEs but keeps the page tables allocated, so empty page tables piled up. The patch went through several rewrites and was merged in 2025. A footnote relays a Hacker News comment from Sopel about Windows: zombie processes can hold at least 32 KiB of page tables, and after 100,000 exited cmd.exe processes, triggered by an AMD iGPU driver bug, RAMMap showed about 3.5 GiB of page tables.

The author argues the pattern is common and cites two database articles. In 2021 Jobin Augustine at Percona reproduced Postgres OOM kills on a 192 GB machine with shared buffers at 138 GiB and only 80 connections: page tables grew from 45 MiB to over 25 GiB as backends touched more of the cache, and the machine started swapping. Huge pages fixed it, with page tables staying at 61 MiB for the same workload. In 2026 Kaushik Iska at ClickHouse measured it directly on a 128 GiB machine with 32 GiB of shared buffers. Each backend scanning a 15.6 GiB cached table added 31 MiB of page tables. At 200 connections page tables took 6.1 GiB, versus 111 MiB with huge pages.

Huge pages have costs. Reserving them needs unfragmented memory, and ClickHouse notes that on a loaded machine "the same request routinely fails". THP can stall on allocation. In 2016 Mel Gorman proposed that the kernel stop defragmenting for THP by default, saying it was "time to throw in the towel". The post also mentions Nelson Elhage's writing on possible memory leaks and CPU usage problems from THP.

The last section covers NUMA machines. A thread on another node has to walk page tables in remote memory on TLB misses. Mitosis (ASPLOS 2020) showed remote page tables can slow an application as much as remote data, and the Hydra paper (USENIX ATC 2024) reproduced this on an 8-socket machine with 8 TB of RAM. Mitosis replicates the whole page table tree on every node, which costs page table size times the number of nodes and means every change updates every copy. Hydra copies a PTE to a node only when one of that node's threads faults on it, and each page table page keeps a list of nodes holding a copy. The conclusion: you either keep one copy and pay for remote reads on TLB misses, or keep a copy per node and pay to keep them in sync.

Key facts

  • Page tables are about 1/512 of mapped memory when each page is mapped once, but cost about N/512 when N processes map the same page; at 512 processes they equal the data itself.
  • Real cases: 64 GiB x86 machines in 2002, an Oracle server with 512 GB RAM and 1500+ clients on a 300 GB SGA in 2022, and a 590 GiB RSS process with 110 GiB of page tables in 2021.
  • Postgres at Percona (2021) saw page tables grow from 45 MiB to over 25 GiB with 80 connections; huge pages held them at 61 MiB. ClickHouse (2026) measured 6.1 GiB at 200 connections versus 111 MiB with huge pages.
  • Huge pages are not free: they need unfragmented memory, and THP can stall on allocation.
  • On NUMA machines the choice is one page-table copy with remote reads on TLB misses, or a copy per node that must be kept in sync (Mitosis replicates everything, Hydra copies a PTE only when a node faults on it).

Why it matters

Page tables are usually treated as noise: about 1/512 of mapped memory, as Roman Gushchin argued in 2020. The post shows where that estimate breaks. When many processes map the same shared memory, each carries its own copy of the PTEs, and the overhead scales with the process count. The same pattern appears in a 2002 kernel case, a 2022 Oracle report and 2021 and 2026 Postgres write-ups, so it is a recurring failure and not a one-off. The NUMA section adds a second cost, latency, and shows there is no free option.

Who it affects

Operators of databases where many backend processes attach to one large shared memory segment, such as Postgres shared buffers or an Oracle SGA. Also kernel developers working on memory accounting, page-table sharing (mshare) and NUMA replication, and anyone running a single process with allocators like jemalloc or tcmalloc that return memory with madvise(MADV_DONTNEED). The post is aimed at systems readers; it does not concern AI directly.

How to use it

The post gives no step-by-step guide, but the cases point to a few measures. For Postgres-style shared memory, huge pages cut page tables from over 25 GiB to 61 MiB at Percona and from 6.1 GiB to 111 MiB at ClickHouse. The memory.stat page table metric that Shakeel Butt proposed in 2020 is the kind of counter that makes the cost visible, though the source does not say whether that patch was merged. For the MADV_DONTNEED case, the fix described is the kernel patch merged in 2025.

How solid is it

It is a personal blog post and the author is not named in the source text. Its figures come from kernel mailing-list reports and from Percona and ClickHouse articles, and the post attributes each one to the person who reported it. Some statements are the author's own reading, such as the explanation of the Qi Zheng case and the comparison of Mitosis and Hydra. On mshare the author says only that, as far as they checked, no version has been merged. The source gives no figures for the slowdown from remote page tables in Mitosis or Hydra.

Risks and caveats

Huge pages are a remedy with costs: reserving them needs unfragmented memory, and ClickHouse notes the request routinely fails on a loaded machine. THP can stall on allocation, Mel Gorman proposed in 2016 that the kernel stop defragmenting for it by default, and Nelson Elhage wrote about possible memory leaks and CPU usage issues. The 2002 case was, in the author's words, an edge case tied to 32-bit x86 with PAE. The worst-case 878 GB figure assumes every process maps the whole SGA. The source does not say how Khalid Aziz's OOM was ultimately resolved.

“for all but very large cgroups the value will be in the noise of per-cpu counters”

— Roman Gushchin, as quoted in the post, arguing against a page table metric in memory.stat