Summary
A 15-patch RFC posted to the Linux memory-management list proposes a substantial change to Multi-Gen LRU (MGLRU): promote folios according to observed access frequency instead of relying primarily on age and refault feedback. The author reports lower page-cache refault rates and higher throughput across constrained kernel builds, database tests, skewed file-cache workloads, and an Android experiment. The strongest result is not any single percentage; it is evidence that MGLRU can misclassify frequently reused file pages when a workload mixes hot data with large scans.
The proposal, called frequency-guided promotion or MGLRU-FG, is not merged and is explicitly labeled RFC. Its benchmark numbers are author-supplied, some tests expose unresolved behavior, and the patch changes accounting visible through /proc/vmstat and cgroup memory.stat. For operators, the practical response is to reproduce the failure mode on representative memory-constrained workloads, collect refault and pressure data, and preserve a rollback path. This is a potentially important reclaim-policy change, not a kernel upgrade recommendation.
Why Recency Alone Is Not Enough
Linux page reclaim must decide which resident memory is likely to be useful again. MGLRU improves that decision by grouping folios into generations that represent similar access recency. Aging discovers accessed pages and moves them toward younger generations; eviction consumes older generations. The design also divides pages into tiers and uses refault feedback to decide which categories deserve protection.
That machinery is more scalable and information-rich than a simple active/inactive list, but recency and frequency answer different questions. A sequential scan can touch a large volume of data recently while a smaller database index or executable working set is accessed repeatedly. If both appear recent, reclaim still needs to distinguish the one-time stream from the repeatedly reused set.
The official MGLRU documentation describes generations as the common time-based frame and tiers as a way to categorize pages accessed through file descriptors. It also describes a feedback controller that compares refault percentages. The RFC argues that this feedback can arrive too late: a useful page may be evicted and refaulted before the system learns that its tier should be protected. Coarse reference state can also make pages with materially different access frequencies look alike.
What Frequency-Guided Promotion Changes
Accesses become an earlier signal
MGLRU-FG gives each folio a compact referenced count and allows repeated accesses to move it through the generation window. Rather than waiting for eviction-time feedback to protect an entire tier, the proposed mechanism promotes an individual folio as evidence accumulates that it belongs to the hot working set.
That distinction matters for page-cache-heavy services. A database can combine random lookups, sequential scans, memory-mapped files, writeback, and application heap pressure in the same memory cgroup. A policy that recognizes only “recent” pages can retain scan pollution. A frequency signal can preferentially retain the smaller set that receives repeated hits.
The series also consolidates reference-state handling used by MGLRU and related code paths. It touches GUP, madvise, DAMON, khugepaged, transparent huge-page splitting, smaps, Btrfs compression, EROFS, and core reclaim code. That breadth explains why the proposal deserves caution even if its benchmark direction is attractive: this is not an isolated tuning knob.
Accounting semantics also change
The RFC proposes deriving active and inactive accounting from per-folio reference state rather than generation-window position. Its author says this avoids abrupt counter changes during aging and misleading behavior on systems without swap. The work also claims to correct under-accounted Pressure Stall Information (PSI) and improve workingset tracking, especially for file cache.
These are operationally significant changes. Capacity controllers, dashboards, anomaly detectors, and autoscaling rules may consume memory.stat, /proc/vmstat, workingset counters, or PSI. A kernel that classifies activity differently can alter those signals even when application behavior has not changed. A test plan therefore needs to validate both workload performance and the meaning of telemetry.
Reading the Early Benchmark Results Correctly
The RFC reports improvements across several test shapes, but the numbers should be interpreted as hypotheses to reproduce rather than universal forecasts.
In a kernel build restricted to a 3 GiB memory cgroup, the patched tree reduced wall time, CPU time, page-ins, swap traffic, and refaults relative to the unpatched MGLRU tree across the tested swappiness values. With disk swap, file refaults fell by 13% and anonymous refaults by 20% in the reported averages. A zram variant showed a similar direction.
A cached fio workload used Zipf-distributed random reads against data larger than its 16 GiB cgroup. Across the listed distributions, the patch increased average IOPS by 11.5% over baseline MGLRU and reduced throughput-normalized file misses by 11%. This is the most direct evidence for the design claim because a skewed distribution deliberately creates frequently reused hot regions among colder data.
The results are not uniformly conclusive. A MongoDB YCSB workload improved 13.4% over baseline MGLRU but remained below classical LRU in that test. The author suspects a writeback-threshold issue, which means the interaction has not been fully explained. An Android backport reduced median file refaults but slightly increased several anonymous-memory and swap counters, while frame-rate differences were negligible. Those mixed signals reinforce the need to evaluate memory policy as a system rather than selecting one favorable throughput number.
The patch also fixes a stark SQLite-plus-grep case constructed to mix a small, repeatedly scanned database region with a cold file stream larger than memory. That benchmark is useful for demonstrating classification failure, but it is deliberately adversarial. Production value depends on whether an organization’s databases, build farms, artifact caches, search nodes, or analytics jobs create a comparable access pattern.
Where Enterprises Could See an Effect
Consolidated services under memcg pressure
The most relevant environment is not an unconstrained server with ample free RAM. It is a host that deliberately overcommits memory or packs services into cgroups with hard limits. Kubernetes nodes, CI builders, multi-tenant database platforms, and large JVM fleets can all combine useful file cache with anonymous memory and periodic scans. In those systems, better working-set detection can translate into fewer storage reads, less reclaim CPU, and lower tail latency.
Fast storage does not eliminate reclaim cost
NVMe can make refaults look inexpensive in bandwidth averages, but a refault still consumes queue capacity, CPU time, and latency budget. On high-throughput storage, kernel bookkeeping and lock contention may become visible before raw media latency does. Conversely, slow or remote storage can magnify a classification error. Benchmarks should therefore record device utilization and latency distributions, not only application throughput.
Observability pipelines may need recalibration
PSI is designed to quantify time lost because tasks are stalled on CPU, memory, or I/O contention. If the RFC’s accounting fixes are accepted, historical baselines may not be directly comparable across kernel versions. The same applies to active/inactive memory and workingset counters. Fleet rollout analysis should segment metrics by kernel build and avoid interpreting a discontinuity as an application regression without checking counter semantics.
A Production-Grade Evaluation Plan
1. Reproduce the access pattern
Start with a workload that contains both a stable hot set and a larger cold or streaming component. Run it inside the same cgroup limits, swap configuration, storage stack, filesystem, and NUMA placement used in production. A generic fio run is useful for mechanics but cannot replace the application’s real read/write mix.
2. Compare three policies where possible
Compare the fleet kernel, an otherwise identical tree with current MGLRU, and the RFC patch set. Keep configuration, compiler, firmware, and workload inputs fixed. Repeat runs after warm-up, report dispersion, and separate application throughput from kernel CPU consumption. Because the proposal remains under review, test only disposable or isolated systems.
3. Measure reclaim, pressure, and service outcomes
At minimum, collect file and anonymous workingset refaults, page-ins, swap-ins and swap-outs, major faults, reclaim CPU, I/O latency, and memory PSI some and full. Pair those with service-level metrics such as p95 and p99 latency, completed jobs, request errors, and throughput. A reduction in refaults is valuable only if it improves the service or reduces resource cost without shifting harm elsewhere.
4. Test phase changes
Frequency-biased protection can be excellent for a stable hot set yet slow to react when the working set changes. The RFC notes that refault distance is not included in the current series. Tests should rotate datasets, restart services, change tenant mix, trigger backup scans, and move between low and severe pressure. Measure recovery time after each transition.
5. Validate telemetry contracts
Diff /proc/vmstat, cgroup memory.stat, and PSI behavior before using a patched kernel in any automated controller. Confirm that alert thresholds, load shedding, and capacity models still represent the intended condition. Treat monitoring compatibility as an acceptance criterion, not an afterthought.
What to Watch Before Adoption
The review process may alter thresholds, data structures, counter semantics, or the scope of the series. Parts may be split and merged independently. The current patch is based on an mm-new development tree, not a stable distribution kernel, and broad subsystem touch points increase regression-testing requirements.
Infrastructure teams should watch for independent benchmark reports, maintainer feedback on accounting semantics, results under NUMA and heavy writeback, and tests that combine anonymous and file pressure. Distribution backports would require their own validation because reclaim behavior depends on the surrounding memory-management code.
The Architectural Takeaway
MGLRU-FG is best understood as a proposal to shorten the control loop in Linux page reclaim. Existing MGLRU uses generations, tiers, and refault feedback to make more informed eviction decisions. Frequency-guided promotion attempts to act on repeated access before useful data reaches the point of eviction.
If the approach survives review, the benefit will not be a universal double-digit speedup. It will be fewer costly classification errors in workloads where a compact hot set competes with scans, builds, or other high-churn data. That is a credible enterprise problem. For now, the right engineering posture is controlled measurement: validate the working-set model, track telemetry changes, test phase shifts, and keep the RFC boundary clear.
