Skip to content

Improve WSL2 guest memory reclaim - #41096

Merged
Ben Hillis (benhillis) merged 3 commits into
masterfrom
user/benhill/improve-memory-reclaim
Jul 24, 2026
Merged

Ben Hillis (benhillis) merged 3 commits into
masterfrom
user/benhill/improve-memory-reclaim

Conversation

@benhillis

@benhillis Ben Hillis (benhillis) commented Jul 16, 2026 •

Copy link
Copy Markdown
Member

What this changes

This replaces the existing memory reduction loop with a shared idle detector and simpler reclaim policy.

  • CPU idle detection uses aggregate busy time instead of user time only.
  • Free pages are compacted when the VM enters idle and after every reclaim operation.
  • Cache reclaim waits for two minutes of sustained idle, protecting normal pauses between commands.
  • Reclaim and compaction CPU are excluded from the next idle sample so the thread does not reset its own grace period.

Modes

Mode Behavior after two minutes idle
Disabled No memory reduction thread
DropCache Run drop_caches once for the idle period
Gradual Reclaim cold cache toward a 128 MB floor using RAM-scaled 256 MB to 1 GB steps

Gradual uses memory.reclaim with swappiness=0, which prevents proactive reclaim from swapping anonymous memory. If memory.reclaim is unavailable, it falls back to DropCache.

The default remains DropCache.

Why

The current implementation waits approximately ten minutes before dropping cache and only measures user CPU time. This leaves vmmemWSL elevated long after real workloads finish and can treat kernel or I/O work as idle.

The new policy returns memory sooner while preserving cache during short user pauses.

DropCache results

Measured with two clean 16-way Linux kernel builds, two read-heavy source archive runs, and three warm-cache reuse runs per configuration.

Metric Existing DropCache New DropCache
Kernel build: 50% working-set reduction 591s 177s
Source archive: 50% working-set reduction 542s 142s
Warm reread after 60 seconds idle 0.233s 0.220s
Kernel working set after 15 minutes 2.14 GiB 2.12 GiB
Source archive working set after 15 minutes 1.99 GiB 1.98 GiB

The new DropCache policy reaches the existing eventual footprint roughly seven minutes sooner without affecting the 60-second warm-cache case.

Gradual results

A focused test using the same two-minute grace period showed:

  • warm rereads after 60 seconds remained at 0.22-0.23 seconds
  • 1.83 GiB of reclaimable cache remained intact through 120 seconds
  • reclaim began at approximately 129 seconds
  • cache fell to 0.14 GiB by 140 seconds
  • vmmemWSL fell from 3.23 GiB to 1.63 GiB by 150 seconds

Gradual reaches a lower sustained footprint because it continues reclaiming regrown cache toward the floor. DropCache remains the default for compatibility and Gradual remains available for users who prioritize minimum memory usage.

Validation

  • cmake --build . --config Debug -- -m
  • built and deployed the generated MSI
  • verified memory.reclaim swappiness=0 support on the WSL 6.18 kernel
  • collected host vmmemWSL, guest memory, reclaim, compaction, workload timing, and warm-cache reuse data

@benhillis
Ben Hillis (benhillis) requested a review from a team as a code owner July 16, 2026 19:43
Copilot AI review requested due to automatic review settings July 16, 2026 19:43

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR updates WSL2’s guest-side memory reduction logic to reclaim memory more efficiently and safely while the VM is idle, aiming to reduce host working set without hurting short-pause resume performance.

Changes:

  • Replaces the user-CPU ring-buffer heuristic with aggregate busy/idle CPU sampling from /proc/stat.
  • Adds reclaim and compaction behavior tuned for idle periods, including a 2-minute grace period and RAM-scaled reclaim requests via memory.reclaim (with fallback to drop_caches).
  • Improves robustness of /proc/stat and /proc/meminfo parsing to avoid terminating the background thread on malformed input.

Comment thread src/linux/init/util.cpp
@benhillis
Ben Hillis (benhillis) force-pushed the user/benhill/improve-memory-reclaim branch from b1f6752 to f4c4747 Compare July 16, 2026 20:14
Copilot-Session: ccc721ed-5fa2-41b2-a5f9-5a8502cdf364
Copilot AI review requested due to automatic review settings July 17, 2026 23:52

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 1 out of 1 changed files in this pull request and generated 2 comments.

Comment thread src/linux/init/util.cpp
Comment thread src/linux/init/util.cpp
Copilot-Session: ccc721ed-5fa2-41b2-a5f9-5a8502cdf364
Copilot AI review requested due to automatic review settings July 20, 2026 16:55

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 1 out of 1 changed files in this pull request and generated 1 comment.

Comment thread src/linux/init/util.cpp

@dkbennett David Bennett (dkbennett) left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good, recommend adding unit tests for CpuIdleTracker. It would need to be pulled out of the anonymous namespace but could be unit tested for confidence it is correctly tracking these states as expected.

@OneBlue Blue (OneBlue) left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, minor suggestions

Comment thread src/linux/init/util.cpp
return false;
}

unsigned long long fields[8] = {};

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit: This might be a good candidate for using std::regex (can be done in a followup though)

Comment thread src/linux/init/util.cpp
}
else if (!droppedThisIdlePeriod)
{
if (WriteToFile("/proc/sys/vm/drop_caches", "1\n") == 0)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Should we switch this to 3 while we're at it to drop metadata as well ?

@benhillis

Copy link
Copy Markdown
Member Author

Good calls, will address as a follow-up.

@benhillis
Ben Hillis (benhillis) merged commit 49f841c into master Jul 24, 2026
12 checks passed
@benhillis
Ben Hillis (benhillis) deleted the user/benhill/improve-memory-reclaim branch July 24, 2026 22:20
Ben Hillis (benhillis) added a commit that referenced this pull request Jul 27, 2026
* Address follow-ups from PR #41096

- Parse the /proc/stat cpu line with std::regex instead of a manual
  strtoull cursor loop, while retaining the O_CLOEXEC bounded read.
- Use drop_caches=3 in DropCache mode to also drop reclaimable slab
  (dentries/inodes), matching the SReclaimable slab counted by
  GetReclaimableCacheBytes.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 0e37f3d5-68e7-4973-8c27-bf4329c8fd9b

* Restrict /proc/stat field separators to spaces/tabs

ECMAScript \s matches newlines, so on a truncated aggregate cpu line the
optional irq/softirq/steal groups could consume digits from the next line
in the read buffer. Use [ \t] so the match cannot span lines.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 0e37f3d5-68e7-4973-8c27-bf4329c8fd9b

---------

Co-authored-by: Ben Hillis 
Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 0e37f3d5-68e7-4973-8c27-bf4329c8fd9b
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants