Repository navigation
Expand file tree
/
Copy pathKconfig.mem
More file actions
185 lines (159 loc) · 8.06 KB
/
Copy pathKconfig.mem
File metadata and controls
185 lines (159 loc) · 8.06 KB
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
#
# Memory options — where a slab's one contiguous reservation comes from, and
# the two kernel hints laid over it. They belong to this package, and this file
# is the single place they are defined: hpc's own Kconfig sources it into
# "Library options", and a tree that embeds hpc (un) sources the same file
# rather than keeping a copy in step by hand.
#
# The file carries no menu of its own so each package can place these where its
# own menu structure wants them.
#
choice
prompt "Slab reservation backend"
default MEM_SLAB_MMAP
help
How the slab allocators in <mem/slab.h> and <mem/slab_class.h> obtain
the single contiguous reservation they hand blocks out of. Both ask
the host for the same three things - reserve @len bytes once, give
them back, and hand the physical pages of a range back while keeping
the addresses mapped - and this picks who answers.
It is a build-time choice because the reservation is sized once at
init and never resized: that is what keeps a block's address stable
for the lifetime of the slab. For the same reason there is no
realloc() backend and cannot be one - realloc() may move the
reservation, and a move invalidates every block pointer a caller is
holding. A slab that can be reallocated is a different data structure
with a different contract (index-only access, no raw pointers held
across a grow).
Neither choice here is the last word: the SLAB_VM_* hooks in
<mem/slab_vm.h> still take a custom backend - a static arena, a
hugetlbfs or shared-memory segment, a freestanding target with no
mmap() at all - and a definition supplied by the code that wants it
wins over both of these.
config MEM_SLAB_MMAP
bool "mmap() anonymous mapping"
help
Reserve with mmap(MAP_ANON | MAP_NORESERVE), release with munmap(),
and hand pages back on shrink with madvise(MADV_DONTNEED), which
keeps the addresses mapped while the memory returns to the host.
This is the backend the slab was written for and the only one that
gives memory back: a shrink that drops blocks actually lowers RSS.
It is also the only one the two hints below apply to, since they are
madvise() calls over the mapping.
One property worth naming because code comes to depend on it without
meaning to: a fresh anonymous mapping reads as zeroes. The slab does
not rely on that - its bitmap and cache entries come zeroed from
SLAB_MEM_CALLOC and a block is payload the caller initialises - but a
caller that grew up here and quietly assumed a zeroed block will not
survive a switch to the heap.
config MEM_SLAB_MALLOC
bool "malloc()/free() (libc heap)"
help
Reserve with malloc() and release with free(), the same backend
-DSLAB_MALLOC_FREE has always selected by hand. For a target where
mmap() is unavailable or unwanted, and for a build whose allocator is
the thing being instrumented - valgrind, an ASan/heap profiler, an
embedded libc with its own arena - where a slab that reserves behind
the allocator's back is invisible to the tool.
What it gives up is releasing memory. The heap has no equivalent of
madvise(MADV_DONTNEED) over part of a live allocation, so
SLAB_VM_RELEASE is a no-op: shrink still drops blocks from the
committed prefix and rebuilds the free list, the block counts still
move, and the process gives nothing back - RSS stays at its
high-water mark for the life of the slab. Everything else behaves the
same, and the slab unit tests pass on it.
Two smaller differences: heap memory is not zeroed (see above), and
the release grain drops to a page since neither madvise() hint
applies, which is the same grain an mmap build without MEM_HUGEPAGE
uses.
endchoice
choice
prompt "Default slab grow/shrink policy"
default MEM_SLAB_POLICY_GRADUAL
help
The policy a slab gets when its owner does not choose one: what
SLAB_POLICY_DEFAULT() in <mem/slab.h> expands to, and what the
empty policy name resolves to in slab_policy_preset() - which is
where a program's own configuration ("cache-policy = eager")
lands when it names no preset. A slab initialized with an
explicit policy is not affected.
The choice is a factor: how much of the distance between the
current working set and the demand one grow or shrink step
covers. The watermarks are the same in every preset so the
factor is the only difference.
config MEM_SLAB_POLICY_GRADUAL
bool "Gradual - by halves"
help
At 75% usage commit half of the remaining headroom; once usage
has stayed at/under 25% for the dwell below, release half of
the free blocks per interval. Committed walks toward the demand
from either side, halving the distance each step - responsive
without ever betting everything on one sample. The default, and
the right answer for load that moves.
config MEM_SLAB_POLICY_EAGER
bool "Eager - whole steps"
help
The first grow commits everything up to the maximum, and one gc
pass after the dwell returns every free block at once, down to
the minimum or the live set. For work that is bursty and
all-or-nothing, where holding half measures between bursts is
just resident memory.
config MEM_SLAB_POLICY_STATIC
bool "Static - a fixed working set"
help
Commit the whole reservation up front and never grow or shrink.
Memory use is decided at init and stays put: no faults after
startup, no release ever - the policy for a deployment sized by
contract rather than by load.
endchoice
config MEM_SLAB_SHRINK_AFTER
int "Idle dwell before releasing memory (ms)"
depends on !MEM_SLAB_POLICY_STATIC
default 30000
help
How long usage must stay at/under the shrink watermark before
the default policy releases memory. This is hysteresis: it stops
a transient dip from thrashing the working set, and it is also
the rate limit - at most one release per interval. 30 seconds
suits a daemon watching traffic; an interactive tool that should
give memory back promptly wants less.
config MEM_HUGEPAGE
bool "Back slab reservations with transparent huge pages"
depends on MEM_SLAB_MMAP
help
Ask for MADV_HUGEPAGE over each slab reservation, so the kernel backs
it with 2 MB pages where it can. The point is TLB reach: a slab large
enough to be randomly accessed across gigabytes spends real time on TLB
misses that huge pages avoid.
The cost is granularity, and it is not subtle. A huge page is faulted,
and released, 2 MB at a time, so this also makes 2 MB the slab's release
grain: min, max and grow_step are rounded up to a whole 2 MB, and a slab
can no longer shrink below it. A policy asking for a 64 block minimum of
2048 byte blocks gets 1024 blocks - 2 MB - because that is what the
kernel will fault in either way, and the block counts the slab reports
are meant to mean the memory it is holding. Read the rounded policy back
with slab_policy_min()/slab_policy_max().
So: worth it for a slab that stays large and hot, where TLB reach is the
cost that matters. Wrong for many small slabs, which it would round up
to 2 MB apiece.
Requires transparent_hugepage=madvise or always. Where the kernel cannot
honour it the hint fails and is ignored - but the rounding still applies,
because it is a build-time decision.
If unsure, say N.
config MEM_POPULATE
bool "Pre-fault slab blocks when committing them"
depends on MEM_SLAB_MMAP
help
Ask for MADV_POPULATE_WRITE over the blocks each grow step commits, so
one syscall faults the range instead of taking a page fault per page.
This applies only where a block is no larger than a page, and is
compiled out entirely otherwise. Committing a block writes a free-list
node into it, which for a sub-page block already touches every page in
the range - that is the fault-per-page this replaces, measured about
1.7x cheaper for a 2048 byte block. A block spanning several pages is
the opposite case: the node write touches only its first page, so
pre-faulting the rest commits memory the caller may never touch and
measures slower.
Needs Linux 5.14 or newer. On an older kernel the hint fails and blocks
are faulted lazily as before.
If unsure, say N.