xref: /linux/Documentation/admin-guide/mm/userfaultfd.rst (revision 570f7e331f5febb30f1384817463c7e42b65ca7d)
1===========
2Userfaultfd
3===========
4
5Objective
6=========
7
8Userfaults allow the implementation of on-demand paging from userland
9and more generally they allow userland to take control of various
10memory page faults, something otherwise only the kernel code could do.
11
12For example userfaults allows a proper and more optimal implementation
13of the ``PROT_NONE+SIGSEGV`` trick.
14
15Design
16======
17
18Userspace creates a new userfaultfd, initializes it, and registers one or more
19regions of virtual memory with it. Then, any page faults which occur within the
20region(s) result in a message being delivered to the userfaultfd, notifying
21userspace of the fault.
22
23The ``userfaultfd`` (aside from registering and unregistering virtual
24memory ranges) provides two primary functionalities:
25
261) ``read/POLLIN`` protocol to notify a userland thread of the faults
27   happening
28
292) various ``UFFDIO_*`` ioctls that can manage the virtual memory regions
30   registered in the ``userfaultfd`` that allows userland to efficiently
31   resolve the userfaults it receives via 1) or to manage the virtual
32   memory in the background
33
34The real advantage of userfaults if compared to regular virtual memory
35management of mremap/mprotect is that the userfaults in all their
36operations never involve heavyweight structures like vmas (in fact the
37``userfaultfd`` runtime load never takes the mmap_lock for writing).
38Vmas are not suitable for page- (or hugepage) granular fault tracking
39when dealing with virtual address spaces that could span
40Terabytes. Too many vmas would be needed for that.
41
42The ``userfaultfd``, once created, can also be
43passed using unix domain sockets to a manager process, so the same
44manager process could handle the userfaults of a multitude of
45different processes without them being aware about what is going on
46(well of course unless they later try to use the ``userfaultfd``
47themselves on the same region the manager is already tracking, which
48is a corner case that would currently return ``-EBUSY``).
49
50API
51===
52
53Creating a userfaultfd
54----------------------
55
56There are two ways to create a new userfaultfd, each of which provide ways to
57restrict access to this functionality (since historically userfaultfds which
58handle kernel page faults have been a useful tool for exploiting the kernel).
59
60The first way, supported since userfaultfd was introduced, is the
61userfaultfd(2) syscall. Access to this is controlled in several ways:
62
63- Any user can always create a userfaultfd which traps userspace page faults
64  only. Such a userfaultfd can be created using the userfaultfd(2) syscall
65  with the flag UFFD_USER_MODE_ONLY.
66
67- In order to also trap kernel page faults for the address space, either the
68  process needs the CAP_SYS_PTRACE capability, or the system must have
69  vm.unprivileged_userfaultfd set to 1. By default, vm.unprivileged_userfaultfd
70  is set to 0.
71
72The second way, added to the kernel more recently, is by opening
73/dev/userfaultfd and issuing a USERFAULTFD_IOC_NEW ioctl to it. This method
74yields equivalent userfaultfds to the userfaultfd(2) syscall.
75
76Unlike userfaultfd(2), access to /dev/userfaultfd is controlled via normal
77filesystem permissions (user/group/mode), which gives fine grained access to
78userfaultfd specifically, without also granting other unrelated privileges at
79the same time (as e.g. granting CAP_SYS_PTRACE would do). Users who have access
80to /dev/userfaultfd can always create userfaultfds that trap kernel page faults;
81vm.unprivileged_userfaultfd is not considered.
82
83Initializing a userfaultfd
84--------------------------
85
86When first opened the ``userfaultfd`` must be enabled invoking the
87``UFFDIO_API`` ioctl specifying a ``uffdio_api.api`` value set to ``UFFD_API`` (or
88a later API version) which will specify the ``read/POLLIN`` protocol
89userland intends to speak on the ``UFFD`` and the ``uffdio_api.features``
90userland requires. The ``UFFDIO_API`` ioctl if successful (i.e. if the
91requested ``uffdio_api.api`` is spoken also by the running kernel and the
92requested features are going to be enabled) will return into
93``uffdio_api.features`` and ``uffdio_api.ioctls`` two 64bit bitmasks of
94respectively all the available features of the read(2) protocol and
95the generic ioctl available.
96
97The ``uffdio_api.features`` bitmask returned by the ``UFFDIO_API`` ioctl
98defines what memory types are supported by the ``userfaultfd`` and what
99events, except page fault notifications, may be generated:
100
101- The ``UFFD_FEATURE_EVENT_*`` flags indicate that various other events
102  other than page faults are supported. These events are described in more
103  detail below in the `Non-cooperative userfaultfd`_ section.
104
105- ``UFFD_FEATURE_MISSING_HUGETLBFS`` and ``UFFD_FEATURE_MISSING_SHMEM``
106  indicate that the kernel supports ``UFFDIO_REGISTER_MODE_MISSING``
107  registrations for hugetlbfs and shared memory (covering all shmem APIs,
108  i.e. tmpfs, ``IPCSHM``, ``/dev/zero``, ``MAP_SHARED``, ``memfd_create``,
109  etc) virtual memory areas, respectively.
110
111- ``UFFD_FEATURE_MINOR_HUGETLBFS`` indicates that the kernel supports
112  ``UFFDIO_REGISTER_MODE_MINOR`` registration for hugetlbfs virtual memory
113  areas. ``UFFD_FEATURE_MINOR_SHMEM`` is the analogous feature indicating
114  support for shmem virtual memory areas.
115
116- ``UFFD_FEATURE_MOVE`` indicates that the kernel supports moving an
117  existing page contents from userspace.
118
119The userland application should set the feature flags it intends to use
120when invoking the ``UFFDIO_API`` ioctl, to request that those features be
121enabled if supported.
122
123Once the ``userfaultfd`` API has been enabled the ``UFFDIO_REGISTER``
124ioctl should be invoked (if present in the returned ``uffdio_api.ioctls``
125bitmask) to register a memory range in the ``userfaultfd`` by setting the
126uffdio_register structure accordingly. The ``uffdio_register.mode``
127bitmask will specify to the kernel which kind of faults to track for
128the range. The ``UFFDIO_REGISTER`` ioctl will return the
129``uffdio_register.ioctls`` bitmask of ioctls that are suitable to resolve
130userfaults on the range registered. Not all ioctls will necessarily be
131supported for all memory types (e.g. anonymous memory vs. shmem vs.
132hugetlbfs), or all types of intercepted faults.
133
134.. note::
135
136   Re-registering an already-registered range must not drop any of the
137   modes that install per-PTE markers — currently
138   ``UFFDIO_REGISTER_MODE_WP`` and ``UFFDIO_REGISTER_MODE_RWP``. Doing
139   so would strand markers with no flag to describe them, so the call
140   is rejected with ``-EBUSY``; userspace must issue
141   ``UFFDIO_UNREGISTER`` first. This differs from older kernels, which
142   silently replaced the mode bits on re-registration.
143
144Userland can use the ``uffdio_register.ioctls`` to manage the virtual
145address space in the background (to add or potentially also remove
146memory from the ``userfaultfd`` registered range). This means a userfault
147could be triggering just before userland maps in the background the
148user-faulted page.
149
150Resolving Userfaults
151--------------------
152
153There are three basic ways to resolve userfaults:
154
155- ``UFFDIO_COPY`` atomically copies some existing page contents from
156  userspace.
157
158- ``UFFDIO_ZEROPAGE`` atomically zeros the new page.
159
160- ``UFFDIO_CONTINUE`` maps an existing, previously-populated page.
161
162These operations are atomic in the sense that they guarantee nothing can
163see a half-populated page, since readers will keep userfaulting until the
164operation has finished.
165
166By default, these wake up userfaults blocked on the range in question.
167They support a ``UFFDIO_*_MODE_DONTWAKE`` ``mode`` flag, which indicates
168that waking will be done separately at some later time.
169
170Which ioctl to choose depends on the kind of page fault, and what we'd
171like to do to resolve it:
172
173- For ``UFFDIO_REGISTER_MODE_MISSING`` faults, the fault needs to be
174  resolved by either providing a new page (``UFFDIO_COPY``), or mapping
175  the zero page (``UFFDIO_ZEROPAGE``). By default, the kernel would map
176  the zero page for a missing fault. With userfaultfd, userspace can
177  decide what content to provide before the faulting thread continues.
178
179- For ``UFFDIO_REGISTER_MODE_MINOR`` faults, there is an existing page (in
180  the page cache). Userspace has the option of modifying the page's
181  contents before resolving the fault. Once the contents are correct
182  (modified or not), userspace asks the kernel to map the page and let the
183  faulting thread continue with ``UFFDIO_CONTINUE``.
184
185Notes:
186
187- You can tell which kind of fault occurred by examining
188  ``pagefault.flags`` within the ``uffd_msg``, checking for the
189  ``UFFD_PAGEFAULT_FLAG_*`` flags.
190
191- None of the page-delivering ioctls default to the range that you
192  registered with.  You must fill in all fields for the appropriate
193  ioctl struct including the range.
194
195- You get the address of the access that triggered the missing page
196  event out of a struct uffd_msg that you read in the thread from the
197  uffd.  You can supply as many pages as you want with these IOCTLs.
198  Keep in mind that unless you used DONTWAKE then the first of any of
199  those IOCTLs wakes up the faulting thread.
200
201- Be sure to test for all errors including
202  (``pollfd[0].revents & POLLERR``).  This can happen, e.g. when ranges
203  supplied were incorrect.
204
205Write Protect Notifications
206---------------------------
207
208This is equivalent to (but faster than) using mprotect and a SIGSEGV
209signal handler.
210
211Firstly you need to register a range with ``UFFDIO_REGISTER_MODE_WP``.
212Instead of using mprotect(2) you use
213``ioctl(uffd, UFFDIO_WRITEPROTECT, struct *uffdio_writeprotect)``
214while ``mode = UFFDIO_WRITEPROTECT_MODE_WP``
215in the struct passed in.  The range does not default to and does not
216have to be identical to the range you registered with.  You can write
217protect as many ranges as you like (inside the registered range).
218Then, in the thread reading from uffd the struct will have
219``msg.arg.pagefault.flags & UFFD_PAGEFAULT_FLAG_WP`` set. Now you send
220``ioctl(uffd, UFFDIO_WRITEPROTECT, struct *uffdio_writeprotect)``
221again while ``pagefault.mode`` does not have ``UFFDIO_WRITEPROTECT_MODE_WP``
222set. This wakes up the thread which will continue to run with writes. This
223allows you to do the bookkeeping about the write in the uffd reading
224thread before the ioctl.
225
226If you registered with both ``UFFDIO_REGISTER_MODE_MISSING`` and
227``UFFDIO_REGISTER_MODE_WP`` then you need to think about the sequence in
228which you supply a page and undo write protect.  Note that there is a
229difference between writes into a WP area and into a !WP area.  The
230former will have ``UFFD_PAGEFAULT_FLAG_WP`` set, the latter
231``UFFD_PAGEFAULT_FLAG_WRITE``.  The latter did not fail on protection but
232you still need to supply a page when ``UFFDIO_REGISTER_MODE_MISSING`` was
233used.
234
235Userfaultfd write-protect mode currently behave differently on none ptes
236(when e.g. page is missing) over different types of memories.
237
238For anonymous memory, ``ioctl(UFFDIO_WRITEPROTECT)`` will ignore none ptes
239(e.g. when pages are missing and not populated).  For file-backed memories
240like shmem and hugetlbfs, none ptes will be write protected just like a
241present pte.  In other words, there will be a userfaultfd write fault
242message generated when writing to a missing page on file typed memories,
243as long as the page range was write-protected before.  Such a message will
244not be generated on anonymous memories by default.
245
246If the application wants to be able to write protect none ptes on anonymous
247memory, one can pre-populate the memory with e.g. MADV_POPULATE_READ.  On
248newer kernels, one can also detect the feature UFFD_FEATURE_WP_UNPOPULATED
249and set the feature bit in advance to make sure none ptes will also be
250write protected even upon anonymous memory.
251
252When using ``UFFDIO_REGISTER_MODE_WP`` in combination with either
253``UFFDIO_REGISTER_MODE_MISSING`` or ``UFFDIO_REGISTER_MODE_MINOR``, when
254resolving missing / minor faults with ``UFFDIO_COPY`` or ``UFFDIO_CONTINUE``
255respectively, it may be desirable for the new page / mapping to be
256write-protected (so future writes will also result in a WP fault). These ioctls
257support a mode flag (``UFFDIO_COPY_MODE_WP`` or ``UFFDIO_CONTINUE_MODE_WP``
258respectively) to configure the mapping this way.
259
260If the userfaultfd context has ``UFFD_FEATURE_WP_ASYNC`` feature bit set,
261any vma registered with write-protection will work in async mode rather
262than the default sync mode.
263
264In async mode, there will be no message generated when a write operation
265happens, meanwhile the write-protection will be resolved automatically by
266the kernel.  It can be seen as a more accurate version of soft-dirty
267tracking and it can be different in a few ways:
268
269  - The dirty result will not be affected by vma changes (e.g. vma
270    merging) because the dirty is only tracked by the pte.
271
272  - It supports range operations by default, so one can enable tracking on
273    any range of memory as long as page aligned.
274
275  - Dirty information will not get lost if the pte was zapped due to
276    various reasons (e.g. during split of a shmem transparent huge page).
277
278  - Due to a reverted meaning of soft-dirty (page clean when the uffd bit
279    is set; dirty when the uffd bit is cleared), it has different semantics
280    on some of the memory operations.  For example: ``MADV_DONTNEED`` on
281    anonymous (or ``MADV_REMOVE`` on a file mapping) will be treated as
282    dirtying of memory by dropping the uffd bit during the procedure.
283
284The user app can collect the "written/dirty" status by looking up the
285uffd bit for the pages being interested in /proc/pagemap.
286
287The page will not be under track of userfaultfd-wp async mode until the page is
288explicitly write-protected by ``ioctl(UFFDIO_WRITEPROTECT)`` with the mode
289flag ``UFFDIO_WRITEPROTECT_MODE_WP`` set.  Trying to resolve a page fault
290that was tracked by async mode userfaultfd-wp is invalid.
291
292When userfaultfd-wp async mode is used alone, it can be applied to all
293kinds of memory.
294
295Memory Poisioning Emulation
296---------------------------
297
298In response to a fault (either missing or minor), an action userspace can
299take to "resolve" it is to issue a ``UFFDIO_POISON``. This will cause any
300future faulters to either get a SIGBUS, or in KVM's case the guest will
301receive an MCE as if there were hardware memory poisoning.
302
303This is used to emulate hardware memory poisoning. Imagine a VM running on a
304machine which experiences a real hardware memory error. Later, we live migrate
305the VM to another physical machine. Since we want the migration to be
306transparent to the guest, we want that same address range to act as if it was
307still poisoned, even though it's on a new physical host which ostensibly
308doesn't have a memory error in the exact same spot.
309
310Read-Write Protection
311---------------------
312
313``UFFDIO_REGISTER_MODE_RWP`` enables read-write protection tracking on a
314memory range. It is similar to (but faster than) ``mprotect(PROT_NONE)``
315combined with a signal handler; unlike ``mprotect(PROT_NONE)``, RWP only
316traps accesses to *present* PTEs, so accesses to unpopulated addresses in a
317protected range fall through to the normal missing-page path. It uses the
318PROT_NONE hinting mechanism (same as NUMA balancing) to make pages
319inaccessible while keeping them resident in memory. Works on anonymous,
320shmem, and hugetlbfs memory.
321
322RWP is designed for VM memory managers that need to track the working set
323of guest memory for cold page eviction to tiered or remote storage.
324
325**Setup:**
326
3271. Open a userfaultfd and enable ``UFFD_FEATURE_RWP`` via ``UFFDIO_API``.
328   Optionally request ``UFFD_FEATURE_RWP_ASYNC`` as well — it requires
329   ``UFFD_FEATURE_RWP`` to be set in the same ``UFFDIO_API`` call.
330
3312. Register the guest memory range with ``UFFDIO_REGISTER_MODE_RWP``
332   (and ``UFFDIO_REGISTER_MODE_MISSING`` if evicted pages will need to be
333   fetched back from storage).
334
335**Feature availability:**
336
337RWP is built on top of two kernel primitives: a spare PTE bit owned by
338userfaultfd (``CONFIG_HAVE_ARCH_USERFAULTFD_WP``) and architecture support
339for present-but-inaccessible PTEs (``CONFIG_ARCH_HAS_PTE_PROTNONE``). When both
340are available on a 64-bit kernel, the build selects
341``CONFIG_USERFAULTFD_RWP=y`` and the ``VM_UFFD_RWP`` VMA flag becomes
342available.
343
344``UFFD_FEATURE_RWP`` and ``UFFD_FEATURE_RWP_ASYNC`` are unavailable when
345the running kernel or architecture does not support them — for example
34632-bit kernels (where ``VM_UFFD_RWP`` is unavailable), kernels built
347without ``CONFIG_USERFAULTFD_RWP``, and architectures whose ptes cannot
348carry the uffd bit at runtime (e.g. riscv without the ``SVRSW60T59B``
349extension). Requesting an unsupported feature in
350``uffdio_api.features`` makes ``UFFDIO_API`` fail with ``EINVAL`` and
351leaves the userfaultfd context uninitialized; the structure is returned
352zeroed, so the error path cannot be used to discover what the kernel
353supports. The recommended probe sequence is therefore to open a
354throwaway userfaultfd, call ``UFFDIO_API`` once with ``features = 0``,
355inspect the returned bitmask, close that fd, then open the real one
356and call ``UFFDIO_API`` again with only the supported features set.
357
358**Protecting and Unprotecting:**
359
360Use ``UFFDIO_RWPROTECT`` to protect or unprotect a range, mirroring the
361``UFFDIO_WRITEPROTECT`` interface::
362
363    struct uffdio_rwprotect rwp = {
364        .range = { .start = addr, .len = len },
365        .mode = UFFDIO_RWPROTECT_MODE_RWP,  /* protect */
366    };
367    ioctl(uffd, UFFDIO_RWPROTECT, &rwp);
368
369Setting ``UFFDIO_RWPROTECT_MODE_RWP`` sets PROT_NONE on present PTEs in the
370range. Pages stay resident and their physical frames are preserved — only
371access permissions are removed.
372
373Clearing ``UFFDIO_RWPROTECT_MODE_RWP`` restores normal VMA permissions and
374wakes any faulting threads (unless ``UFFDIO_RWPROTECT_MODE_DONTWAKE`` is set).
375
376**Scope of protection:**
377
378RWP protection is a property of *present* PTEs. ``UFFDIO_RWPROTECT`` only
379affects entries that are already populated. Unpopulated addresses within
380the range remain unpopulated; when first accessed they fault through the
381normal missing path (``do_anonymous_page()``, ``do_swap_page()``,
382``finish_fault()``) and the resulting PTE is not RWP-protected. To observe
383the population itself, co-register the range with
384``UFFDIO_REGISTER_MODE_MISSING``.
385
386Protection is preserved across page reclaim: a page swapped out while
387RWP-protected carries the marker on its swap entry, and swap-in restores
388the PROT_NONE state so the first access after swap-in still faults. The
389same applies to pages temporarily replaced by migration entries.
390
391Operations that drop the PTE entirely — ``MADV_DONTNEED`` on anonymous
392memory, hole-punch on shmem, truncation of a file mapping — also drop the
393RWP marker: the next access re-populates the range without protection.
394Unlike WP (which persists via ``PTE_MARKER_UFFD_WP``), there is no
395persistent RWP marker today. The user needs to re-arm the range with
396``UFFDIO_RWPROTECT`` after any operation that explicitly frees PTEs.
397
398**Fault Handling:**
399
400When a protected page is accessed:
401
402- **Sync mode** (default): The faulting thread blocks and a
403  ``UFFD_PAGEFAULT_FLAG_RWP`` message is delivered to the userfaultfd
404  handler. The handler resolves the fault with ``UFFDIO_RWPROTECT``
405  (clearing ``MODE_RWP``), which restores the PTE permissions and wakes
406  the faulting thread.
407
408- **Async mode** (``UFFD_FEATURE_RWP_ASYNC``): The kernel automatically
409  restores PTE permissions and the thread continues without blocking. No
410  message is delivered to the handler.
411
412**Runtime Mode Switching:**
413
414``UFFDIO_SET_MODE`` toggles ``UFFD_FEATURE_RWP_ASYNC`` at runtime, allowing
415the VMM to switch between lightweight async detection and safe sync
416eviction without re-registering. The toggle takes ``mmap_write_lock()``
417and calls ``vma_start_write()`` on each UFFD-armed VMA, draining
418in-flight per-VMA-locked faults before the new mode takes effect.
419
420**Working-set detection with PAGEMAP_SCAN:**
421
422RWP-protected PTEs carry the uffd PTE bit; an access (and, in async mode, its
423auto-resolution) clears it. ``PAGEMAP_SCAN`` reports ``PAGE_IS_ACCESSED`` once
424the bit is clear on a ``VM_UFFD_RWP`` VMA, so a *non-inverted* scan reports the
425pages that were touched during the interval -- the hot set::
426
427    struct pm_scan_arg arg = {
428        .size = sizeof(arg),
429        .start = guest_mem_start,
430        .end = guest_mem_end,
431        .vec = (uint64_t)regions,
432        .vec_len = regions_len,
433        .category_mask = PAGE_IS_ACCESSED,
434        .return_mask = PAGE_IS_ACCESSED,
435    };
436    long n = ioctl(pagemap_fd, PAGEMAP_SCAN, &arg);
437
438The returned ``page_region`` array lists the hot ranges. ``PAGE_IS_ACCESSED``
439is set on an accessed page whether it is still present or has since been
440swapped out, so the hot scan needs no ``PAGE_IS_PRESENT`` filter -- unpopulated
441holes carry neither bit and are excluded on their own.
442
443Track the hot set and reclaim everything else from the backing file (see the
444workflow below). Do **not** invert the scan to enumerate "cold" pages
445directly: an inverted scan reports only the ``VM_UFFD_RWP`` PTEs that are still
446protected, i.e. the resident portion of *this* VMA. For a file mapping the
447working set spans the whole file -- pages that live in the page cache but are
448not mapped into this VMA (a pre-populated tmpfs file, or memory populated
449through another mapping) are ``pte_none`` here, never appear in the scan, and
450would never be considered for eviction even though they occupy memory. Driving
451eviction from "file offsets minus the hot set" avoids that blind spot; a cold
452PTE scan cannot. To additionally record the *first* access to a cached but
453unmapped page (e.g. pre-populated content) as hot, co-register the range with
454``UFFDIO_REGISTER_MODE_MINOR``: such accesses then fault as minor faults
455instead of mapping the page silently.
456
457**Cleanup:**
458
459When the userfaultfd is closed or the range is unregistered, all PROT_NONE
460PTEs are automatically restored to their normal VMA permissions. This
461prevents pages from becoming permanently inaccessible.
462
463**VMM Working Set Tracking Workflow:**
464
465A typical VMM lifecycle for cold page eviction to tiered storage. Two
466mappings of the same shmem (or hugetlbfs) file are used: ``guest_mem`` is
467the RWP-registered mapping that vCPUs access through, and ``io_mem`` is a
468private mapping for VMM-side I/O. Reading ``io_mem`` does not go through
469the RWP-protected PTEs of ``guest_mem``, so the VMM's own ``pwrite()``
470never traps on its own ::
471
472    /* One-time setup */
473    fd = memfd_create("guest", MFD_CLOEXEC);
474    ftruncate(fd, guest_size);
475    guest_mem = mmap(NULL, guest_size, PROT_READ | PROT_WRITE,
476                     MAP_SHARED, fd, 0);  /* vCPU view, RWP-registered */
477    io_mem    = mmap(NULL, guest_size, PROT_READ | PROT_WRITE,
478                     MAP_SHARED, fd, 0);  /* VMM I/O view, unprotected */
479
480    uffd = userfaultfd(O_CLOEXEC | O_NONBLOCK);
481    struct uffdio_api api = {
482        .api = UFFD_API,
483        .features = UFFD_FEATURE_RWP | UFFD_FEATURE_RWP_ASYNC,
484    };
485    ioctl(uffd, UFFDIO_API, &api);
486    if (!(api.features & UFFD_FEATURE_RWP))
487        /* RWP unavailable on this kernel/arch -- fall back. */
488    ioctl(uffd, UFFDIO_REGISTER, &(struct uffdio_register){
489        .range = { guest_mem, guest_size },
490        .mode = UFFDIO_REGISTER_MODE_RWP |
491                UFFDIO_REGISTER_MODE_MISSING,
492    });
493
494    /* Tracking loop */
495    while (vm_running) {
496        /* 1. Detection phase (async -- no vCPU stalls) */
497        ioctl(uffd, UFFDIO_RWPROTECT, &(struct uffdio_rwprotect){
498            .range = full_range,
499            .mode = UFFDIO_RWPROTECT_MODE_RWP });
500        sleep(tracking_interval);
501
502        /*
503         * 2. Switch to sync BEFORE scanning. In async mode a vCPU
504         * access races eviction: it would auto-resolve and mark the
505         * page hot just as the VMM writes it out and punches it,
506         * losing the update. Sync mode makes such accesses block and
507         * be delivered, freezing the hot snapshot for the rest of the
508         * iteration.
509         */
510        ioctl(uffd, UFFDIO_SET_MODE,
511              &(struct uffdio_set_mode){
512                  .disable = UFFD_FEATURE_RWP_ASYNC });
513
514        /* 3. Read the hot set: pages touched this interval. */
515        ioctl(pagemap_fd, PAGEMAP_SCAN, &(struct pm_scan_arg){
516            .category_mask = PAGE_IS_ACCESSED,
517            .return_mask = PAGE_IS_ACCESSED,
518            ...
519        });
520
521        /*
522         * 4. Reclaim the file offsets that are NOT in the hot set.
523         * Driving this from the file's offset space (rather than from a
524         * cold PTE scan) also reclaims pages that are cached but not
525         * mapped into guest_mem, e.g. pre-populated content.
526         */
527        for each non-hot offset range:
528            /* Read from io_mem -- bypasses RWP, no fault. */
529            pwrite(storage_fd, (char *)io_mem + off, len, off);
530            /* Drop the page from the shared file. */
531            fallocate(fd, FALLOC_FL_PUNCH_HOLE | FALLOC_FL_KEEP_SIZE,
532                      off, len);
533            /*
534             * Wake any vCPU blocked on the RWP fault for this range:
535             * fallocate() does not iterate ctx->fault_pending_wqh.
536             */
537            ioctl(uffd, UFFDIO_WAKE, &(struct uffdio_range){
538                .start = (uintptr_t)guest_mem + off, .len = len });
539
540        /* 5. Resume async tracking */
541        ioctl(uffd, UFFDIO_SET_MODE,
542              &(struct uffdio_set_mode){
543                  .enable = UFFD_FEATURE_RWP_ASYNC });
544    }
545
546During step 4, a vCPU that accesses a ``guest_mem`` offset being evicted
547blocks with a ``UFFD_PAGEFAULT_FLAG_RWP`` fault while the eviction is in
548progress. After ``fallocate()`` punches the page out and ``UFFDIO_WAKE``
549fires, the vCPU retries the access, faults as ``MISSING``, and the
550handler resolves it with ``UFFDIO_COPY`` from storage.
551
552This workflow targets shmem and hugetlbfs (both support a private
553``io_mem`` mapping over the same fd). Anonymous-memory backings need a
554different inner-loop strategy because the VMM has no way to read the
555page without going through the RWP-protected mapping.
556
557QEMU/KVM
558========
559
560QEMU/KVM is using the ``userfaultfd`` syscall to implement postcopy live
561migration. Postcopy live migration is one form of memory
562externalization consisting of a virtual machine running with part or
563all of its memory residing on a different node in the cloud. The
564``userfaultfd`` abstraction is generic enough that not a single line of
565KVM kernel code had to be modified in order to add postcopy live
566migration to QEMU.
567
568Guest async page faults, ``FOLL_NOWAIT`` and all other ``GUP*`` features work
569just fine in combination with userfaults. Userfaults trigger async
570page faults in the guest scheduler so those guest processes that
571aren't waiting for userfaults (i.e. network bound) can keep running in
572the guest vcpus.
573
574It is generally beneficial to run one pass of precopy live migration
575just before starting postcopy live migration, in order to avoid
576generating userfaults for readonly guest regions.
577
578The implementation of postcopy live migration currently uses one
579single bidirectional socket but in the future two different sockets
580will be used (to reduce the latency of the userfaults to the minimum
581possible without having to decrease ``/proc/sys/net/ipv4/tcp_wmem``).
582
583The QEMU in the source node writes all pages that it knows are missing
584in the destination node, into the socket, and the migration thread of
585the QEMU running in the destination node runs ``UFFDIO_COPY|ZEROPAGE``
586ioctls on the ``userfaultfd`` in order to map the received pages into the
587guest (``UFFDIO_ZEROCOPY`` is used if the source page was a zero page).
588
589A different postcopy thread in the destination node listens with
590poll() to the ``userfaultfd`` in parallel. When a ``POLLIN`` event is
591generated after a userfault triggers, the postcopy thread read() from
592the ``userfaultfd`` and receives the fault address (or ``-EAGAIN`` in case the
593userfault was already resolved and waken by a ``UFFDIO_COPY|ZEROPAGE`` run
594by the parallel QEMU migration thread).
595
596After the QEMU postcopy thread (running in the destination node) gets
597the userfault address it writes the information about the missing page
598into the socket. The QEMU source node receives the information and
599roughly "seeks" to that page address and continues sending all
600remaining missing pages from that new page offset. Soon after that
601(just the time to flush the tcp_wmem queue through the network) the
602migration thread in the QEMU running in the destination node will
603receive the page that triggered the userfault and it'll map it as
604usual with the ``UFFDIO_COPY|ZEROPAGE`` (without actually knowing if it
605was spontaneously sent by the source or if it was an urgent page
606requested through a userfault).
607
608By the time the userfaults start, the QEMU in the destination node
609doesn't need to keep any per-page state bitmap relative to the live
610migration around and a single per-page bitmap has to be maintained in
611the QEMU running in the source node to know which pages are still
612missing in the destination node. The bitmap in the source node is
613checked to find which missing pages to send in round robin and we seek
614over it when receiving incoming userfaults. After sending each page of
615course the bitmap is updated accordingly. It's also useful to avoid
616sending the same page twice (in case the userfault is read by the
617postcopy thread just before ``UFFDIO_COPY|ZEROPAGE`` runs in the migration
618thread).
619
620Non-cooperative userfaultfd
621===========================
622
623When the ``userfaultfd`` is monitored by an external manager, the manager
624must be able to track changes in the process virtual memory
625layout. Userfaultfd can notify the manager about such changes using
626the same read(2) protocol as for the page fault notifications. The
627manager has to explicitly enable these events by setting appropriate
628bits in ``uffdio_api.features`` passed to ``UFFDIO_API`` ioctl:
629
630``UFFD_FEATURE_EVENT_FORK``
631	enable ``userfaultfd`` hooks for fork(). When this feature is
632	enabled, the ``userfaultfd`` context of the parent process is
633	duplicated into the newly created process. The manager
634	receives ``UFFD_EVENT_FORK`` with file descriptor of the new
635	``userfaultfd`` context in the ``uffd_msg.fork``.
636
637``UFFD_FEATURE_EVENT_REMAP``
638	enable notifications about mremap() calls. When the
639	non-cooperative process moves a virtual memory area to a
640	different location, the manager will receive
641	``UFFD_EVENT_REMAP``. The ``uffd_msg.remap`` will contain the old and
642	new addresses of the area and its original length.
643
644``UFFD_FEATURE_EVENT_REMOVE``
645	enable notifications about madvise(MADV_REMOVE) and
646	madvise(MADV_DONTNEED) calls. The event ``UFFD_EVENT_REMOVE`` will
647	be generated upon these calls to madvise(). The ``uffd_msg.remove``
648	will contain start and end addresses of the removed area.
649
650``UFFD_FEATURE_EVENT_UNMAP``
651	enable notifications about memory unmapping. The manager will
652	get ``UFFD_EVENT_UNMAP`` with ``uffd_msg.remove`` containing start and
653	end addresses of the unmapped area.
654
655Although the ``UFFD_FEATURE_EVENT_REMOVE`` and ``UFFD_FEATURE_EVENT_UNMAP``
656are pretty similar, they quite differ in the action expected from the
657``userfaultfd`` manager. In the former case, the virtual memory is
658removed, but the area is not, the area remains monitored by the
659``userfaultfd``, and if a page fault occurs in that area it will be
660delivered to the manager. The proper resolution for such page fault is
661to zeromap the faulting address. However, in the latter case, when an
662area is unmapped, either explicitly (with munmap() system call), or
663implicitly (e.g. during mremap()), the area is removed and in turn the
664``userfaultfd`` context for such area disappears too and the manager will
665not get further userland page faults from the removed area. Still, the
666notification is required in order to prevent manager from using
667``UFFDIO_COPY`` on the unmapped area.
668
669Unlike userland page faults which have to be synchronous and require
670explicit or implicit wakeup, all the events are delivered
671asynchronously and the non-cooperative process resumes execution as
672soon as manager executes read(). The ``userfaultfd`` manager should
673carefully synchronize calls to ``UFFDIO_COPY`` with the events
674processing. To aid the synchronization, the ``UFFDIO_COPY`` ioctl will
675return ``-ENOSPC`` when the monitored process exits at the time of
676``UFFDIO_COPY``, and ``-ENOENT``, when the non-cooperative process has changed
677its virtual memory layout simultaneously with outstanding ``UFFDIO_COPY``
678operation.
679
680The current asynchronous model of the event delivery is optimal for
681single threaded non-cooperative ``userfaultfd`` manager implementations. A
682synchronous event delivery model can be added later as a new
683``userfaultfd`` feature to facilitate multithreading enhancements of the
684non cooperative manager, for example to allow ``UFFDIO_COPY`` ioctls to
685run in parallel to the event reception. Single threaded
686implementations should continue to use the current async event
687delivery model instead.
688