1=========== 2Userfaultfd 3=========== 4 5Objective 6========= 7 8Userfaults allow the implementation of on-demand paging from userland 9and more generally they allow userland to take control of various 10memory page faults, something otherwise only the kernel code could do. 11 12For example userfaults allows a proper and more optimal implementation 13of the ``PROT_NONE+SIGSEGV`` trick. 14 15Design 16====== 17 18Userspace creates a new userfaultfd, initializes it, and registers one or more 19regions of virtual memory with it. Then, any page faults which occur within the 20region(s) result in a message being delivered to the userfaultfd, notifying 21userspace of the fault. 22 23The ``userfaultfd`` (aside from registering and unregistering virtual 24memory ranges) provides two primary functionalities: 25 261) ``read/POLLIN`` protocol to notify a userland thread of the faults 27 happening 28 292) various ``UFFDIO_*`` ioctls that can manage the virtual memory regions 30 registered in the ``userfaultfd`` that allows userland to efficiently 31 resolve the userfaults it receives via 1) or to manage the virtual 32 memory in the background 33 34The real advantage of userfaults if compared to regular virtual memory 35management of mremap/mprotect is that the userfaults in all their 36operations never involve heavyweight structures like vmas (in fact the 37``userfaultfd`` runtime load never takes the mmap_lock for writing). 38Vmas are not suitable for page- (or hugepage) granular fault tracking 39when dealing with virtual address spaces that could span 40Terabytes. Too many vmas would be needed for that. 41 42The ``userfaultfd``, once created, can also be 43passed using unix domain sockets to a manager process, so the same 44manager process could handle the userfaults of a multitude of 45different processes without them being aware about what is going on 46(well of course unless they later try to use the ``userfaultfd`` 47themselves on the same region the manager is already tracking, which 48is a corner case that would currently return ``-EBUSY``). 49 50API 51=== 52 53Creating a userfaultfd 54---------------------- 55 56There are two ways to create a new userfaultfd, each of which provide ways to 57restrict access to this functionality (since historically userfaultfds which 58handle kernel page faults have been a useful tool for exploiting the kernel). 59 60The first way, supported since userfaultfd was introduced, is the 61userfaultfd(2) syscall. Access to this is controlled in several ways: 62 63- Any user can always create a userfaultfd which traps userspace page faults 64 only. Such a userfaultfd can be created using the userfaultfd(2) syscall 65 with the flag UFFD_USER_MODE_ONLY. 66 67- In order to also trap kernel page faults for the address space, either the 68 process needs the CAP_SYS_PTRACE capability, or the system must have 69 vm.unprivileged_userfaultfd set to 1. By default, vm.unprivileged_userfaultfd 70 is set to 0. 71 72The second way, added to the kernel more recently, is by opening 73/dev/userfaultfd and issuing a USERFAULTFD_IOC_NEW ioctl to it. This method 74yields equivalent userfaultfds to the userfaultfd(2) syscall. 75 76Unlike userfaultfd(2), access to /dev/userfaultfd is controlled via normal 77filesystem permissions (user/group/mode), which gives fine grained access to 78userfaultfd specifically, without also granting other unrelated privileges at 79the same time (as e.g. granting CAP_SYS_PTRACE would do). Users who have access 80to /dev/userfaultfd can always create userfaultfds that trap kernel page faults; 81vm.unprivileged_userfaultfd is not considered. 82 83Initializing a userfaultfd 84-------------------------- 85 86When first opened the ``userfaultfd`` must be enabled invoking the 87``UFFDIO_API`` ioctl specifying a ``uffdio_api.api`` value set to ``UFFD_API`` (or 88a later API version) which will specify the ``read/POLLIN`` protocol 89userland intends to speak on the ``UFFD`` and the ``uffdio_api.features`` 90userland requires. The ``UFFDIO_API`` ioctl if successful (i.e. if the 91requested ``uffdio_api.api`` is spoken also by the running kernel and the 92requested features are going to be enabled) will return into 93``uffdio_api.features`` and ``uffdio_api.ioctls`` two 64bit bitmasks of 94respectively all the available features of the read(2) protocol and 95the generic ioctl available. 96 97The ``uffdio_api.features`` bitmask returned by the ``UFFDIO_API`` ioctl 98defines what memory types are supported by the ``userfaultfd`` and what 99events, except page fault notifications, may be generated: 100 101- The ``UFFD_FEATURE_EVENT_*`` flags indicate that various other events 102 other than page faults are supported. These events are described in more 103 detail below in the `Non-cooperative userfaultfd`_ section. 104 105- ``UFFD_FEATURE_MISSING_HUGETLBFS`` and ``UFFD_FEATURE_MISSING_SHMEM`` 106 indicate that the kernel supports ``UFFDIO_REGISTER_MODE_MISSING`` 107 registrations for hugetlbfs and shared memory (covering all shmem APIs, 108 i.e. tmpfs, ``IPCSHM``, ``/dev/zero``, ``MAP_SHARED``, ``memfd_create``, 109 etc) virtual memory areas, respectively. 110 111- ``UFFD_FEATURE_MINOR_HUGETLBFS`` indicates that the kernel supports 112 ``UFFDIO_REGISTER_MODE_MINOR`` registration for hugetlbfs virtual memory 113 areas. ``UFFD_FEATURE_MINOR_SHMEM`` is the analogous feature indicating 114 support for shmem virtual memory areas. 115 116- ``UFFD_FEATURE_MOVE`` indicates that the kernel supports moving an 117 existing page contents from userspace. 118 119The userland application should set the feature flags it intends to use 120when invoking the ``UFFDIO_API`` ioctl, to request that those features be 121enabled if supported. 122 123Once the ``userfaultfd`` API has been enabled the ``UFFDIO_REGISTER`` 124ioctl should be invoked (if present in the returned ``uffdio_api.ioctls`` 125bitmask) to register a memory range in the ``userfaultfd`` by setting the 126uffdio_register structure accordingly. The ``uffdio_register.mode`` 127bitmask will specify to the kernel which kind of faults to track for 128the range. The ``UFFDIO_REGISTER`` ioctl will return the 129``uffdio_register.ioctls`` bitmask of ioctls that are suitable to resolve 130userfaults on the range registered. Not all ioctls will necessarily be 131supported for all memory types (e.g. anonymous memory vs. shmem vs. 132hugetlbfs), or all types of intercepted faults. 133 134.. note:: 135 136 Re-registering an already-registered range must not drop any of the 137 modes that install per-PTE markers — currently 138 ``UFFDIO_REGISTER_MODE_WP`` and ``UFFDIO_REGISTER_MODE_RWP``. Doing 139 so would strand markers with no flag to describe them, so the call 140 is rejected with ``-EBUSY``; userspace must issue 141 ``UFFDIO_UNREGISTER`` first. This differs from older kernels, which 142 silently replaced the mode bits on re-registration. 143 144Userland can use the ``uffdio_register.ioctls`` to manage the virtual 145address space in the background (to add or potentially also remove 146memory from the ``userfaultfd`` registered range). This means a userfault 147could be triggering just before userland maps in the background the 148user-faulted page. 149 150Resolving Userfaults 151-------------------- 152 153There are three basic ways to resolve userfaults: 154 155- ``UFFDIO_COPY`` atomically copies some existing page contents from 156 userspace. 157 158- ``UFFDIO_ZEROPAGE`` atomically zeros the new page. 159 160- ``UFFDIO_CONTINUE`` maps an existing, previously-populated page. 161 162These operations are atomic in the sense that they guarantee nothing can 163see a half-populated page, since readers will keep userfaulting until the 164operation has finished. 165 166By default, these wake up userfaults blocked on the range in question. 167They support a ``UFFDIO_*_MODE_DONTWAKE`` ``mode`` flag, which indicates 168that waking will be done separately at some later time. 169 170Which ioctl to choose depends on the kind of page fault, and what we'd 171like to do to resolve it: 172 173- For ``UFFDIO_REGISTER_MODE_MISSING`` faults, the fault needs to be 174 resolved by either providing a new page (``UFFDIO_COPY``), or mapping 175 the zero page (``UFFDIO_ZEROPAGE``). By default, the kernel would map 176 the zero page for a missing fault. With userfaultfd, userspace can 177 decide what content to provide before the faulting thread continues. 178 179- For ``UFFDIO_REGISTER_MODE_MINOR`` faults, there is an existing page (in 180 the page cache). Userspace has the option of modifying the page's 181 contents before resolving the fault. Once the contents are correct 182 (modified or not), userspace asks the kernel to map the page and let the 183 faulting thread continue with ``UFFDIO_CONTINUE``. 184 185Notes: 186 187- You can tell which kind of fault occurred by examining 188 ``pagefault.flags`` within the ``uffd_msg``, checking for the 189 ``UFFD_PAGEFAULT_FLAG_*`` flags. 190 191- None of the page-delivering ioctls default to the range that you 192 registered with. You must fill in all fields for the appropriate 193 ioctl struct including the range. 194 195- You get the address of the access that triggered the missing page 196 event out of a struct uffd_msg that you read in the thread from the 197 uffd. You can supply as many pages as you want with these IOCTLs. 198 Keep in mind that unless you used DONTWAKE then the first of any of 199 those IOCTLs wakes up the faulting thread. 200 201- Be sure to test for all errors including 202 (``pollfd[0].revents & POLLERR``). This can happen, e.g. when ranges 203 supplied were incorrect. 204 205Write Protect Notifications 206--------------------------- 207 208This is equivalent to (but faster than) using mprotect and a SIGSEGV 209signal handler. 210 211Firstly you need to register a range with ``UFFDIO_REGISTER_MODE_WP``. 212Instead of using mprotect(2) you use 213``ioctl(uffd, UFFDIO_WRITEPROTECT, struct *uffdio_writeprotect)`` 214while ``mode = UFFDIO_WRITEPROTECT_MODE_WP`` 215in the struct passed in. The range does not default to and does not 216have to be identical to the range you registered with. You can write 217protect as many ranges as you like (inside the registered range). 218Then, in the thread reading from uffd the struct will have 219``msg.arg.pagefault.flags & UFFD_PAGEFAULT_FLAG_WP`` set. Now you send 220``ioctl(uffd, UFFDIO_WRITEPROTECT, struct *uffdio_writeprotect)`` 221again while ``pagefault.mode`` does not have ``UFFDIO_WRITEPROTECT_MODE_WP`` 222set. This wakes up the thread which will continue to run with writes. This 223allows you to do the bookkeeping about the write in the uffd reading 224thread before the ioctl. 225 226If you registered with both ``UFFDIO_REGISTER_MODE_MISSING`` and 227``UFFDIO_REGISTER_MODE_WP`` then you need to think about the sequence in 228which you supply a page and undo write protect. Note that there is a 229difference between writes into a WP area and into a !WP area. The 230former will have ``UFFD_PAGEFAULT_FLAG_WP`` set, the latter 231``UFFD_PAGEFAULT_FLAG_WRITE``. The latter did not fail on protection but 232you still need to supply a page when ``UFFDIO_REGISTER_MODE_MISSING`` was 233used. 234 235Userfaultfd write-protect mode currently behave differently on none ptes 236(when e.g. page is missing) over different types of memories. 237 238For anonymous memory, ``ioctl(UFFDIO_WRITEPROTECT)`` will ignore none ptes 239(e.g. when pages are missing and not populated). For file-backed memories 240like shmem and hugetlbfs, none ptes will be write protected just like a 241present pte. In other words, there will be a userfaultfd write fault 242message generated when writing to a missing page on file typed memories, 243as long as the page range was write-protected before. Such a message will 244not be generated on anonymous memories by default. 245 246If the application wants to be able to write protect none ptes on anonymous 247memory, one can pre-populate the memory with e.g. MADV_POPULATE_READ. On 248newer kernels, one can also detect the feature UFFD_FEATURE_WP_UNPOPULATED 249and set the feature bit in advance to make sure none ptes will also be 250write protected even upon anonymous memory. 251 252When using ``UFFDIO_REGISTER_MODE_WP`` in combination with either 253``UFFDIO_REGISTER_MODE_MISSING`` or ``UFFDIO_REGISTER_MODE_MINOR``, when 254resolving missing / minor faults with ``UFFDIO_COPY`` or ``UFFDIO_CONTINUE`` 255respectively, it may be desirable for the new page / mapping to be 256write-protected (so future writes will also result in a WP fault). These ioctls 257support a mode flag (``UFFDIO_COPY_MODE_WP`` or ``UFFDIO_CONTINUE_MODE_WP`` 258respectively) to configure the mapping this way. 259 260If the userfaultfd context has ``UFFD_FEATURE_WP_ASYNC`` feature bit set, 261any vma registered with write-protection will work in async mode rather 262than the default sync mode. 263 264In async mode, there will be no message generated when a write operation 265happens, meanwhile the write-protection will be resolved automatically by 266the kernel. It can be seen as a more accurate version of soft-dirty 267tracking and it can be different in a few ways: 268 269 - The dirty result will not be affected by vma changes (e.g. vma 270 merging) because the dirty is only tracked by the pte. 271 272 - It supports range operations by default, so one can enable tracking on 273 any range of memory as long as page aligned. 274 275 - Dirty information will not get lost if the pte was zapped due to 276 various reasons (e.g. during split of a shmem transparent huge page). 277 278 - Due to a reverted meaning of soft-dirty (page clean when the uffd bit 279 is set; dirty when the uffd bit is cleared), it has different semantics 280 on some of the memory operations. For example: ``MADV_DONTNEED`` on 281 anonymous (or ``MADV_REMOVE`` on a file mapping) will be treated as 282 dirtying of memory by dropping the uffd bit during the procedure. 283 284The user app can collect the "written/dirty" status by looking up the 285uffd bit for the pages being interested in /proc/pagemap. 286 287The page will not be under track of userfaultfd-wp async mode until the page is 288explicitly write-protected by ``ioctl(UFFDIO_WRITEPROTECT)`` with the mode 289flag ``UFFDIO_WRITEPROTECT_MODE_WP`` set. Trying to resolve a page fault 290that was tracked by async mode userfaultfd-wp is invalid. 291 292When userfaultfd-wp async mode is used alone, it can be applied to all 293kinds of memory. 294 295Memory Poisioning Emulation 296--------------------------- 297 298In response to a fault (either missing or minor), an action userspace can 299take to "resolve" it is to issue a ``UFFDIO_POISON``. This will cause any 300future faulters to either get a SIGBUS, or in KVM's case the guest will 301receive an MCE as if there were hardware memory poisoning. 302 303This is used to emulate hardware memory poisoning. Imagine a VM running on a 304machine which experiences a real hardware memory error. Later, we live migrate 305the VM to another physical machine. Since we want the migration to be 306transparent to the guest, we want that same address range to act as if it was 307still poisoned, even though it's on a new physical host which ostensibly 308doesn't have a memory error in the exact same spot. 309 310Read-Write Protection 311--------------------- 312 313``UFFDIO_REGISTER_MODE_RWP`` enables read-write protection tracking on a 314memory range. It is similar to (but faster than) ``mprotect(PROT_NONE)`` 315combined with a signal handler; unlike ``mprotect(PROT_NONE)``, RWP only 316traps accesses to *present* PTEs, so accesses to unpopulated addresses in a 317protected range fall through to the normal missing-page path. It uses the 318PROT_NONE hinting mechanism (same as NUMA balancing) to make pages 319inaccessible while keeping them resident in memory. Works on anonymous, 320shmem, and hugetlbfs memory. 321 322RWP is designed for VM memory managers that need to track the working set 323of guest memory for cold page eviction to tiered or remote storage. 324 325**Setup:** 326 3271. Open a userfaultfd and enable ``UFFD_FEATURE_RWP`` via ``UFFDIO_API``. 328 Optionally request ``UFFD_FEATURE_RWP_ASYNC`` as well — it requires 329 ``UFFD_FEATURE_RWP`` to be set in the same ``UFFDIO_API`` call. 330 3312. Register the guest memory range with ``UFFDIO_REGISTER_MODE_RWP`` 332 (and ``UFFDIO_REGISTER_MODE_MISSING`` if evicted pages will need to be 333 fetched back from storage). 334 335**Feature availability:** 336 337RWP is built on top of two kernel primitives: a spare PTE bit owned by 338userfaultfd (``CONFIG_HAVE_ARCH_USERFAULTFD_WP``) and architecture support 339for present-but-inaccessible PTEs (``CONFIG_ARCH_HAS_PTE_PROTNONE``). When both 340are available on a 64-bit kernel, the build selects 341``CONFIG_USERFAULTFD_RWP=y`` and the ``VM_UFFD_RWP`` VMA flag becomes 342available. 343 344``UFFD_FEATURE_RWP`` and ``UFFD_FEATURE_RWP_ASYNC`` are unavailable when 345the running kernel or architecture does not support them — for example 34632-bit kernels (where ``VM_UFFD_RWP`` is unavailable), kernels built 347without ``CONFIG_USERFAULTFD_RWP``, and architectures whose ptes cannot 348carry the uffd bit at runtime (e.g. riscv without the ``SVRSW60T59B`` 349extension). Requesting an unsupported feature in 350``uffdio_api.features`` makes ``UFFDIO_API`` fail with ``EINVAL`` and 351leaves the userfaultfd context uninitialized; the structure is returned 352zeroed, so the error path cannot be used to discover what the kernel 353supports. The recommended probe sequence is therefore to open a 354throwaway userfaultfd, call ``UFFDIO_API`` once with ``features = 0``, 355inspect the returned bitmask, close that fd, then open the real one 356and call ``UFFDIO_API`` again with only the supported features set. 357 358**Protecting and Unprotecting:** 359 360Use ``UFFDIO_RWPROTECT`` to protect or unprotect a range, mirroring the 361``UFFDIO_WRITEPROTECT`` interface:: 362 363 struct uffdio_rwprotect rwp = { 364 .range = { .start = addr, .len = len }, 365 .mode = UFFDIO_RWPROTECT_MODE_RWP, /* protect */ 366 }; 367 ioctl(uffd, UFFDIO_RWPROTECT, &rwp); 368 369Setting ``UFFDIO_RWPROTECT_MODE_RWP`` sets PROT_NONE on present PTEs in the 370range. Pages stay resident and their physical frames are preserved — only 371access permissions are removed. 372 373Clearing ``UFFDIO_RWPROTECT_MODE_RWP`` restores normal VMA permissions and 374wakes any faulting threads (unless ``UFFDIO_RWPROTECT_MODE_DONTWAKE`` is set). 375 376**Scope of protection:** 377 378RWP protection is a property of *present* PTEs. ``UFFDIO_RWPROTECT`` only 379affects entries that are already populated. Unpopulated addresses within 380the range remain unpopulated; when first accessed they fault through the 381normal missing path (``do_anonymous_page()``, ``do_swap_page()``, 382``finish_fault()``) and the resulting PTE is not RWP-protected. To observe 383the population itself, co-register the range with 384``UFFDIO_REGISTER_MODE_MISSING``. 385 386Protection is preserved across page reclaim: a page swapped out while 387RWP-protected carries the marker on its swap entry, and swap-in restores 388the PROT_NONE state so the first access after swap-in still faults. The 389same applies to pages temporarily replaced by migration entries. 390 391Operations that drop the PTE entirely — ``MADV_DONTNEED`` on anonymous 392memory, hole-punch on shmem, truncation of a file mapping — also drop the 393RWP marker: the next access re-populates the range without protection. 394Unlike WP (which persists via ``PTE_MARKER_UFFD_WP``), there is no 395persistent RWP marker today. The user needs to re-arm the range with 396``UFFDIO_RWPROTECT`` after any operation that explicitly frees PTEs. 397 398**Fault Handling:** 399 400When a protected page is accessed: 401 402- **Sync mode** (default): The faulting thread blocks and a 403 ``UFFD_PAGEFAULT_FLAG_RWP`` message is delivered to the userfaultfd 404 handler. The handler resolves the fault with ``UFFDIO_RWPROTECT`` 405 (clearing ``MODE_RWP``), which restores the PTE permissions and wakes 406 the faulting thread. 407 408- **Async mode** (``UFFD_FEATURE_RWP_ASYNC``): The kernel automatically 409 restores PTE permissions and the thread continues without blocking. No 410 message is delivered to the handler. 411 412**Runtime Mode Switching:** 413 414``UFFDIO_SET_MODE`` toggles ``UFFD_FEATURE_RWP_ASYNC`` at runtime, allowing 415the VMM to switch between lightweight async detection and safe sync 416eviction without re-registering. The toggle takes ``mmap_write_lock()`` 417and calls ``vma_start_write()`` on each UFFD-armed VMA, draining 418in-flight per-VMA-locked faults before the new mode takes effect. 419 420**Working-set detection with PAGEMAP_SCAN:** 421 422RWP-protected PTEs carry the uffd PTE bit; an access (and, in async mode, its 423auto-resolution) clears it. ``PAGEMAP_SCAN`` reports ``PAGE_IS_ACCESSED`` once 424the bit is clear on a ``VM_UFFD_RWP`` VMA, so a *non-inverted* scan reports the 425pages that were touched during the interval -- the hot set:: 426 427 struct pm_scan_arg arg = { 428 .size = sizeof(arg), 429 .start = guest_mem_start, 430 .end = guest_mem_end, 431 .vec = (uint64_t)regions, 432 .vec_len = regions_len, 433 .category_mask = PAGE_IS_ACCESSED, 434 .return_mask = PAGE_IS_ACCESSED, 435 }; 436 long n = ioctl(pagemap_fd, PAGEMAP_SCAN, &arg); 437 438The returned ``page_region`` array lists the hot ranges. ``PAGE_IS_ACCESSED`` 439is set on an accessed page whether it is still present or has since been 440swapped out, so the hot scan needs no ``PAGE_IS_PRESENT`` filter -- unpopulated 441holes carry neither bit and are excluded on their own. 442 443Track the hot set and reclaim everything else from the backing file (see the 444workflow below). Do **not** invert the scan to enumerate "cold" pages 445directly: an inverted scan reports only the ``VM_UFFD_RWP`` PTEs that are still 446protected, i.e. the resident portion of *this* VMA. For a file mapping the 447working set spans the whole file -- pages that live in the page cache but are 448not mapped into this VMA (a pre-populated tmpfs file, or memory populated 449through another mapping) are ``pte_none`` here, never appear in the scan, and 450would never be considered for eviction even though they occupy memory. Driving 451eviction from "file offsets minus the hot set" avoids that blind spot; a cold 452PTE scan cannot. To additionally record the *first* access to a cached but 453unmapped page (e.g. pre-populated content) as hot, co-register the range with 454``UFFDIO_REGISTER_MODE_MINOR``: such accesses then fault as minor faults 455instead of mapping the page silently. 456 457**Cleanup:** 458 459When the userfaultfd is closed or the range is unregistered, all PROT_NONE 460PTEs are automatically restored to their normal VMA permissions. This 461prevents pages from becoming permanently inaccessible. 462 463**VMM Working Set Tracking Workflow:** 464 465A typical VMM lifecycle for cold page eviction to tiered storage. Two 466mappings of the same shmem (or hugetlbfs) file are used: ``guest_mem`` is 467the RWP-registered mapping that vCPUs access through, and ``io_mem`` is a 468private mapping for VMM-side I/O. Reading ``io_mem`` does not go through 469the RWP-protected PTEs of ``guest_mem``, so the VMM's own ``pwrite()`` 470never traps on its own :: 471 472 /* One-time setup */ 473 fd = memfd_create("guest", MFD_CLOEXEC); 474 ftruncate(fd, guest_size); 475 guest_mem = mmap(NULL, guest_size, PROT_READ | PROT_WRITE, 476 MAP_SHARED, fd, 0); /* vCPU view, RWP-registered */ 477 io_mem = mmap(NULL, guest_size, PROT_READ | PROT_WRITE, 478 MAP_SHARED, fd, 0); /* VMM I/O view, unprotected */ 479 480 uffd = userfaultfd(O_CLOEXEC | O_NONBLOCK); 481 struct uffdio_api api = { 482 .api = UFFD_API, 483 .features = UFFD_FEATURE_RWP | UFFD_FEATURE_RWP_ASYNC, 484 }; 485 ioctl(uffd, UFFDIO_API, &api); 486 if (!(api.features & UFFD_FEATURE_RWP)) 487 /* RWP unavailable on this kernel/arch -- fall back. */ 488 ioctl(uffd, UFFDIO_REGISTER, &(struct uffdio_register){ 489 .range = { guest_mem, guest_size }, 490 .mode = UFFDIO_REGISTER_MODE_RWP | 491 UFFDIO_REGISTER_MODE_MISSING, 492 }); 493 494 /* Tracking loop */ 495 while (vm_running) { 496 /* 1. Detection phase (async -- no vCPU stalls) */ 497 ioctl(uffd, UFFDIO_RWPROTECT, &(struct uffdio_rwprotect){ 498 .range = full_range, 499 .mode = UFFDIO_RWPROTECT_MODE_RWP }); 500 sleep(tracking_interval); 501 502 /* 503 * 2. Switch to sync BEFORE scanning. In async mode a vCPU 504 * access races eviction: it would auto-resolve and mark the 505 * page hot just as the VMM writes it out and punches it, 506 * losing the update. Sync mode makes such accesses block and 507 * be delivered, freezing the hot snapshot for the rest of the 508 * iteration. 509 */ 510 ioctl(uffd, UFFDIO_SET_MODE, 511 &(struct uffdio_set_mode){ 512 .disable = UFFD_FEATURE_RWP_ASYNC }); 513 514 /* 3. Read the hot set: pages touched this interval. */ 515 ioctl(pagemap_fd, PAGEMAP_SCAN, &(struct pm_scan_arg){ 516 .category_mask = PAGE_IS_ACCESSED, 517 .return_mask = PAGE_IS_ACCESSED, 518 ... 519 }); 520 521 /* 522 * 4. Reclaim the file offsets that are NOT in the hot set. 523 * Driving this from the file's offset space (rather than from a 524 * cold PTE scan) also reclaims pages that are cached but not 525 * mapped into guest_mem, e.g. pre-populated content. 526 */ 527 for each non-hot offset range: 528 /* Read from io_mem -- bypasses RWP, no fault. */ 529 pwrite(storage_fd, (char *)io_mem + off, len, off); 530 /* Drop the page from the shared file. */ 531 fallocate(fd, FALLOC_FL_PUNCH_HOLE | FALLOC_FL_KEEP_SIZE, 532 off, len); 533 /* 534 * Wake any vCPU blocked on the RWP fault for this range: 535 * fallocate() does not iterate ctx->fault_pending_wqh. 536 */ 537 ioctl(uffd, UFFDIO_WAKE, &(struct uffdio_range){ 538 .start = (uintptr_t)guest_mem + off, .len = len }); 539 540 /* 5. Resume async tracking */ 541 ioctl(uffd, UFFDIO_SET_MODE, 542 &(struct uffdio_set_mode){ 543 .enable = UFFD_FEATURE_RWP_ASYNC }); 544 } 545 546During step 4, a vCPU that accesses a ``guest_mem`` offset being evicted 547blocks with a ``UFFD_PAGEFAULT_FLAG_RWP`` fault while the eviction is in 548progress. After ``fallocate()`` punches the page out and ``UFFDIO_WAKE`` 549fires, the vCPU retries the access, faults as ``MISSING``, and the 550handler resolves it with ``UFFDIO_COPY`` from storage. 551 552This workflow targets shmem and hugetlbfs (both support a private 553``io_mem`` mapping over the same fd). Anonymous-memory backings need a 554different inner-loop strategy because the VMM has no way to read the 555page without going through the RWP-protected mapping. 556 557QEMU/KVM 558======== 559 560QEMU/KVM is using the ``userfaultfd`` syscall to implement postcopy live 561migration. Postcopy live migration is one form of memory 562externalization consisting of a virtual machine running with part or 563all of its memory residing on a different node in the cloud. The 564``userfaultfd`` abstraction is generic enough that not a single line of 565KVM kernel code had to be modified in order to add postcopy live 566migration to QEMU. 567 568Guest async page faults, ``FOLL_NOWAIT`` and all other ``GUP*`` features work 569just fine in combination with userfaults. Userfaults trigger async 570page faults in the guest scheduler so those guest processes that 571aren't waiting for userfaults (i.e. network bound) can keep running in 572the guest vcpus. 573 574It is generally beneficial to run one pass of precopy live migration 575just before starting postcopy live migration, in order to avoid 576generating userfaults for readonly guest regions. 577 578The implementation of postcopy live migration currently uses one 579single bidirectional socket but in the future two different sockets 580will be used (to reduce the latency of the userfaults to the minimum 581possible without having to decrease ``/proc/sys/net/ipv4/tcp_wmem``). 582 583The QEMU in the source node writes all pages that it knows are missing 584in the destination node, into the socket, and the migration thread of 585the QEMU running in the destination node runs ``UFFDIO_COPY|ZEROPAGE`` 586ioctls on the ``userfaultfd`` in order to map the received pages into the 587guest (``UFFDIO_ZEROCOPY`` is used if the source page was a zero page). 588 589A different postcopy thread in the destination node listens with 590poll() to the ``userfaultfd`` in parallel. When a ``POLLIN`` event is 591generated after a userfault triggers, the postcopy thread read() from 592the ``userfaultfd`` and receives the fault address (or ``-EAGAIN`` in case the 593userfault was already resolved and waken by a ``UFFDIO_COPY|ZEROPAGE`` run 594by the parallel QEMU migration thread). 595 596After the QEMU postcopy thread (running in the destination node) gets 597the userfault address it writes the information about the missing page 598into the socket. The QEMU source node receives the information and 599roughly "seeks" to that page address and continues sending all 600remaining missing pages from that new page offset. Soon after that 601(just the time to flush the tcp_wmem queue through the network) the 602migration thread in the QEMU running in the destination node will 603receive the page that triggered the userfault and it'll map it as 604usual with the ``UFFDIO_COPY|ZEROPAGE`` (without actually knowing if it 605was spontaneously sent by the source or if it was an urgent page 606requested through a userfault). 607 608By the time the userfaults start, the QEMU in the destination node 609doesn't need to keep any per-page state bitmap relative to the live 610migration around and a single per-page bitmap has to be maintained in 611the QEMU running in the source node to know which pages are still 612missing in the destination node. The bitmap in the source node is 613checked to find which missing pages to send in round robin and we seek 614over it when receiving incoming userfaults. After sending each page of 615course the bitmap is updated accordingly. It's also useful to avoid 616sending the same page twice (in case the userfault is read by the 617postcopy thread just before ``UFFDIO_COPY|ZEROPAGE`` runs in the migration 618thread). 619 620Non-cooperative userfaultfd 621=========================== 622 623When the ``userfaultfd`` is monitored by an external manager, the manager 624must be able to track changes in the process virtual memory 625layout. Userfaultfd can notify the manager about such changes using 626the same read(2) protocol as for the page fault notifications. The 627manager has to explicitly enable these events by setting appropriate 628bits in ``uffdio_api.features`` passed to ``UFFDIO_API`` ioctl: 629 630``UFFD_FEATURE_EVENT_FORK`` 631 enable ``userfaultfd`` hooks for fork(). When this feature is 632 enabled, the ``userfaultfd`` context of the parent process is 633 duplicated into the newly created process. The manager 634 receives ``UFFD_EVENT_FORK`` with file descriptor of the new 635 ``userfaultfd`` context in the ``uffd_msg.fork``. 636 637``UFFD_FEATURE_EVENT_REMAP`` 638 enable notifications about mremap() calls. When the 639 non-cooperative process moves a virtual memory area to a 640 different location, the manager will receive 641 ``UFFD_EVENT_REMAP``. The ``uffd_msg.remap`` will contain the old and 642 new addresses of the area and its original length. 643 644``UFFD_FEATURE_EVENT_REMOVE`` 645 enable notifications about madvise(MADV_REMOVE) and 646 madvise(MADV_DONTNEED) calls. The event ``UFFD_EVENT_REMOVE`` will 647 be generated upon these calls to madvise(). The ``uffd_msg.remove`` 648 will contain start and end addresses of the removed area. 649 650``UFFD_FEATURE_EVENT_UNMAP`` 651 enable notifications about memory unmapping. The manager will 652 get ``UFFD_EVENT_UNMAP`` with ``uffd_msg.remove`` containing start and 653 end addresses of the unmapped area. 654 655Although the ``UFFD_FEATURE_EVENT_REMOVE`` and ``UFFD_FEATURE_EVENT_UNMAP`` 656are pretty similar, they quite differ in the action expected from the 657``userfaultfd`` manager. In the former case, the virtual memory is 658removed, but the area is not, the area remains monitored by the 659``userfaultfd``, and if a page fault occurs in that area it will be 660delivered to the manager. The proper resolution for such page fault is 661to zeromap the faulting address. However, in the latter case, when an 662area is unmapped, either explicitly (with munmap() system call), or 663implicitly (e.g. during mremap()), the area is removed and in turn the 664``userfaultfd`` context for such area disappears too and the manager will 665not get further userland page faults from the removed area. Still, the 666notification is required in order to prevent manager from using 667``UFFDIO_COPY`` on the unmapped area. 668 669Unlike userland page faults which have to be synchronous and require 670explicit or implicit wakeup, all the events are delivered 671asynchronously and the non-cooperative process resumes execution as 672soon as manager executes read(). The ``userfaultfd`` manager should 673carefully synchronize calls to ``UFFDIO_COPY`` with the events 674processing. To aid the synchronization, the ``UFFDIO_COPY`` ioctl will 675return ``-ENOSPC`` when the monitored process exits at the time of 676``UFFDIO_COPY``, and ``-ENOENT``, when the non-cooperative process has changed 677its virtual memory layout simultaneously with outstanding ``UFFDIO_COPY`` 678operation. 679 680The current asynchronous model of the event delivery is optimal for 681single threaded non-cooperative ``userfaultfd`` manager implementations. A 682synchronous event delivery model can be added later as a new 683``userfaultfd`` feature to facilitate multithreading enhancements of the 684non cooperative manager, for example to allow ``UFFDIO_COPY`` ioctls to 685run in parallel to the event reception. Single threaded 686implementations should continue to use the current async event 687delivery model instead. 688