1*3160c8deSAhmed S. Darwish.. SPDX-License-Identifier: GPL-2.0 2*3160c8deSAhmed S. Darwish 3*3160c8deSAhmed S. Darwish============================== 4*3160c8deSAhmed S. DarwishReal-Time Kernel configuration 5*3160c8deSAhmed S. Darwish============================== 6*3160c8deSAhmed S. Darwish 7*3160c8deSAhmed S. Darwish.. contents:: Table of Contents 8*3160c8deSAhmed S. Darwish :depth: 3 9*3160c8deSAhmed S. Darwish :local: 10*3160c8deSAhmed S. Darwish 11*3160c8deSAhmed S. DarwishIntroduction 12*3160c8deSAhmed S. Darwish============ 13*3160c8deSAhmed S. Darwish 14*3160c8deSAhmed S. DarwishThis document lists the kernel configuration options that might affect a 15*3160c8deSAhmed S. Darwishreal-time kernel's worst-case latency. It is intended for system integrators. 16*3160c8deSAhmed S. Darwish 17*3160c8deSAhmed S. DarwishConfiguration options 18*3160c8deSAhmed S. Darwish===================== 19*3160c8deSAhmed S. Darwish 20*3160c8deSAhmed S. Darwish.. Please keep the configuration listings alphabetically ordered 21*3160c8deSAhmed S. Darwish 22*3160c8deSAhmed S. DarwishCPU frequency governors 23*3160c8deSAhmed S. Darwish----------------------- 24*3160c8deSAhmed S. Darwish 25*3160c8deSAhmed S. Darwish``CONFIG_CPU_FREQ`` 26*3160c8deSAhmed S. Darwish^^^^^^^^^^^^^^^^^^^ 27*3160c8deSAhmed S. Darwish 28*3160c8deSAhmed S. Darwish:Expectation: enabled 29*3160c8deSAhmed S. Darwish:Severity: *high* 30*3160c8deSAhmed S. Darwish 31*3160c8deSAhmed S. DarwishThe CPU frequency scaling subsystem ensures that the processor can operate at 32*3160c8deSAhmed S. Darwishits maximum supported frequency. While, in general, bootloaders are tasked 33*3160c8deSAhmed S. Darwishwith setting the CPU clock to the highest speed on boot, some do not. It is 34*3160c8deSAhmed S. Darwishthus desirable to keep this option enabled. 35*3160c8deSAhmed S. Darwish 36*3160c8deSAhmed S. Darwish.. caution:: 37*3160c8deSAhmed S. Darwish 38*3160c8deSAhmed S. Darwish A real-time kernel is not about being "as fast as possible", however 39*3160c8deSAhmed S. Darwish real-time requirements may demand that the CPU is clocked at a particular 40*3160c8deSAhmed S. Darwish speed. 41*3160c8deSAhmed S. Darwish 42*3160c8deSAhmed S. Darwish``CONFIG_CPU_FREQ_DEFAULT_GOV_PERFORMANCE`` 43*3160c8deSAhmed S. Darwish^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ 44*3160c8deSAhmed S. Darwish 45*3160c8deSAhmed S. Darwish:Expectation: enabled 46*3160c8deSAhmed S. Darwish:Severity: *high* 47*3160c8deSAhmed S. Darwish 48*3160c8deSAhmed S. DarwishReal-Time workloads expect a fixed CPU frequency during execution. Using the 49*3160c8deSAhmed S. Darwishperformance governor is an easy way to achieve that purely from kernel 50*3160c8deSAhmed S. Darwishconfiguration. 51*3160c8deSAhmed S. Darwish 52*3160c8deSAhmed S. DarwishThis is not an absolute rule. Some setups might prefer to clock the CPU to 53*3160c8deSAhmed S. Darwishlower speeds due to thermal packaging or other requirements. The key is that 54*3160c8deSAhmed S. Darwishthe CPU frequency remains constant once set. 55*3160c8deSAhmed S. Darwish 56*3160c8deSAhmed S. DarwishNon-performance CPU frequency governors 57*3160c8deSAhmed S. Darwish^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ 58*3160c8deSAhmed S. Darwish 59*3160c8deSAhmed S. Darwish:Expectation: disabled 60*3160c8deSAhmed S. Darwish:Severity: *medium* 61*3160c8deSAhmed S. Darwish 62*3160c8deSAhmed S. DarwishTo ensure reproducible system latency measurements, disable the 63*3160c8deSAhmed S. Darwishnon-``PERFORMANCE`` CPU frequency governors whenever possible. This avoids 64*3160c8deSAhmed S. Darwishthe risk of unknown userspace tasks implicitly or explicitly setting a 65*3160c8deSAhmed S. Darwishdifferent CPU frequency governor, and thereby changing latency behavior while 66*3160c8deSAhmed S. Darwishthe system is running. 67*3160c8deSAhmed S. Darwish 68*3160c8deSAhmed S. DarwishIf disabling other frequency governors is not an option, use a governor that 69*3160c8deSAhmed S. Darwishkeeps the CPU frequency fixed. For example, 70*3160c8deSAhmed S. Darwish``CONFIG_CPU_FREQ_DEFAULT_GOV_USERSPACE`` can be enabled when userspace is 71*3160c8deSAhmed S. Darwishresponsible for setting a *stable* frequency during system initialization. 72*3160c8deSAhmed S. Darwish 73*3160c8deSAhmed S. DarwishIf a low CPU frequency is desired, then 74*3160c8deSAhmed S. Darwish``CONFIG_CPU_FREQ_DEFAULT_GOV_POWERSAVE`` can be set. 75*3160c8deSAhmed S. Darwish 76*3160c8deSAhmed S. DarwishThe ``ONDEMAND`` governor should not be enabled on a real-time system. Its 77*3160c8deSAhmed S. Darwishfrequency changes depend on workload behavior and can significantly harm 78*3160c8deSAhmed S. Darwishdeterminism. 79*3160c8deSAhmed S. Darwish 80*3160c8deSAhmed S. DarwishFor more information, see Documentation/admin-guide/pm/cpufreq.rst 81*3160c8deSAhmed S. Darwish 82*3160c8deSAhmed S. Darwish``CONFIG_CPU_IDLE`` 83*3160c8deSAhmed S. Darwish------------------- 84*3160c8deSAhmed S. Darwish 85*3160c8deSAhmed S. Darwish:Expectation: enabled 86*3160c8deSAhmed S. Darwish:Severity: *info* 87*3160c8deSAhmed S. Darwish 88*3160c8deSAhmed S. DarwishCPU idle states (C-states) allow the processor to enter low-power modes during 89*3160c8deSAhmed S. Darwishperiods of inactivity. Very-low CPU idle states may require flushing the CPU 90*3160c8deSAhmed S. Darwishcaches and lowering or disabling the clocking. This can lower power 91*3160c8deSAhmed S. Darwishconsumption, but it also increases the entry and exit latency from such 92*3160c8deSAhmed S. Darwishstates. 93*3160c8deSAhmed S. Darwish 94*3160c8deSAhmed S. DarwishWhile disabling this option eliminates cpuidle-related latencies, doing so can 95*3160c8deSAhmed S. Darwishsignificantly impact hardware longevity, warranty, and thermal behavior. 96*3160c8deSAhmed S. DarwishUsers should cap the maximum C-state to C1 instead. For ACPI platforms, this 97*3160c8deSAhmed S. Darwishcan be achieved by using the boot parameter [1]_:: 98*3160c8deSAhmed S. Darwish 99*3160c8deSAhmed S. Darwish processor.max_cstate=1 100*3160c8deSAhmed S. Darwish 101*3160c8deSAhmed S. DarwishHigher C-states can be acceptable depending on the user workload's latency 102*3160c8deSAhmed S. Darwishrequirements. For ACPI-based platforms, use the ``cpupower idle-info`` 103*3160c8deSAhmed S. Darwishcommand to inspect the available idle states. 104*3160c8deSAhmed S. Darwish 105*3160c8deSAhmed S. DarwishFor more information, please see: 106*3160c8deSAhmed S. Darwish 107*3160c8deSAhmed S. Darwish- ``linux/tools/power/cpupower`` 108*3160c8deSAhmed S. Darwish- Documentation/admin-guide/pm/cpuidle.rst 109*3160c8deSAhmed S. Darwish- Documentation/admin-guide/pm/index.rst 110*3160c8deSAhmed S. Darwish 111*3160c8deSAhmed S. Darwish``CONFIG_DRM`` 112*3160c8deSAhmed S. Darwish-------------- 113*3160c8deSAhmed S. Darwish 114*3160c8deSAhmed S. Darwish:Expectation: disabled 115*3160c8deSAhmed S. Darwish:Severity: *info* 116*3160c8deSAhmed S. Darwish 117*3160c8deSAhmed S. DarwishGPU-accelerated workloads can share system resources with the CPU, including 118*3160c8deSAhmed S. Darwishlast-level cache (LLC) and memory bandwidth. Modern integrated GPUs optimize 119*3160c8deSAhmed S. Darwishgraphics performance at the expense of CPU determinism. 120*3160c8deSAhmed S. Darwish 121*3160c8deSAhmed S. DarwishExamples of affected platforms: 122*3160c8deSAhmed S. Darwish 123*3160c8deSAhmed S. Darwish- Intel processors with integrated graphics (Gen9 and later) 124*3160c8deSAhmed S. Darwish- AMD APUs with Radeon Graphics 125*3160c8deSAhmed S. Darwish- Xilinx Zynq UltraScale+ MPSoC EG/EV series 126*3160c8deSAhmed S. Darwish 127*3160c8deSAhmed S. DarwishIf graphics workloads must run alongside real-time tasks, users must conduct 128*3160c8deSAhmed S. Darwishthorough stress testing using tools like ``glmark2`` while measuring the 129*3160c8deSAhmed S. Darwishoverall system latency. 130*3160c8deSAhmed S. Darwish 131*3160c8deSAhmed S. DarwishFor more information, please check: 132*3160c8deSAhmed S. Darwish 133*3160c8deSAhmed S. Darwish- Documentation/core-api/real-time/hardware.rst ("Regarding hardware" section) 134*3160c8deSAhmed S. Darwish- Documentation/filesystems/resctrl.rst 135*3160c8deSAhmed S. Darwish- `Real-Time and Graphics: A Contradiction? <https://web.archive.org/web/20221025085614/https://linutronix.de/PDF/Realtime_and_graphics-acontradiction2021.pdf>`_ 136*3160c8deSAhmed S. Darwish 137*3160c8deSAhmed S. Darwish``CONFIG_EFI_DISABLE_RUNTIME`` 138*3160c8deSAhmed S. Darwish------------------------------ 139*3160c8deSAhmed S. Darwish 140*3160c8deSAhmed S. Darwish:Expectation: enabled 141*3160c8deSAhmed S. Darwish:Severity: *medium* 142*3160c8deSAhmed S. Darwish 143*3160c8deSAhmed S. DarwishEFI is the standard boot and firmware interface for multiple architectures. 144*3160c8deSAhmed S. DarwishEFI runtime services provide callback functions to be called from the kernel; 145*3160c8deSAhmed S. Darwishe.g., as utilized by (``CONFIG_EFI_VARS*``) or (``CONFIG_RTC_DRV_EFI``). For 146*3160c8deSAhmed S. Darwishthe former, the kernel calls into EFI to update the EFI variables. 147*3160c8deSAhmed S. Darwish 148*3160c8deSAhmed S. DarwishCalling into EFI means invoking firmware callbacks. During such invocations, 149*3160c8deSAhmed S. Darwishthe system might not be able to react to interrupts and will thus not be able 150*3160c8deSAhmed S. Darwishto perform a context switch. This can cause significant latency spikes for 151*3160c8deSAhmed S. Darwishthe real-time system. 152*3160c8deSAhmed S. Darwish 153*3160c8deSAhmed S. Darwish``CONFIG_PREEMPT_RT`` enables this option by default. If this option is 154*3160c8deSAhmed S. Darwishmanually disabled at build time, the following boot parameter [1]_ may be used 155*3160c8deSAhmed S. Darwishto disable EFI runtime at boot up:: 156*3160c8deSAhmed S. Darwish 157*3160c8deSAhmed S. Darwish efi=noruntime 158*3160c8deSAhmed S. Darwish 159*3160c8deSAhmed S. DarwishAlternatively, confine EFI runtime service calls to a housekeeping CPU by 160*3160c8deSAhmed S. Darwishrestricting the ``efi_runtime`` workqueue CPU affinity. For example, set that 161*3160c8deSAhmed S. Darwishworkqueue's affinity to CPU #0 and pin your RT tasks to a different CPU range. 162*3160c8deSAhmed S. DarwishSee Documentation/core-api/workqueue.rst 163*3160c8deSAhmed S. Darwish 164*3160c8deSAhmed S. Darwish``CONFIG_NO_HZ`` / ``CONFIG_NO_HZ_FULL`` 165*3160c8deSAhmed S. Darwish---------------------------------------- 166*3160c8deSAhmed S. Darwish 167*3160c8deSAhmed S. Darwish:Expectation: disabled 168*3160c8deSAhmed S. Darwish:Severity: *medium* 169*3160c8deSAhmed S. Darwish 170*3160c8deSAhmed S. DarwishTickless operation can increase kernel-to-userspace transition latency due to 171*3160c8deSAhmed S. Darwishthe extra accounting and state book-keeping. 172*3160c8deSAhmed S. Darwish 173*3160c8deSAhmed S. Darwish*Guidance by real-time workload type:* 174*3160c8deSAhmed S. Darwish 175*3160c8deSAhmed S. Darwish- For periodic workloads; e.g., control loops executing every 100 µs, avoid 176*3160c8deSAhmed S. Darwish ``NO_HZ`` modes. Consistent kernel ticks are preferable. 177*3160c8deSAhmed S. Darwish 178*3160c8deSAhmed S. Darwish- For computation-intensive workloads; e.g. extended userspace execution, 179*3160c8deSAhmed S. Darwish ``NO_HZ_FULL`` may be beneficial. In such cases, users should offload the 180*3160c8deSAhmed S. Darwish kernel housekeeping to dedicated CPUs and isolate compute cores. 181*3160c8deSAhmed S. Darwish 182*3160c8deSAhmed S. DarwishSee also Documentation/timers/no_hz.rst 183*3160c8deSAhmed S. Darwish 184*3160c8deSAhmed S. Darwish``CONFIG_PREEMPT_RT`` 185*3160c8deSAhmed S. Darwish--------------------- 186*3160c8deSAhmed S. Darwish 187*3160c8deSAhmed S. Darwish:Expectation: enabled 188*3160c8deSAhmed S. Darwish:Severity: **fatal** 189*3160c8deSAhmed S. Darwish 190*3160c8deSAhmed S. DarwishThis option must be enabled, or the resulting kernel will not be fully 191*3160c8deSAhmed S. Darwishpreemptible and real-time capable. 192*3160c8deSAhmed S. Darwish 193*3160c8deSAhmed S. Darwish``CONFIG_TRACING`` (and tracing options) 194*3160c8deSAhmed S. Darwish---------------------------------------- 195*3160c8deSAhmed S. Darwish 196*3160c8deSAhmed S. Darwish:Expectation: enabled 197*3160c8deSAhmed S. Darwish:Severity: *info* 198*3160c8deSAhmed S. Darwish 199*3160c8deSAhmed S. DarwishShipping kernels with tracing support enabled (but not actively running) is 200*3160c8deSAhmed S. Darwishhighly recommended. This will allow the users to extract more information if 201*3160c8deSAhmed S. Darwishlatency problems arise. Nonetheless, some tracers do incur latency overhead 202*3160c8deSAhmed S. Darwishjust by being enabled. 203*3160c8deSAhmed S. Darwish 204*3160c8deSAhmed S. Darwish.. caution:: 205*3160c8deSAhmed S. Darwish 206*3160c8deSAhmed S. Darwish Users should *not* make use of tracers or trace events during production 207*3160c8deSAhmed S. Darwish real-time kernel operation as they can add considerable overhead and degrade 208*3160c8deSAhmed S. Darwish the system's latency. 209*3160c8deSAhmed S. Darwish 210*3160c8deSAhmed S. Darwish``CONFIG_IRQSOFF_TRACER`` and ``CONFIG_PREEMPT_TRACER`` 211*3160c8deSAhmed S. Darwish^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ 212*3160c8deSAhmed S. Darwish 213*3160c8deSAhmed S. Darwish:Expectation: disabled 214*3160c8deSAhmed S. Darwish:Severity: *high* 215*3160c8deSAhmed S. Darwish 216*3160c8deSAhmed S. DarwishThese tracers do incur measurable latency overhead even when tracing is not 217*3160c8deSAhmed S. Darwishcurrently active. 218*3160c8deSAhmed S. Darwish 219*3160c8deSAhmed S. DarwishKernel Debug Options 220*3160c8deSAhmed S. Darwish==================== 221*3160c8deSAhmed S. Darwish 222*3160c8deSAhmed S. DarwishMost kernel debug options add runtime overhead that increases the worst-case 223*3160c8deSAhmed S. Darwishlatency. 224*3160c8deSAhmed S. Darwish 225*3160c8deSAhmed S. Darwish.. caution:: 226*3160c8deSAhmed S. Darwish 227*3160c8deSAhmed S. Darwish During development and early testing, users are encouraged to run their 228*3160c8deSAhmed S. Darwish real-time workloads and peripherals with lockdep (:ref:`lockdep`) and other 229*3160c8deSAhmed S. Darwish kernel debug options enabled, for a considerable amount of time. Such 230*3160c8deSAhmed S. Darwish workloads might trigger kernel code paths that were not triggered during the 231*3160c8deSAhmed S. Darwish internal Linux real-time kernel development, thus helping to uncover locking 232*3160c8deSAhmed S. Darwish and other types of kernel bugs. 233*3160c8deSAhmed S. Darwish 234*3160c8deSAhmed S. Darwish``CONFIG_DEBUG_ATOMIC_SLEEP`` 235*3160c8deSAhmed S. Darwish----------------------------- 236*3160c8deSAhmed S. Darwish 237*3160c8deSAhmed S. Darwish:Expectation: allowed 238*3160c8deSAhmed S. Darwish 239*3160c8deSAhmed S. DarwishThis sanity check catches common kernel programming errors with a tolerable 240*3160c8deSAhmed S. Darwishlatency cost. It also increases overall scheduling as each ``might_sleep()`` 241*3160c8deSAhmed S. Darwishcan lead to a context switch. 242*3160c8deSAhmed S. Darwish 243*3160c8deSAhmed S. Darwish``CONFIG_DEBUG_BUGVERBOSE`` and ``CONFIG_DEBUG_INFO*`` 244*3160c8deSAhmed S. Darwish------------------------------------------------------ 245*3160c8deSAhmed S. Darwish 246*3160c8deSAhmed S. Darwish:Expectation: allowed 247*3160c8deSAhmed S. Darwish 248*3160c8deSAhmed S. DarwishThese options increase the kernel image size but have no latency impact. They 249*3160c8deSAhmed S. Darwishare also essential for meaningful BUG logs, crash dumps, and profiling. 250*3160c8deSAhmed S. Darwish 251*3160c8deSAhmed S. Darwish``CONFIG_DEBUG_FS`` 252*3160c8deSAhmed S. Darwish------------------- 253*3160c8deSAhmed S. Darwish 254*3160c8deSAhmed S. Darwish:Expectation: allowed 255*3160c8deSAhmed S. Darwish 256*3160c8deSAhmed S. DarwishThis is safe to include in real-time kernels, *provided that debugfs is not 257*3160c8deSAhmed S. Darwishaccessed during production runtime*. 258*3160c8deSAhmed S. Darwish 259*3160c8deSAhmed S. Darwish``CONFIG_DEBUG_KERNEL`` 260*3160c8deSAhmed S. Darwish----------------------- 261*3160c8deSAhmed S. Darwish 262*3160c8deSAhmed S. Darwish:Expectation: allowed 263*3160c8deSAhmed S. Darwish 264*3160c8deSAhmed S. DarwishMeta-option which allows debug features to be enabled. It has no runtime 265*3160c8deSAhmed S. Darwishimpact, but beware of any debug features that it may have implicitly enabled. 266*3160c8deSAhmed S. Darwish 267*3160c8deSAhmed S. Darwish``CONFIG_LOCKUP_DETECTOR`` 268*3160c8deSAhmed S. Darwish-------------------------- 269*3160c8deSAhmed S. Darwish 270*3160c8deSAhmed S. Darwish:Expectation: disabled 271*3160c8deSAhmed S. Darwish:Severity: *high* 272*3160c8deSAhmed S. Darwish 273*3160c8deSAhmed S. DarwishThe lockup detector creates kernel timer callbacks that execute every few 274*3160c8deSAhmed S. Darwishseconds, in hard-IRQ context, even on real-time kernels. These periodic 275*3160c8deSAhmed S. Darwishinterrupts can cause latency spikes. 276*3160c8deSAhmed S. Darwish 277*3160c8deSAhmed S. DarwishUsers should use hardware watchdogs instead, which will provide a similar 278*3160c8deSAhmed S. Darwishfunctionality without the software-induced latency. 279*3160c8deSAhmed S. Darwish 280*3160c8deSAhmed S. Darwish.. _lockdep: 281*3160c8deSAhmed S. Darwish 282*3160c8deSAhmed S. Darwish``CONFIG_PROVE_LOCKING`` 283*3160c8deSAhmed S. Darwish------------------------ 284*3160c8deSAhmed S. Darwish 285*3160c8deSAhmed S. Darwish:Expectation: disabled 286*3160c8deSAhmed S. Darwish:Severity: *high* 287*3160c8deSAhmed S. Darwish 288*3160c8deSAhmed S. DarwishProving the correctness of all kernel locking adds substantial overhead and 289*3160c8deSAhmed S. Darwishsignificantly increases worst-case latency. 290*3160c8deSAhmed S. Darwish 291*3160c8deSAhmed S. DarwishSummary 292*3160c8deSAhmed S. Darwish======= 293*3160c8deSAhmed S. Darwish 294*3160c8deSAhmed S. DarwishThere is no "one size fits all" solution for configuring a real-time Linux 295*3160c8deSAhmed S. Darwishsystem. Beginning with the system real-time requirements, integrators must 296*3160c8deSAhmed S. Darwishconsider the features and functions of the system's hardware, kernel, and 297*3160c8deSAhmed S. Darwishuserspace. All such components must be properly configured in order to 298*3160c8deSAhmed S. Darwishestablish and constrain the system's maximum latency. 299*3160c8deSAhmed S. Darwish 300*3160c8deSAhmed S. DarwishWith that in mind, any incorrect real-time kernel configuration could cause a 301*3160c8deSAhmed S. Darwishnew maximum latency that shows up at the wrong time and is catastrophic for 302*3160c8deSAhmed S. Darwishthe real-time system's latency. 303*3160c8deSAhmed S. Darwish 304*3160c8deSAhmed S. DarwishReferences 305*3160c8deSAhmed S. Darwish========== 306*3160c8deSAhmed S. Darwish 307*3160c8deSAhmed S. Darwish.. [1] See Documentation/admin-guide/kernel-parameters.rst 308