xref: /linux/Documentation/core-api/real-time/kernel-configuration.rst (revision 72fdff1416e280e2baaa3cca69574defb998437e)
1*3160c8deSAhmed S. Darwish.. SPDX-License-Identifier: GPL-2.0
2*3160c8deSAhmed S. Darwish
3*3160c8deSAhmed S. Darwish==============================
4*3160c8deSAhmed S. DarwishReal-Time Kernel configuration
5*3160c8deSAhmed S. Darwish==============================
6*3160c8deSAhmed S. Darwish
7*3160c8deSAhmed S. Darwish.. contents:: Table of Contents
8*3160c8deSAhmed S. Darwish   :depth: 3
9*3160c8deSAhmed S. Darwish   :local:
10*3160c8deSAhmed S. Darwish
11*3160c8deSAhmed S. DarwishIntroduction
12*3160c8deSAhmed S. Darwish============
13*3160c8deSAhmed S. Darwish
14*3160c8deSAhmed S. DarwishThis document lists the kernel configuration options that might affect a
15*3160c8deSAhmed S. Darwishreal-time kernel's worst-case latency.  It is intended for system integrators.
16*3160c8deSAhmed S. Darwish
17*3160c8deSAhmed S. DarwishConfiguration options
18*3160c8deSAhmed S. Darwish=====================
19*3160c8deSAhmed S. Darwish
20*3160c8deSAhmed S. Darwish.. Please keep the configuration listings alphabetically ordered
21*3160c8deSAhmed S. Darwish
22*3160c8deSAhmed S. DarwishCPU frequency governors
23*3160c8deSAhmed S. Darwish-----------------------
24*3160c8deSAhmed S. Darwish
25*3160c8deSAhmed S. Darwish``CONFIG_CPU_FREQ``
26*3160c8deSAhmed S. Darwish^^^^^^^^^^^^^^^^^^^
27*3160c8deSAhmed S. Darwish
28*3160c8deSAhmed S. Darwish:Expectation: enabled
29*3160c8deSAhmed S. Darwish:Severity: *high*
30*3160c8deSAhmed S. Darwish
31*3160c8deSAhmed S. DarwishThe CPU frequency scaling subsystem ensures that the processor can operate at
32*3160c8deSAhmed S. Darwishits maximum supported frequency.  While, in general, bootloaders are tasked
33*3160c8deSAhmed S. Darwishwith setting the CPU clock to the highest speed on boot, some do not.  It is
34*3160c8deSAhmed S. Darwishthus desirable to keep this option enabled.
35*3160c8deSAhmed S. Darwish
36*3160c8deSAhmed S. Darwish.. caution::
37*3160c8deSAhmed S. Darwish
38*3160c8deSAhmed S. Darwish  A real-time kernel is not about being "as fast as possible", however
39*3160c8deSAhmed S. Darwish  real-time requirements may demand that the CPU is clocked at a particular
40*3160c8deSAhmed S. Darwish  speed.
41*3160c8deSAhmed S. Darwish
42*3160c8deSAhmed S. Darwish``CONFIG_CPU_FREQ_DEFAULT_GOV_PERFORMANCE``
43*3160c8deSAhmed S. Darwish^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
44*3160c8deSAhmed S. Darwish
45*3160c8deSAhmed S. Darwish:Expectation: enabled
46*3160c8deSAhmed S. Darwish:Severity: *high*
47*3160c8deSAhmed S. Darwish
48*3160c8deSAhmed S. DarwishReal-Time workloads expect a fixed CPU frequency during execution.  Using the
49*3160c8deSAhmed S. Darwishperformance governor is an easy way to achieve that purely from kernel
50*3160c8deSAhmed S. Darwishconfiguration.
51*3160c8deSAhmed S. Darwish
52*3160c8deSAhmed S. DarwishThis is not an absolute rule.  Some setups might prefer to clock the CPU to
53*3160c8deSAhmed S. Darwishlower speeds due to thermal packaging or other requirements.  The key is that
54*3160c8deSAhmed S. Darwishthe CPU frequency remains constant once set.
55*3160c8deSAhmed S. Darwish
56*3160c8deSAhmed S. DarwishNon-performance CPU frequency governors
57*3160c8deSAhmed S. Darwish^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
58*3160c8deSAhmed S. Darwish
59*3160c8deSAhmed S. Darwish:Expectation: disabled
60*3160c8deSAhmed S. Darwish:Severity: *medium*
61*3160c8deSAhmed S. Darwish
62*3160c8deSAhmed S. DarwishTo ensure reproducible system latency measurements, disable the
63*3160c8deSAhmed S. Darwishnon-``PERFORMANCE`` CPU frequency governors whenever possible.  This avoids
64*3160c8deSAhmed S. Darwishthe risk of unknown userspace tasks implicitly or explicitly setting a
65*3160c8deSAhmed S. Darwishdifferent CPU frequency governor, and thereby changing latency behavior while
66*3160c8deSAhmed S. Darwishthe system is running.
67*3160c8deSAhmed S. Darwish
68*3160c8deSAhmed S. DarwishIf disabling other frequency governors is not an option, use a governor that
69*3160c8deSAhmed S. Darwishkeeps the CPU frequency fixed.  For example,
70*3160c8deSAhmed S. Darwish``CONFIG_CPU_FREQ_DEFAULT_GOV_USERSPACE`` can be enabled when userspace is
71*3160c8deSAhmed S. Darwishresponsible for setting a *stable* frequency during system initialization.
72*3160c8deSAhmed S. Darwish
73*3160c8deSAhmed S. DarwishIf a low CPU frequency is desired, then
74*3160c8deSAhmed S. Darwish``CONFIG_CPU_FREQ_DEFAULT_GOV_POWERSAVE`` can be set.
75*3160c8deSAhmed S. Darwish
76*3160c8deSAhmed S. DarwishThe ``ONDEMAND`` governor should not be enabled on a real-time system.  Its
77*3160c8deSAhmed S. Darwishfrequency changes depend on workload behavior and can significantly harm
78*3160c8deSAhmed S. Darwishdeterminism.
79*3160c8deSAhmed S. Darwish
80*3160c8deSAhmed S. DarwishFor more information, see Documentation/admin-guide/pm/cpufreq.rst
81*3160c8deSAhmed S. Darwish
82*3160c8deSAhmed S. Darwish``CONFIG_CPU_IDLE``
83*3160c8deSAhmed S. Darwish-------------------
84*3160c8deSAhmed S. Darwish
85*3160c8deSAhmed S. Darwish:Expectation: enabled
86*3160c8deSAhmed S. Darwish:Severity: *info*
87*3160c8deSAhmed S. Darwish
88*3160c8deSAhmed S. DarwishCPU idle states (C-states) allow the processor to enter low-power modes during
89*3160c8deSAhmed S. Darwishperiods of inactivity.  Very-low CPU idle states may require flushing the CPU
90*3160c8deSAhmed S. Darwishcaches and lowering or disabling the clocking.  This can lower power
91*3160c8deSAhmed S. Darwishconsumption, but it also increases the entry and exit latency from such
92*3160c8deSAhmed S. Darwishstates.
93*3160c8deSAhmed S. Darwish
94*3160c8deSAhmed S. DarwishWhile disabling this option eliminates cpuidle-related latencies, doing so can
95*3160c8deSAhmed S. Darwishsignificantly impact hardware longevity, warranty, and thermal behavior.
96*3160c8deSAhmed S. DarwishUsers should cap the maximum C-state to C1 instead.  For ACPI platforms, this
97*3160c8deSAhmed S. Darwishcan be achieved by using the boot parameter [1]_::
98*3160c8deSAhmed S. Darwish
99*3160c8deSAhmed S. Darwish  processor.max_cstate=1
100*3160c8deSAhmed S. Darwish
101*3160c8deSAhmed S. DarwishHigher C-states can be acceptable depending on the user workload's latency
102*3160c8deSAhmed S. Darwishrequirements.  For ACPI-based platforms, use the ``cpupower idle-info``
103*3160c8deSAhmed S. Darwishcommand to inspect the available idle states.
104*3160c8deSAhmed S. Darwish
105*3160c8deSAhmed S. DarwishFor more information, please see:
106*3160c8deSAhmed S. Darwish
107*3160c8deSAhmed S. Darwish- ``linux/tools/power/cpupower``
108*3160c8deSAhmed S. Darwish- Documentation/admin-guide/pm/cpuidle.rst
109*3160c8deSAhmed S. Darwish- Documentation/admin-guide/pm/index.rst
110*3160c8deSAhmed S. Darwish
111*3160c8deSAhmed S. Darwish``CONFIG_DRM``
112*3160c8deSAhmed S. Darwish--------------
113*3160c8deSAhmed S. Darwish
114*3160c8deSAhmed S. Darwish:Expectation: disabled
115*3160c8deSAhmed S. Darwish:Severity: *info*
116*3160c8deSAhmed S. Darwish
117*3160c8deSAhmed S. DarwishGPU-accelerated workloads can share system resources with the CPU, including
118*3160c8deSAhmed S. Darwishlast-level cache (LLC) and memory bandwidth.  Modern integrated GPUs optimize
119*3160c8deSAhmed S. Darwishgraphics performance at the expense of CPU determinism.
120*3160c8deSAhmed S. Darwish
121*3160c8deSAhmed S. DarwishExamples of affected platforms:
122*3160c8deSAhmed S. Darwish
123*3160c8deSAhmed S. Darwish- Intel processors with integrated graphics (Gen9 and later)
124*3160c8deSAhmed S. Darwish- AMD APUs with Radeon Graphics
125*3160c8deSAhmed S. Darwish- Xilinx Zynq UltraScale+ MPSoC EG/EV series
126*3160c8deSAhmed S. Darwish
127*3160c8deSAhmed S. DarwishIf graphics workloads must run alongside real-time tasks, users must conduct
128*3160c8deSAhmed S. Darwishthorough stress testing using tools like ``glmark2`` while measuring the
129*3160c8deSAhmed S. Darwishoverall system latency.
130*3160c8deSAhmed S. Darwish
131*3160c8deSAhmed S. DarwishFor more information, please check:
132*3160c8deSAhmed S. Darwish
133*3160c8deSAhmed S. Darwish- Documentation/core-api/real-time/hardware.rst ("Regarding hardware" section)
134*3160c8deSAhmed S. Darwish- Documentation/filesystems/resctrl.rst
135*3160c8deSAhmed S. Darwish- `Real-Time and Graphics: A Contradiction? <https://web.archive.org/web/20221025085614/https://linutronix.de/PDF/Realtime_and_graphics-acontradiction2021.pdf>`_
136*3160c8deSAhmed S. Darwish
137*3160c8deSAhmed S. Darwish``CONFIG_EFI_DISABLE_RUNTIME``
138*3160c8deSAhmed S. Darwish------------------------------
139*3160c8deSAhmed S. Darwish
140*3160c8deSAhmed S. Darwish:Expectation: enabled
141*3160c8deSAhmed S. Darwish:Severity: *medium*
142*3160c8deSAhmed S. Darwish
143*3160c8deSAhmed S. DarwishEFI is the standard boot and firmware interface for multiple architectures.
144*3160c8deSAhmed S. DarwishEFI runtime services provide callback functions to be called from the kernel;
145*3160c8deSAhmed S. Darwishe.g., as utilized by (``CONFIG_EFI_VARS*``) or (``CONFIG_RTC_DRV_EFI``).  For
146*3160c8deSAhmed S. Darwishthe former, the kernel calls into EFI to update the EFI variables.
147*3160c8deSAhmed S. Darwish
148*3160c8deSAhmed S. DarwishCalling into EFI means invoking firmware callbacks.  During such invocations,
149*3160c8deSAhmed S. Darwishthe system might not be able to react to interrupts and will thus not be able
150*3160c8deSAhmed S. Darwishto perform a context switch.  This can cause significant latency spikes for
151*3160c8deSAhmed S. Darwishthe real-time system.
152*3160c8deSAhmed S. Darwish
153*3160c8deSAhmed S. Darwish``CONFIG_PREEMPT_RT`` enables this option by default.  If this option is
154*3160c8deSAhmed S. Darwishmanually disabled at build time, the following boot parameter [1]_ may be used
155*3160c8deSAhmed S. Darwishto disable EFI runtime at boot up::
156*3160c8deSAhmed S. Darwish
157*3160c8deSAhmed S. Darwish  efi=noruntime
158*3160c8deSAhmed S. Darwish
159*3160c8deSAhmed S. DarwishAlternatively, confine EFI runtime service calls to a housekeeping CPU by
160*3160c8deSAhmed S. Darwishrestricting the ``efi_runtime`` workqueue CPU affinity.  For example, set that
161*3160c8deSAhmed S. Darwishworkqueue's affinity to CPU #0 and pin your RT tasks to a different CPU range.
162*3160c8deSAhmed S. DarwishSee Documentation/core-api/workqueue.rst
163*3160c8deSAhmed S. Darwish
164*3160c8deSAhmed S. Darwish``CONFIG_NO_HZ`` / ``CONFIG_NO_HZ_FULL``
165*3160c8deSAhmed S. Darwish----------------------------------------
166*3160c8deSAhmed S. Darwish
167*3160c8deSAhmed S. Darwish:Expectation: disabled
168*3160c8deSAhmed S. Darwish:Severity: *medium*
169*3160c8deSAhmed S. Darwish
170*3160c8deSAhmed S. DarwishTickless operation can increase kernel-to-userspace transition latency due to
171*3160c8deSAhmed S. Darwishthe extra accounting and state book-keeping.
172*3160c8deSAhmed S. Darwish
173*3160c8deSAhmed S. Darwish*Guidance by real-time workload type:*
174*3160c8deSAhmed S. Darwish
175*3160c8deSAhmed S. Darwish- For periodic workloads; e.g., control loops executing every 100 µs, avoid
176*3160c8deSAhmed S. Darwish  ``NO_HZ`` modes.  Consistent kernel ticks are preferable.
177*3160c8deSAhmed S. Darwish
178*3160c8deSAhmed S. Darwish- For computation-intensive workloads; e.g. extended userspace execution,
179*3160c8deSAhmed S. Darwish  ``NO_HZ_FULL`` may be beneficial.  In such cases, users should offload the
180*3160c8deSAhmed S. Darwish  kernel housekeeping to dedicated CPUs and isolate compute cores.
181*3160c8deSAhmed S. Darwish
182*3160c8deSAhmed S. DarwishSee also Documentation/timers/no_hz.rst
183*3160c8deSAhmed S. Darwish
184*3160c8deSAhmed S. Darwish``CONFIG_PREEMPT_RT``
185*3160c8deSAhmed S. Darwish---------------------
186*3160c8deSAhmed S. Darwish
187*3160c8deSAhmed S. Darwish:Expectation: enabled
188*3160c8deSAhmed S. Darwish:Severity: **fatal**
189*3160c8deSAhmed S. Darwish
190*3160c8deSAhmed S. DarwishThis option must be enabled, or the resulting kernel will not be fully
191*3160c8deSAhmed S. Darwishpreemptible and real-time capable.
192*3160c8deSAhmed S. Darwish
193*3160c8deSAhmed S. Darwish``CONFIG_TRACING`` (and tracing options)
194*3160c8deSAhmed S. Darwish----------------------------------------
195*3160c8deSAhmed S. Darwish
196*3160c8deSAhmed S. Darwish:Expectation: enabled
197*3160c8deSAhmed S. Darwish:Severity: *info*
198*3160c8deSAhmed S. Darwish
199*3160c8deSAhmed S. DarwishShipping kernels with tracing support enabled (but not actively running) is
200*3160c8deSAhmed S. Darwishhighly recommended.  This will allow the users to extract more information if
201*3160c8deSAhmed S. Darwishlatency problems arise.  Nonetheless, some tracers do incur latency overhead
202*3160c8deSAhmed S. Darwishjust by being enabled.
203*3160c8deSAhmed S. Darwish
204*3160c8deSAhmed S. Darwish.. caution::
205*3160c8deSAhmed S. Darwish
206*3160c8deSAhmed S. Darwish  Users should *not* make use of tracers or trace events during production
207*3160c8deSAhmed S. Darwish  real-time kernel operation as they can add considerable overhead and degrade
208*3160c8deSAhmed S. Darwish  the system's latency.
209*3160c8deSAhmed S. Darwish
210*3160c8deSAhmed S. Darwish``CONFIG_IRQSOFF_TRACER`` and ``CONFIG_PREEMPT_TRACER``
211*3160c8deSAhmed S. Darwish^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
212*3160c8deSAhmed S. Darwish
213*3160c8deSAhmed S. Darwish:Expectation: disabled
214*3160c8deSAhmed S. Darwish:Severity: *high*
215*3160c8deSAhmed S. Darwish
216*3160c8deSAhmed S. DarwishThese tracers do incur measurable latency overhead even when tracing is not
217*3160c8deSAhmed S. Darwishcurrently active.
218*3160c8deSAhmed S. Darwish
219*3160c8deSAhmed S. DarwishKernel Debug Options
220*3160c8deSAhmed S. Darwish====================
221*3160c8deSAhmed S. Darwish
222*3160c8deSAhmed S. DarwishMost kernel debug options add runtime overhead that increases the worst-case
223*3160c8deSAhmed S. Darwishlatency.
224*3160c8deSAhmed S. Darwish
225*3160c8deSAhmed S. Darwish.. caution::
226*3160c8deSAhmed S. Darwish
227*3160c8deSAhmed S. Darwish  During development and early testing, users are encouraged to run their
228*3160c8deSAhmed S. Darwish  real-time workloads and peripherals with lockdep (:ref:`lockdep`) and other
229*3160c8deSAhmed S. Darwish  kernel debug options enabled, for a considerable amount of time.  Such
230*3160c8deSAhmed S. Darwish  workloads might trigger kernel code paths that were not triggered during the
231*3160c8deSAhmed S. Darwish  internal Linux real-time kernel development, thus helping to uncover locking
232*3160c8deSAhmed S. Darwish  and other types of kernel bugs.
233*3160c8deSAhmed S. Darwish
234*3160c8deSAhmed S. Darwish``CONFIG_DEBUG_ATOMIC_SLEEP``
235*3160c8deSAhmed S. Darwish-----------------------------
236*3160c8deSAhmed S. Darwish
237*3160c8deSAhmed S. Darwish:Expectation: allowed
238*3160c8deSAhmed S. Darwish
239*3160c8deSAhmed S. DarwishThis sanity check catches common kernel programming errors with a tolerable
240*3160c8deSAhmed S. Darwishlatency cost.  It also increases overall scheduling as each ``might_sleep()``
241*3160c8deSAhmed S. Darwishcan lead to a context switch.
242*3160c8deSAhmed S. Darwish
243*3160c8deSAhmed S. Darwish``CONFIG_DEBUG_BUGVERBOSE`` and ``CONFIG_DEBUG_INFO*``
244*3160c8deSAhmed S. Darwish------------------------------------------------------
245*3160c8deSAhmed S. Darwish
246*3160c8deSAhmed S. Darwish:Expectation: allowed
247*3160c8deSAhmed S. Darwish
248*3160c8deSAhmed S. DarwishThese options increase the kernel image size but have no latency impact.  They
249*3160c8deSAhmed S. Darwishare also essential for meaningful BUG logs, crash dumps, and profiling.
250*3160c8deSAhmed S. Darwish
251*3160c8deSAhmed S. Darwish``CONFIG_DEBUG_FS``
252*3160c8deSAhmed S. Darwish-------------------
253*3160c8deSAhmed S. Darwish
254*3160c8deSAhmed S. Darwish:Expectation: allowed
255*3160c8deSAhmed S. Darwish
256*3160c8deSAhmed S. DarwishThis is safe to include in real-time kernels, *provided that debugfs is not
257*3160c8deSAhmed S. Darwishaccessed during production runtime*.
258*3160c8deSAhmed S. Darwish
259*3160c8deSAhmed S. Darwish``CONFIG_DEBUG_KERNEL``
260*3160c8deSAhmed S. Darwish-----------------------
261*3160c8deSAhmed S. Darwish
262*3160c8deSAhmed S. Darwish:Expectation: allowed
263*3160c8deSAhmed S. Darwish
264*3160c8deSAhmed S. DarwishMeta-option which allows debug features to be enabled.  It has no runtime
265*3160c8deSAhmed S. Darwishimpact, but beware of any debug features that it may have implicitly enabled.
266*3160c8deSAhmed S. Darwish
267*3160c8deSAhmed S. Darwish``CONFIG_LOCKUP_DETECTOR``
268*3160c8deSAhmed S. Darwish--------------------------
269*3160c8deSAhmed S. Darwish
270*3160c8deSAhmed S. Darwish:Expectation: disabled
271*3160c8deSAhmed S. Darwish:Severity: *high*
272*3160c8deSAhmed S. Darwish
273*3160c8deSAhmed S. DarwishThe lockup detector creates kernel timer callbacks that execute every few
274*3160c8deSAhmed S. Darwishseconds, in hard-IRQ context, even on real-time kernels.  These periodic
275*3160c8deSAhmed S. Darwishinterrupts can cause latency spikes.
276*3160c8deSAhmed S. Darwish
277*3160c8deSAhmed S. DarwishUsers should use hardware watchdogs instead, which will provide a similar
278*3160c8deSAhmed S. Darwishfunctionality without the software-induced latency.
279*3160c8deSAhmed S. Darwish
280*3160c8deSAhmed S. Darwish.. _lockdep:
281*3160c8deSAhmed S. Darwish
282*3160c8deSAhmed S. Darwish``CONFIG_PROVE_LOCKING``
283*3160c8deSAhmed S. Darwish------------------------
284*3160c8deSAhmed S. Darwish
285*3160c8deSAhmed S. Darwish:Expectation: disabled
286*3160c8deSAhmed S. Darwish:Severity: *high*
287*3160c8deSAhmed S. Darwish
288*3160c8deSAhmed S. DarwishProving the correctness of all kernel locking adds substantial overhead and
289*3160c8deSAhmed S. Darwishsignificantly increases worst-case latency.
290*3160c8deSAhmed S. Darwish
291*3160c8deSAhmed S. DarwishSummary
292*3160c8deSAhmed S. Darwish=======
293*3160c8deSAhmed S. Darwish
294*3160c8deSAhmed S. DarwishThere is no "one size fits all" solution for configuring a real-time Linux
295*3160c8deSAhmed S. Darwishsystem.  Beginning with the system real-time requirements, integrators must
296*3160c8deSAhmed S. Darwishconsider the features and functions of the system's hardware, kernel, and
297*3160c8deSAhmed S. Darwishuserspace.  All such components must be properly configured in order to
298*3160c8deSAhmed S. Darwishestablish and constrain the system's maximum latency.
299*3160c8deSAhmed S. Darwish
300*3160c8deSAhmed S. DarwishWith that in mind, any incorrect real-time kernel configuration could cause a
301*3160c8deSAhmed S. Darwishnew maximum latency that shows up at the wrong time and is catastrophic for
302*3160c8deSAhmed S. Darwishthe real-time system's latency.
303*3160c8deSAhmed S. Darwish
304*3160c8deSAhmed S. DarwishReferences
305*3160c8deSAhmed S. Darwish==========
306*3160c8deSAhmed S. Darwish
307*3160c8deSAhmed S. Darwish.. [1] See Documentation/admin-guide/kernel-parameters.rst
308