| #
11260c33 |
| 20-Aug-2026 |
Linus Torvalds <torvalds@linux-foundation.org> |
Merge tag 'sched_ext-for-7.3' of git://git.kernel.org/pub/scm/linux/kernel/git/tj/sched_ext
Pull sched_ext updates from Tejun Heo: "Most of this cycle completes the enqueue-path support for hierarc
Merge tag 'sched_ext-for-7.3' of git://git.kernel.org/pub/scm/linux/kernel/git/tj/sched_ext
Pull sched_ext updates from Tejun Heo: "Most of this cycle completes the enqueue-path support for hierarchical sub-scheduling, which makes sub-scheduler support feature complete: a root BPF scheduler can now hand a cgroup subtree over to a nested sub-scheduler together with revocable CPU grants, and the sub-scheduler owns all scheduling decisions for its tasks on those CPUs.
Development volume was high and a number of changes plugging holes in the new support landed late in the cycle. Also included are core scheduling fixes that were completed too late for the v7.2 release and are routed through this pull request.
Sub-scheduler CPU delegation:
- Parent schedulers now grant and revoke per-CPU capabilities (enqueueing, preemption, CPU frequency control) on their children, enforced on every path a scheduler can reach a CPU through. Previously only dispatching could be delegated; this lets sub-schedulers fully schedule their CPUs.
- Rescue execution: a task whose scheduler doesn't have access to the CPUs the task needs to run on starved until the watchdog ejected the whole scheduler. The kernel now runs such tasks directly on a small bandwidth budget, turning a scheduler-killing failure into bounded degradation.
- Cgroup integration: tasks migrating across a sub-scheduler boundary weren't re-homed to the new owner, causing wrong-scheduler scheduling and a use-after-free. Sub-schedulers now take over their cgroup subtree and receive its cgroup callbacks.
- Arena objects now cross the kernel/BPF boundary as typed pointer arguments, translated transparently by the BPF tree's new arena argument support, replacing untyped arguments with manual translation.
- scx_qmap now demonstrates full hierarchical sub-scheduling.
Other fixes and updates:
- Robustness improvements: the abort path is now NMI-safe, fixing deadlocks when errors are raised from NMI context and making hardlockup recovery direct. Reenqueue loops that could monopolize a CPU ahead of the watchdog now eject the offending scheduler, and stalls are blamed on the scheduler actually responsible.
- Hardening: BPF-writable arena memory is validated before kernel use, and task slice and vtime writes got explicit synchronization rules, closing corruption vectors open to buggy or malicious schedulers.
- Core scheduling: sched_ext dispatching can drop the rq lock inside the core-wide pick, which let interleaving selections corrupt each other's state and hard-hang the machine. The selection now restarts when the lock was released. The task ordering callback was also invoked with its arguments swapped, and the default ordering is updated to work across sub-scheduler boundaries. The fixes are marked for stable.
- Other fixes headed for stable: a task init leak on fork failure during enable, tooling compat macros that silently failed to detect newer kernels, and a crash on reenqueueing against a destroyed dispatch queue.
- Tooling: scx_pair moves off deprecated callbacks, and the deprecated scx_bpf_cpu_rq() kfunc is removed"
* tag 'sched_ext-for-7.3' of git://git.kernel.org/pub/scm/linux/kernel/git/tj/sched_ext: (144 commits) sched_ext: Drop the dead SCX_DEQ_CORE_SCHED_EXEC test in dequeue_task_scx() sched_ext: Make core-sched task ordering hierarchy-aware sched_ext: Use runnable_at for the default core-sched task ordering sched_ext: Fix inverted ops.core_sched_before() invocation sched_ext: Move the config-off sub-cap kfunc stubs into sub.c sched_ext: Rename balance-era identifiers to dispatch terms sched_ext: Drop the stale keep_prev fixup in dispatch_pick() sched_ext: Keep kick_sync waiting on the rq's own CPU sched_ext: Make SCHED_CLASS_EXT select GENERIC_ALLOCATOR sched_ext/scx_flatcg: Fix cvtime true-up on slice expiry sched_ext: Don't BUG_ON a destroyed DSQ in process_deferred_reenq_users sched_ext: Fix scx_bpf_dsq_move_to_local___v2 compat detection sched_ext: Make scx_bpf_events() read the calling scheduler's counters sched_ext: Drop unlocked scx_rq_clock_invalidate() from scx_root_disable() selftests/sched_ext: Fix flaky ddsp failure tests on busy systems selftests/sched_ext: Make numa idle validation race-free sched_ext: Fix scx_bpf_dsq_reenq___compat kfunc extern prototype sched_ext/scx_flatcg: expire cached hweights on weight changes sched_ext: Fix exit_task leak on fork failure during enable sched_ext: fix stale references in doc comments ...
show more ...
|
|
Revision tags: v7.2, v7.2-rc7 |
|
| #
5fd50174 |
| 03-Aug-2026 |
Tejun Heo <tj@kernel.org> |
sched_ext: Add bandwidth-limited rescue execution for stranded tasks
A local DSQ insert lacking the needed caps is diverted to the reject DSQ and bounced back through ops.enqueue() so the scheduler
sched_ext: Add bandwidth-limited rescue execution for stranded tasks
A local DSQ insert lacking the needed caps is diverted to the reject DSQ and bounced back through ops.enqueue() so the scheduler can re-decide. That recovery assumes the scheduler has somewhere legal to send the task. When it doesn't, e.g. when the task's affinity is restricted to cids delegated away, the task starves until the stall watchdog ejects the scheduler. An exiting task is worse - it skips ops.enqueue() and the rejection becomes a self-requeuing cycle that burns the CPU until the watchdog fires.
Add SCX_ENQ_RESCUE, a fallback modifier on local DSQ inserts. When the insert would be rejected for missing caps, the kernel takes over and runs the task on the target CPU without consulting the owning scheduler. The kernel sets the flag itself when enqueueing an exiting task.
Rescue is a last-resort forward-progress backstop with a persistent disadvantage, not a way around cap enforcement. A per-CPU token bucket accrues rescue_bandwidth_ppt (default 2%) of CPU time and rescues run one at a time in arrival order. Each is granted a slice of the rescue_quantum_us (default 5ms) quantum divided across the waiters, waits at the tail of the local DSQ claiming no priority, and rejoins its scheduler as a fresh arrival once the slice is served.
The schedulers keep their normal control over an admitted rescuee and may preempt or reslice it. Service is measured on CPU time actually received, so neither shortens the rescue. Prolonged denial escalates - the remaining slice turns into protected execution (SCX_TASK_PROTECTED) and the rescuee preempts the current task. Escalation is paced by the same bucket, and delivered service converges on the configured bandwidth no matter how aggressively the schedulers dispatch.
Both knobs are root-only and SCX_RESCUE_DISABLE turns rescue off, making SCX_ENQ_RESCUE inserts reject as usual.
v2: - Add SCX_OPS_OPEN() fix-ups for the new ops fields so cpu-form schedulers setting them still load on older kernels. (Andrea)
Signed-off-by: Tejun Heo <tj@kernel.org> Reviewed-by: Andrea Righi <arighi@nvidia.com>
show more ...
|
| #
78f8d726 |
| 03-Aug-2026 |
Tejun Heo <tj@kernel.org> |
sched_ext: Make SCX_ENQ_IGNORE_CAPS waive the preemption cap too
SCX_ENQ_IGNORE_CAPS is kernel-internal and marks a placement the kernel forces. scx_caps_for_enq() waives the enqueue cap for it, but
sched_ext: Make SCX_ENQ_IGNORE_CAPS waive the preemption cap too
SCX_ENQ_IGNORE_CAPS is kernel-internal and marks a placement the kernel forces. scx_caps_for_enq() waives the enqueue cap for it, but a PREEMPT insert still picks up the preemption cap requirement from scx_caps_for_preempt(). Update scx_caps_for_preempt() to take enq_flags and require nothing when SCX_ENQ_IGNORE_CAPS is set.
Signed-off-by: Tejun Heo <tj@kernel.org> Reviewed-by: Andrea Righi <arighi@nvidia.com>
show more ...
|
| #
8b3b8522 |
| 03-Aug-2026 |
Tejun Heo <tj@kernel.org> |
sched_ext: Rename scx_local_or_reject_dsq() to scx_resolve_local_dsq()
The following rescue execution addition gives the function a third possible destination, making a name that enumerates the outc
sched_ext: Rename scx_local_or_reject_dsq() to scx_resolve_local_dsq()
The following rescue execution addition gives the function a third possible destination, making a name that enumerates the outcomes a poor fit. Rename to the destination-neutral scx_resolve_local_dsq(). No functional changes.
Signed-off-by: Tejun Heo <tj@kernel.org> Reviewed-by: Andrea Righi <arighi@nvidia.com>
show more ...
|
|
Revision tags: v7.2-rc6, v7.2-rc5 |
|
| #
79474420 |
| 24-Jul-2026 |
Tejun Heo <tj@kernel.org> |
sched_ext: Add scx_cgroup_sched() for cgrp->scx_sched reads
cgrp->scx_sched is __rcu and published with rcu_assign_pointer() but every reader loads it with a plain access, so sparse flags all of the
sched_ext: Add scx_cgroup_sched() for cgrp->scx_sched reads
cgrp->scx_sched is __rcu and published with rcu_assign_pointer() but every reader loads it with a plain access, so sparse flags all of them. The reads are lock-protected: enable/disable paths rewrite the field under all of scx_enable_mutex, scx_fork_rwsem and cgroup_mutex, and cgroup creation inherits the parent's sched under cgroup_mutex before the new cgroup is reachable, so holding any one of the three locks makes the read stable.
Add scx_cgroup_sched() which states the protection with rcu_dereference_check() and convert the readers. No functional changes.
Signed-off-by: Tejun Heo <tj@kernel.org>
show more ...
|
|
Revision tags: v7.2-rc4 |
|
| #
8946dbd3 |
| 16-Jul-2026 |
Tejun Heo <tj@kernel.org> |
sched_ext: Add the scx_has_subs static key and gate sub-sched hot paths
With CONFIG_EXT_SUB_SCHED=y but no sub-scheduler attached - the common case - hot paths still pay for sub-sched bookkeeping. G
sched_ext: Add the scx_has_subs static key and gate sub-sched hot paths
With CONFIG_EXT_SUB_SCHED=y but no sub-scheduler attached - the common case - hot paths still pay for sub-sched bookkeeping. Gate it behind __scx_has_subs, a static key counting live sub-schedulers, so that a root-only system stops paying.
Most conversions are simple skip-if-no-sub tests. scx_idle_notify() is special - it's a hierarchy walk, so give it a fast path which notifies the root directly using the same tests as the walk. A pending SCX_RQ_SUB_IDLE_RENOTIFY can be ignored as no sub can be owed one and the caller clears the flag either way.
Suggested-by: Andrea Righi <arighi@nvidia.com> Signed-off-by: Tejun Heo <tj@kernel.org> Reviewed-by: Andrea Righi <arighi@nvidia.com>
show more ...
|
| #
7f480f34 |
| 16-Jul-2026 |
Tejun Heo <tj@kernel.org> |
sched_ext: Move scx_dispatch_sched() to a new inlines.h
scx_dispatch_sched() is common dispatch machinery and looks out of place in sub.h, but it needs scx_cpu_arg() from cid.h and can't move into i
sched_ext: Move scx_dispatch_sched() to a new inlines.h
scx_dispatch_sched() is common dispatch machinery and looks out of place in sub.h, but it needs scx_cpu_arg() from cid.h and can't move into internal.h without creating a circular include. Add inlines.h on top of internal.h and cid.h, and move the function there. The function was sub.h's only cid.h user, so drop that include. Pure code move, no functional change.
v2: Host the function in a new inlines.h instead of at internal.h's tail, which formed a circular include with cid.h. Drop sub.h's now-unused cid.h include. (sashiko AI)
Signed-off-by: Tejun Heo <tj@kernel.org> Reviewed-by: Andrea Righi <arighi@nvidia.com>
show more ...
|
| #
75c268ed |
| 14-Jul-2026 |
Tejun Heo <tj@kernel.org> |
sched_ext: Replay ecaps notifications suppressed by bypass
scx_process_sync_ecaps() consumes ecaps syncs while the sched is bypassing without delivering ops.sub_ecaps_updated(), leaving reported_eca
sched_ext: Replay ecaps notifications suppressed by bypass
scx_process_sync_ecaps() consumes ecaps syncs while the sched is bypassing without delivering ops.sub_ecaps_updated(), leaving reported_ecaps stale. Nothing re-queued a sync when bypass lifted, so a cid whose caps never change again would never be notified. Attach-time initial grants hit this every time: they are consumed during the enable bypass window, so a sched never learned its initial effective caps through the callback.
Re-queue a sync for every (sched, cpu) with an undelivered delta at the per-cpu bypass exit in scx_bypass(), next to the idle renotify catch-up. The next balance on the cpu then delivers the pending delta with proper dispatch context.
Signed-off-by: Tejun Heo <tj@kernel.org> Reviewed-by: Andrea Righi <arighi@nvidia.com>
show more ...
|
| #
6ea3be36 |
| 14-Jul-2026 |
Tejun Heo <tj@kernel.org> |
sched_ext: Gate kicks on SCX_CAP_BASE and preemption on SCX_CAP_PREEMPT
A kick forces a scheduling event on the target cpu, and a preemption also evicts the running task. Gate both on caps. Any kick
sched_ext: Gate kicks on SCX_CAP_BASE and preemption on SCX_CAP_PREEMPT
A kick forces a scheduling event on the target cpu, and a preemption also evicts the running task. Gate both on caps. Any kick requires baseline access on the cid, and preempting a task the sub-sched does not own - whether by a SCX_ENQ_PREEMPT insert or a SCX_KICK_PREEMPT kick - requires the new SCX_CAP_PREEMPT. Gating either alone would leave a hole - the weakest cap authorizing preempting kicks, or plain kicks disturbing cpus the kicker has no access to.
Preempting the sched's own subtree is always allowed, and the cap extends the right to any task on the cid. PREEMPT implies ENQ, and so ENQ_IMMED.
A preempting insert tests the running task under the target rq lock and is rejected and reenqueued unless the victim is in the inserter's subtree or it holds PREEMPT. A migration-disabled task is admitted regardless, but with SCX_ENQ_PREEMPT stripped.
Kicks are enforced on the delivery path, where the effective caps can be read coherently under the target rq's lock. A kick from a sub-sched lacking SCX_CAP_BASE on the cid is dropped, and a SCX_KICK_PREEMPT kick without PREEMPT for a task outside the kicker's subtree degrades to a plain reschedule.
Unlike the enqueue caps, PREEMPT is checked only at the instant of the insert or kick, never as a standing property of a queued task.
v2: Clear SCX_ENQ_PREEMPT on the offline and migration_pending force-admits.
Signed-off-by: Tejun Heo <tj@kernel.org> Reviewed-by: Andrea Righi <arighi@nvidia.com>
show more ...
|
| #
701b7bca |
| 14-Jul-2026 |
Tejun Heo <tj@kernel.org> |
sched_ext: Add the SCX_CAP_ENQ cap
Add SCX_CAP_ENQ, which gates inserting tasks onto a cid's local DSQ. Unlike IMMED enqueue, plain enqueues can pile up, so ENQ is the stronger cap and implies ENQ_I
sched_ext: Add the SCX_CAP_ENQ cap
Add SCX_CAP_ENQ, which gates inserting tasks onto a cid's local DSQ. Unlike IMMED enqueue, plain enqueues can pile up, so ENQ is the stronger cap and implies ENQ_IMMED. Losing ENQ also triggers the reenq scan. The scan tests each queued task and the running task against the cap each needs via scx_caps_for_task(), so an ENQ-only loss reenqueues plain tasks, evicting a running one, while IMMED tasks, which need only ENQ_IMMED, stay put.
Signed-off-by: Tejun Heo <tj@kernel.org> Reviewed-by: Andrea Righi <arighi@nvidia.com>
show more ...
|
| #
46a85ae6 |
| 14-Jul-2026 |
Tejun Heo <tj@kernel.org> |
sched_ext: Tie cpu occupancy to SCX_CAP_BASE through the task slice
A task's slice grants it cpu occupancy - how long it holds its cpu. In a sub-scheduler hierarchy cpu access is delegated through r
sched_ext: Tie cpu occupancy to SCX_CAP_BASE through the task slice
A task's slice grants it cpu occupancy - how long it holds its cpu. In a sub-scheduler hierarchy cpu access is delegated through revocable capabilities, so a task's occupancy must follow them. Only its own scheduler sets its slice, and extending the slice is allowed only while that scheduler holds baseline cpu access (SCX_CAP_BASE) on the cpu. Otherwise a scheduler could keep occupying a cpu it has been denied simply by handing out long slices.
The cap check reads effective caps, which are coherent only under the task's rq lock, and the kernel decrements the slice under that lock as the task runs, so a running task's slice can be changed only there while a queued task's can be set directly. Make scx_bpf_task_set_slice() apply the slice under the rq lock. Synchronously when the caller already holds it, otherwise by stashing it in the new p->scx.slice_oob, tagged with the scheduler's id so a request that outlived a reassignment is dropped. Whether the caller holds @p's current rq lock is tested with p->scx.runnable_cpu.
Revocation is enforced through the same grant. When a cpu's effective caps lose SCX_CAP_BASE, the cap-revoke reenq scan also checks the running task and zeroes its slice to evict it. The scan runs as a balance callback after the pick, so this catches both the task that was running when the revoke landed and a capless task the pick just promoted off the local DSQ. The paths that keep a task on its cpu - holding on to the last runnable task in balance, the ENQ_LAST reinsertion and the slice refill on pick - skip tasks lacking baseline access. A migration-disabled task is exempt, mirroring its capless admission on insert.
v4: Test rq ownership with p->scx.runnable_cpu, closing a remote-wakeup TOCTOU. (sashiko AI) v3: Keep a pending out-of-band slice request across refill and preserve. (sashiko AI) v2: Only write slice directly when @p is queued on the held rq. (sashiko AI)
Signed-off-by: Tejun Heo <tj@kernel.org> Reviewed-by: Andrea Righi <arighi@nvidia.com>
show more ...
|
| #
147d1885 |
| 14-Jul-2026 |
Tejun Heo <tj@kernel.org> |
sched_ext: Add the SCX_CAP_ENQ_IMMED cap
Replace the __SCX_CAP_DUMMY placeholder with SCX_CAP_ENQ_IMMED, which gates inserting IMMED tasks onto a cid's local DSQ. An IMMED enqueue is guaranteed to e
sched_ext: Add the SCX_CAP_ENQ_IMMED cap
Replace the __SCX_CAP_DUMMY placeholder with SCX_CAP_ENQ_IMMED, which gates inserting IMMED tasks onto a cid's local DSQ. An IMMED enqueue is guaranteed to either get its task running on the cpu at once or hand it back to the scheduler, so IMMED work can never pile up on the cpu's queue and a cpu can be shared across sub-scheds through IMMED access without any of them swamping it.
That makes ENQ_IMMED the natural baseline, the minimal cap to make any use of a cpu. SCX_CAP_BASE aliases it so gates on basic cpu access can state the intention instead of naming ENQ_IMMED.
Enforcement covers inserts and queued tasks. An insert without the cap is diverted to the reject DSQ, and queued tasks are reenqueued when the cap is lost. scx_bpf_sub_dispatch() skips a child that lacks the cap on the cpu, as its inserts would only be rejected. Vacating the running task on cap loss lands in a later patch.
Signed-off-by: Tejun Heo <tj@kernel.org> Reviewed-by: Andrea Righi <arighi@nvidia.com>
show more ...
|
| #
8b175234 |
| 14-Jul-2026 |
Tejun Heo <tj@kernel.org> |
sched_ext: Add SCX_ENQ_IGNORE_CAPS for in-place restore
A SAVE/RESTORE requeue re-inserts a running task in place and is immediately followed by set_next_task_scx(). It is not a real scheduling even
sched_ext: Add SCX_ENQ_IGNORE_CAPS for in-place restore
A SAVE/RESTORE requeue re-inserts a running task in place and is immediately followed by set_next_task_scx(). It is not a real scheduling event: the task is already admitted to its cid and must return to the local DSQ unconditionally.
scx_caps_for_enq() maps an enqueue to the cap its local-DSQ insert requires. Add SCX_ENQ_IGNORE_CAPS, set it on the RESTORE-in-place branch of enqueue_task_scx(), and have scx_caps_for_enq() require no caps for it, so the cid admission gate never diverts an in-place restore to the reject DSQ.
Signed-off-by: Tejun Heo <tj@kernel.org> Reviewed-by: Andrea Righi <arighi@nvidia.com>
show more ...
|
| #
75a8c820 |
| 14-Jul-2026 |
Tejun Heo <tj@kernel.org> |
sched_ext: Add reject DSQ for cap-rejected dispatches
When a sub-scheduler dispatches a task to a CPU it lacks the required capability on, the task must be rejected rather than allowed to run.
Add
sched_ext: Add reject DSQ for cap-rejected dispatches
When a sub-scheduler dispatches a task to a CPU it lacks the required capability on, the task must be rejected rather than allowed to run.
Add the machinery for that. Each rq gets a reject DSQ, a kernel-internal holding queue that is never run and that the BPF scheduler cannot reach. An insert that must be refused is diverted there instead of the local DSQ, and a deferred requeue then hands the parked tasks back to the BPF scheduler to re-decide. A cap revoke extends this to already-queued tasks. When the revoke reaches the cpu's effective caps, the cpu scans its local DSQ and reenqueues the tasks that no longer qualify.
A migration-disabled task must run on its cpu, so a capless one is admitted anyway and counted in the new SCX_EV_SUB_FORCED_ADMIT event.
This is preparation for the actual sub-sched cap enforcement. The divert is wired but inert here.
v2: Admit offline-rq and migration_pending inserts to local, not reject. (sashiko AI)
Signed-off-by: Tejun Heo <tj@kernel.org> Reviewed-by: Andrea Righi <arighi@nvidia.com>
show more ...
|
| #
b81a6c01 |
| 14-Jul-2026 |
Tejun Heo <tj@kernel.org> |
sched_ext: Add sub_ecaps_updated() effective-cap change notifier
A sub-scheduler that gains or loses effective caps on a cpu may want to act on it right away - e.g. place or preempt on a newly usabl
sched_ext: Add sub_ecaps_updated() effective-cap change notifier
A sub-scheduler that gains or loses effective caps on a cpu may want to act on it right away - e.g. place or preempt on a newly usable cpu. The existing ops.sub_caps_updated() doesn't fit as it is delivered asynchronously to scheduling operations and can arrive before the per-cpu effective caps go live.
Add ops.sub_ecaps_updated(cid, before, after), a cid-form callback fired from scx_process_sync_ecaps() when a sub-sched's effective caps on a cid change. It runs in dispatch context so the sched can insert, kick or preempt on the cid directly. @before is the caps as of the last delivery.
Cpu hotplug rides the same machinery. Going down zeroes each sched's ecaps on the cpu's cid, with queued syncs discarded at consumption while the cpu is inactive. Coming back up queues a sync for every sched. reported_ecaps is kept across the down/up cycle, so the resync fires the callback only if ownership actually changed while the cpu was down.
v2: Compute cid below the active-cpu guard; discard queued syncs on !cpu_active(). (sashiko AI)
Signed-off-by: Tejun Heo <tj@kernel.org> Reviewed-by: Andrea Righi <arighi@nvidia.com>
show more ...
|
| #
56fdc35b |
| 14-Jul-2026 |
Tejun Heo <tj@kernel.org> |
sched_ext: Maintain per-cpu effective cap copies for single-read checks
Checking a sched's caps on a cid would need to test several cap bits against caps[] to account for implied caps. Also, caps[]
sched_ext: Maintain per-cpu effective cap copies for single-read checks
Checking a sched's caps on a cid would need to test several cap bits against caps[] to account for implied caps. Also, caps[] modifications aren't synchronized against scheduling operations on each cpu, which can lead to awkward race conditions.
Collect them per cpu instead. caps[] under pshard->lock stays the target configuration. scx_sched_pcpu->ecaps is added, the transposed effective copy: the set of cap bits the sched holds on that cpu which can be accessed with a single read. It is stable under the rq lock. It can also be read locklessly with READ_ONCE().
Grant and revoke only mutate caps[]. They queue a sync request on the target cpu's rq->scx.ecaps_to_sync and kick it, and the cpu recomputes the queued scheds' ecaps from caps[] in balance_one() under its own rq lock. A dying sched runs the sync directly to retire its queued request before freeing. As held references can defer the freeing past the enclosing root scheduler's lifetime, root enable discards leftover sync requests before going live.
Signed-off-by: Tejun Heo <tj@kernel.org> Reviewed-by: Andrea Righi <arighi@nvidia.com>
show more ...
|
| #
5f2a9a4c |
| 14-Jul-2026 |
Tejun Heo <tj@kernel.org> |
sched_ext: Add coalescing sub_caps_updated() notifier for sub-schedulers
Wire up ops_cid.sub_caps_updated() to notify sub-scheds of cap changes.
Three constraints shape the design:
1. Static mem
sched_ext: Add coalescing sub_caps_updated() notifier for sub-schedulers
Wire up ops_cid.sub_caps_updated() to notify sub-scheds of cap changes.
Three constraints shape the design:
1. Static memory. Deliveries use a fixed-size buffer, both for runtime efficiency and so notifications can't be lost under memory pressure.
2. High-frequency updates. Grant/revoke can mutate caps in bursts, and the notifier path must absorb that without amplifying it.
3. Recursive grant/revoke from the callback. A child receiving a notification can call grant/revoke on its own children, which can cascade recursively down its subtree.
(1) and (2) lead to coalescing into a fixed payload. Each delivery carries a single (cmask, caps) pair covering every change since the previous one. Direction (set vs cleared) isn't encoded as it doesn't fit in the fixed-size summary. The callback queries scx_bpf_sub_caps() for current state. Only one delivery is in flight per shard. Further changes fold into the same buffer and ship as the next callback, so a shard's callbacks fire in order.
(3) leads to deferred delivery. Events accumulate during grant/revoke and are delivered after the shard lock is released.
v2: - Request a private stack for ops.sub_caps_updated(). (sashiko AI) - Build cmask_arena_out via scx_cmask_ref, not by re-reading its header.
Signed-off-by: Tejun Heo <tj@kernel.org> Reviewed-by: Andrea Righi <arighi@nvidia.com>
show more ...
|
| #
86094b95 |
| 14-Jul-2026 |
Tejun Heo <tj@kernel.org> |
sched_ext: Add per-shard cap delegation for sub-schedulers
Caps are per-cid permissions parents delegate to direct children via scx_bpf_sub_grant() / scx_bpf_sub_revoke(). A child's cap set is alway
sched_ext: Add per-shard cap delegation for sub-schedulers
Caps are per-cid permissions parents delegate to direct children via scx_bpf_sub_grant() / scx_bpf_sub_revoke(). A child's cap set is always a subset of its parent's. Sub-scheds check their caps locally, and cross-sched communication is needed only when the delegation set itself changes.
Caps will be used to implement sub-sched scheduling on the enqueue path. Picking a cid for a task at a leaf depends on which cids the leaf is allowed to use, and resolving that programmatically on every enqueue would mean a cross-sched round-trip call chain, possibly retrying if the request can't be granted as-is. The dispatch path is different - it runs as top-down recursion via scx_bpf_sub_dispatch().
Locking is per shard. cid space is split into shards, and each sub-sched has its own pshard->lock for each shard. Operations are broken up on shard boundaries. Different shards never contend. Shards are expected to be topology-aligned and likely to serve as the locality unit when cids are allocated to schedulers, so per-shard lock granularity scales naturally with the allocation pattern.
This patch adds the framework with a single dummy cap. Real caps land in later patches.
The enable path is reordered for pshards. scx_arena_pool_init() moves ahead of scx_link_sched() so the pshards are allocated before the sched becomes reachable - scx_alloc_pshards() skips allocation when the arena pool isn't initialized.
- scx_bpf_sub_grant(): Per-cid all-or-nothing grant to direct child. - scx_bpf_sub_revoke(): Clear caps on @cmask across @child and its subtree. - scx_bpf_sub_caps(): Lockless snapshot of caps on a cid range.
/sys/kernel/sched_ext/SCHED/caps shows the caps each scheduler currently holds.
v4: Move the pshard[] full build/publish and the err_disable scx_error() recording to earlier patches. (sashiko AI) v3: Build pshard[] fully before publishing it, read it with READ_ONCE. (sashiko AI) v2: Validate ops before scx_link_sched() publishes the sub. (sashiko AI)
Signed-off-by: Tejun Heo <tj@kernel.org> Reviewed-by: Andrea Righi <arighi@nvidia.com>
show more ...
|
| #
bbda59d8 |
| 14-Jul-2026 |
Tejun Heo <tj@kernel.org> |
sched_ext: Add scx_skip_subtree_pre()
Factor the sibling/ancestor portion of scx_next_descendant_pre() out as scx_skip_subtree_pre(), a pre-order walk primitive that skips @pos's subtree, and call i
sched_ext: Add scx_skip_subtree_pre()
Factor the sibling/ancestor portion of scx_next_descendant_pre() out as scx_skip_subtree_pre(), a pre-order walk primitive that skips @pos's subtree, and call it from scx_next_descendant_pre(). Same locking rules as the existing primitive.
Used in a follow-up to fast-skip subtrees that have nothing to do during a descendant walk.
Signed-off-by: Tejun Heo <tj@kernel.org> Reviewed-by: Andrea Righi <arighi@nvidia.com>
show more ...
|
| #
70f8b178 |
| 14-Jul-2026 |
Tejun Heo <tj@kernel.org> |
sched_ext: RCU-protect the sub-sched tree's children/sibling lists
Future kfuncs need to walk descendants without scx_sched_lock. Make the walker RCU-safe so that they can. A sub-sched's fields are
sched_ext: RCU-protect the sub-sched tree's children/sibling lists
Future kfuncs need to walk descendants without scx_sched_lock. Make the walker RCU-safe so that they can. A sub-sched's fields are initialized before it is linked, so a walk that observes a linked node also observes its setup. In-place changes after linking carry their own ordering.
Switch the children/sibling list ops to RCU and expand the descendant walker to accept rcu_read_lock as a valid read-side context. Walkers that mutate keep scx_sched_lock.
A sub-sched can be linked while an ancestor is bypassing, after the bypass walk that propagates the depth has passed its parent. Bypass state is a per-cpu flag plus a depth count and can't be established atomically at link time, so refuse to link under a bypassing ancestor. Take scx_bypass_lock across linking to check the parent's bypass state coherently.
v3: Reject linking under a bypassing ancestor instead of inheriting bypass_depth. (sashiko AI) v2: Inherit bypass_depth before publishing @sch on the RCU sibling list.
Signed-off-by: Tejun Heo <tj@kernel.org> Reviewed-by: Andrea Righi <arighi@nvidia.com>
show more ...
|
| #
8dba3bbd |
| 14-Jul-2026 |
Tejun Heo <tj@kernel.org> |
sched_ext: Add per-shard scx_sched storage scaffolding
Add struct scx_pshard and sch->pshard[] indexed by shard_idx, each entry allocated on its shard's NUMA node from scx_shard_node[si]. The struct
sched_ext: Add per-shard scx_sched storage scaffolding
Add struct scx_pshard and sch->pshard[] indexed by shard_idx, each entry allocated on its shard's NUMA node from scx_shard_node[si]. The struct starts empty (one dummy field). Follow-up patches will grow it as shard-local state lands. Only cid-type schedulers with an arena pool get pshards.
Allocation happens after ops.init_cids() returns so any scx_bpf_cid_override() it issues has finalized scx_nr_cid_shards and scx_shard_node[]. sch->nr_pshards records the array size for the async RCU free path, which may run after a later scheduler's scx_cid_init() has rewritten the global.
v3: Build and publish pshard[] fully-formed here rather than a later patch. v2: Free the partially-allocated pshard array on alloc failure. (sashiko AI)
Signed-off-by: Tejun Heo <tj@kernel.org> Reviewed-by: Andrea Righi <arighi@nvidia.com>
show more ...
|
|
Revision tags: v7.2-rc3 |
|
| #
cbcda14b |
| 10-Jul-2026 |
Pat Somaru <patso@likewhatevs.io> |
sched_ext: Add tracepoint for scheduler exit
sched_ext schedulers have state in BPF programs and kernel. scx_dump provides kernel state and BPF program state on error, but this is static in what it
sched_ext: Add tracepoint for scheduler exit
sched_ext schedulers have state in BPF programs and kernel. scx_dump provides kernel state and BPF program state on error, but this is static in what it can provide.
Add a sched_ext_exit tracepoint in scx_claim_exit() so that BPF programs can dynamically inspect scheduler specific state at the moment of exit. Pass the exiting scx_sched so attached programs can read its state, and, since exits propagate through a hierarchy of sub-schedulers, identify which scheduler each event belongs to.
Signed-off-by: Pat Somaru <patso@likewhatevs.io> Signed-off-by: Tejun Heo <tj@kernel.org>
show more ...
|
|
Revision tags: v7.2-rc2 |
|
| #
daf8e166 |
| 01-Jul-2026 |
Tejun Heo <tj@kernel.org> |
sched_ext: Split sub-scheduler implementation into sub.c
The sub-scheduler implementation has grown and will continue to expand. Move the sub-scheduler functions from ext.c into a new kernel/sched/e
sched_ext: Split sub-scheduler implementation into sub.c
The sub-scheduler implementation has grown and will continue to expand. Move the sub-scheduler functions from ext.c into a new kernel/sched/ext/sub.c. sub.h holds the prototypes and the !CONFIG_EXT_SUB_SCHED no-op stubs.
scx_dispatch_sched() is shared: balance_one() in ext.c and the scx_bpf_sub_dispatch() kfunc in sub.c both call it, and the latter re-enters it as sub-scheduler dispatch nests. It moves into sub.h as a static __always_inline so both callers keep it inlined and per-level stack stays bounded across the recursion. The event macros it uses move to internal.h.
No functional change.
Signed-off-by: Tejun Heo <tj@kernel.org> Reviewed-by: Andrea Righi <arighi@nvidia.com>
show more ...
|