A Linux CPU scheduler written in BPF that figures out which of your threads are urgent — without anyone telling it.
In 2023 SteamOS still ran CFS, and tuning the
scheduler meant a handful of debugfs knobs. Igalia asked a blunter
question:
Most of what strace saw was Wine
synchronisation — tasks waiting on each other.
Figure: Changwoo Min, OSS NA 2024 · Igalia VaporMark
sched_ext reached mainline in 6.12, over a year later
A steady 144 FPS and 144 FPS with one 40 ms hitch per second have the same average — only one is playable.
The frames that ruin it were spent waiting in a run queue, not rendering.
Same workload, two schedulers — kept anonymous as A and B. Measured and published by Changwoo Min, OSS NA 2024. B peaks higher and averages only a little lower; its worst 1% is where the stutter people complain about lives.
Each segment is coloured by the act of this talk that covers it — criticality and deadline, topology, power — so this map is also the route ahead. Today: 904 commits, 42 contributors, ~12,100 lines.
Delay B and you delay everything downstream of it.
How often the task sleeps waiting on something external — I/O, a timer, a packet, a mutex.
How often the task wakes somebody else.
Average invariant runtime per schedule. Shorter is more critical
In the code it is log-compressed, then squared: lat_cri = ( log₂(wait_ft·wake_ft) + log₂(rt_ft·w_ft) )2
LAVD reads tasks as a connected graph, not in isolation: urgency flows forward to whoever you wake, and backward to whoever is blocking you.
Everything the scheduler knows collapses into one number — and lat_cri divides, so ten times more critical means a tenth of the distance to wait. The clock keeps advancing, so even the least critical task eventually wins: priority that decays into fairness.
Every queued task should get a turn within 10 ms, across all the active CPUs — so the slice is that budget divided by how many are waiting. The clamps stop it becoming either a context-switch storm or a monopoly.
p's domain that p may run on.lat_cri
and a later predicted finish, or you evict a task that was about to release the CPU
anyway.A latency-critical task is only as fast as whoever holds the lock it wants — so LAVD traces futex operations to find the holder, then protects and boosts it.
Each compute domain is one L3 cache, so a domain's queue holds tasks that already share a cache. But no CPU should sit idle while another domain has work waiting — so a CPU that runs dry may pull from beyond its own domain.
LAVD is a domain-DSQ scheduler: every domain has one dispatch queue, shared by the CPUs inside it.
An underloaded domain. It periodically pulls tasks from stealees to keep its own CPUs busy.
An overloaded domain, carrying more queued load than its fair share. Stealers may pull from it.
Figure adapted from “Evolving sched_ext: Resource Control, Topology Awareness, and Energy Efficiency for Modern Systems” — Changwoo Min and Gavin Guo, April 2026.
The target is not an average. A domain with 40% of the machine's capacity should carry 40% of the queued load — so a domain holding more load than its neighbour can still be underloaded, and steal from it.
Some CPUs are interrupted by things the scheduler does not control — IRQ, RT, deadline, steal. Vulnerable tasks are queued away from them; steady CPUs may still drain the turbulent queue, so the boundary stays soft.
Turbulent DSQs: David Dai, March 2026.
Always taking the idle core maximises work conservation — and on a memory-bound service it thrashes the TLB. Migrating away from hot per-thread caches cost about 3% of IPC. So sometimes it is cheaper to wait.
Task warmth: David Dai, July 2026.
At 20% system load, the obvious placement is 20% on each of 16 CPUs. What actually happens:
Low performance and high power. The worst of both.
Compaction inverts it: fewer CPUs, driven properly, and let the rest sleep deeply.
Identical total work, 8 CPUs, ~25% load. Every flicker on the top half is a C-state entry and exit. The bottom half pays it a handful of times instead of continuously.
Fewer cores, driven properly, beat every core running half-idle. The only question left is how few will carry the load.
The signal LAVD measures — who is on the critical path — was never gaming-specific. Anything where tasks wait on each other has the same shape.
SteamOS and handhelds: latency-critical, communication-heavy, and running on a battery.
Meta have explored scx_lavd as a default fleet scheduler — for the range of hardware and use cases where they do not need a specialised one.
reported via PhoronixInference servers like vLLM and SGLang live or die on keeping the GPU busy. We are looking at whether the target should be GPU idle time, not CPU latency alone.
Two ways in, and both continue what you have already seen: categorise the workload, because the problem definition follows from what it is — then migrate selectively, per workload. Warmth already decides whether to migrate and turbulent DSQs decide where; next is deciding for whom.
Read the source — main.bpf.c opens with a 180-line design document
written by the person who built it.
github.com/sched-ext/scx
Documentation/scheduler/sched-ext.rst
Changwoo Min · Igalia
sponsored by Valve