scx_lavd — Latency-criticality Aware Virtual Deadline
COSCUP 2026Taipeisched_ext / Linux kernel

scx_lavd

Latency-criticalityAwareVirtualDeadline

A Linux CPU scheduler written in BPF that figures out which of your threads are urgent — without anyone telling it.

Speaker: Gavin Guo · Igalia scx_lavd author: Changwoo Min · Igalia, sponsored by Valve ~12,100 lines · 904 commits · since March 2024
§1OriginIgalia, 2023

It started with strace and a hunch

In 2023 SteamOS still ran CFS, and tuning the scheduler meant a handful of debugfs knobs. Igalia asked a blunter question:

Can we build a gaming-optimised scheduler on sched_ext?
~80%of syscalls were futex

Most of what strace saw was Wine synchronisation — tasks waiting on each other.

graphics · Proton / DXVK Task worker thr[7825] futex_wait (10202) Task worker thr[7826] futex_wait (10438) Task worker thr[7827] futex_wait (10354) Task worker thr[7828] futex_wait (9404) dxvk-cs [7922] dxvk-submit [7920] futex_wait (30970) audio · Wine / PipeWire pipewire-pulse [1703] winepulse_main [7991] winepulse_timer [7993] FAudio_AudioCli [7992] poll (17091) epoll (22574) futex_wait (10145) poll (20102) futex_wait (10213) edge label = the call the woken task was blocked in, and how many times it was woken

Figure: Changwoo Min, OSS NA 2024 · Igalia VaporMark

sched_ext reached mainline in 6.12, over a year later

§2MotivationWhy a new scheduler

Average throughput is a lie your users can feel

A steady 144 FPS and 144 FPS with one 40 ms hitch per second have the same average — only one is playable.

The frames that ruin it were spent waiting in a run queue, not rendering.

tail latency 1% lows frame pacing wake-up latency

Same workload, two schedulers — kept anonymous as A and B. Measured and published by Changwoo Min, OSS NA 2024. B peaks higher and averages only a little lower; its worst 1% is where the stutter people complain about lives.

§3Historysince March 2024

Two and a half years, one path

Each segment is coloured by the act of this talk that covers it — criticality and deadline, topology, power — so this map is also the route ahead. Today: 904 commits, 42 contributors, ~12,100 lines.

§4Latency criticalityWho wakes whom

Waker and Wakee are the keys

Task graph · event-driven systems are graphs A producer wakes B in the middle wakes C consumer wait_freq · B blocks on A wake_freq · B unblocks C

Delay B and you delay everything downstream of it.

§5Latency criticalityWhat gets measured

Three signals

wait_freq

How often the task sleeps waiting on something external — I/O, a timer, a packet, a mutex.

wake_freq

How often the task wakes somebody else.

runtime

Average invariant runtime per schedule. Shorter is more critical

High wait and high wake means the task sits in the middle of a chain — the most expensive place to introduce delay.
§6Latency criticalitylat_cri.bpf.c

Turning behaviour into a number

lat_cri ∝ wait_freq × wake_freqruntime

In the code it is log-compressed, then squared: lat_cri = ( log₂(wait_ft·wake_ft) + log₂(rt_ft·w_ft) )2

A task that waits often, that wakes others often, and that runs only briefly sits in the middle of a chain — so it scores highest, and the scheduler runs it first.
§7Latency criticalityContext awareness

Urgency is contagious — Latency Inheritance.

key press propagation chain lat_cri carried forward along the chain key press you press input interrupt chain start input thread game simulation render submit display refresh irq chain end frame on screen you see

LAVD reads tasks as a connected graph, not in isolation: urgency flows forward to whoever you wake, and backward to whoever is blocking you.

Bounded inheritance forward · keeps the chain moving waker lat_cri 900 wakee lat_cri 200 backward · lifts the blocker
§8Deadlinecalc_when_to_run()

How do we deduce a task's deadline?

deadline_delta = adjusted_runtime × greedy_penaltylat_cri vdeadline = logical_clk − compete_window + delta

Everything the scheduler knows collapses into one number — and lat_cri divides, so ten times more critical means a tenth of the distance to wait. The clock keeps advancing, so even the least critical task eventually wins: priority that decays into fairness.

§9DeadlineHow long, not just when

How long should a task run?

slice = targeted latency 10 ms × active CPUsqueued tasksclamped to [500 µs, 5 ms]

Every queued task should get a turn within 10 ms, across all the active CPUs — so the slice is that budget divided by how many are waiting. The clamps stop it becoming either a context-switch storm or a monopoly.

5 ms 500 µs slice boosted · up to 500 ms 10 ms × cpus ÷ n 1 runnable tasks 40 clamped floor
§10Deadlinepreempt.bpf.c

Preemption mechanism

  • Two random probes, never a full scan — a random start in a random direction stops every CPU converging on the same victim: the thundering herd.
  • Only inside the task's own compute domain — the CPUs of p's domain that p may run on.
  • A candidate must be losing on both counts — lower lat_cri and a later predicted finish, or you evict a task that was about to release the CPU anyway.
  • No IPI. The victim is not interrupted — its time slice is shortened so it yields at its next scheduling point.
  • Never a lock holder, and only the top 1.56% of tasks may preempt at all.
§11Deadlinelock.bpf.c

Who holds the lock?

priority inversion, and the fix without — priority inversion with lavd — holder protected and boosted critical task lock holder blocked on the lock not scheduled nothing marks it urgent, so it waits needs the lock latency — hostage to an unscheduled task protected · boosted · releases sooner latency — the holder released first time →

A latency-critical task is only as fast as whoever holds the lock it wants — so LAVD traces futex operations to find the holder, then protects and boosts it.

§12Topologyops.enqueue() → ops.dispatch()

Dispatch Queue Layout

Each compute domain is one L3 cache, so a domain's queue holds tasks that already share a cache. But no CPU should sit idle while another domain has work waiting — so a CPU that runs dry may pull from beyond its own domain.

§13Topologystealer · stealee

Load balancer — stealer vs. stealee

LAVD is a domain-DSQ scheduler: every domain has one dispatch queue, shared by the CPUs inside it.

Stealer

An underloaded domain. It periodically pulls tasks from stealees to keep its own CPUs busy.

Stealee

An overloaded domain, carrying more queued load than its fair share. Stealers may pull from it.

Figure adapted from “Evolving sched_ext: Resource Control, Topology Awareness, and Energy Efficiency for Modern Systems” — Changwoo Min and Gavin Guo, April 2026.

§14Topologyload_invr

How to calculate the queue load?

  • Task count says nothing about work. A domain with two heavy tasks is correctly recognised as busier than one with ten small ones.
  • invr is invariant runtime — scaled by CPU capacity and frequency, so loads are comparable across heterogeneous cores.
load = queued_load_invr + util_invr queued_load_invr = Σ size_of(task), over everything queued in the domain util_invr = how busy the domain's CPUs are right now
One domain's load, bottom to top task size C task size B task size A util_invr queued_load_invr load of this domain heavier
§15Topologybalance.bpf.c

How much load should a domain carry?

fair_share = total_queued_load_invr × domain_capacitytotal_capacity

The target is not an average. A domain with 40% of the machine's capacity should carry 40% of the queued load — so a domain holding more load than its neighbour can still be underloaded, and steal from it.

§16Topologyturbulent DSQ

Which queue does a task belong in?

Some CPUs are interrupted by things the scheduler does not control — IRQ, RT, deadline, steal. Vulnerable tasks are queued away from them; steady CPUs may still drain the turbulent queue, so the boundary stays soft.

Turbulent DSQs: David Dai, March 2026.

§17PowerCPU stickiness

Should a task wait for its old CPU?

Always taking the idle core maximises work conservation — and on a memory-bound service it thrashes the TLB. Migrating away from hot per-thread caches cost about 3% of IPC. So sometimes it is cheaper to wait.

Task warmth: David Dai, July 2026.

§18PowerThe counter-intuitive one

Spreading work across all cores can be the worst thing to do

At 20% system load, the obvious placement is 20% on each of 16 CPUs. What actually happens:

  • Every core sits low, so the governor holds every core at a low frequency.
  • Every core keeps entering and leaving C-states, paying transition energy and exit latency constantly.

Low performance and high power. The worst of both.

Compaction inverts it: fewer CPUs, driven properly, and let the rest sleep deeply.

higher clocks on active cores deeper C-states elsewhere fewer P/C transitions

Identical total work, 8 CPUs, ~25% load. Every flicker on the top half is a C-state entry and exit. The bottom half pays it a handful of times instead of continuously.

§19Powercalc_nr_active_cpus()

Core Compaction: How many cores are enough?

Fewer cores, driven properly, beat every core running half-idle. The only question left is how few will carry the load.

req_cap = nr_cpus_online × avg_util_invr req_cap += 25% ← spike headroom for cpu in preference_order: sum += effective_capacity(cpu) × 50% if sum ≥ req_cap: break the CPUs walked become the ACTIVE set; the rest go to OVERFLOW
§20Beyond gamingDirection

From games, to fleets, to GPUs

The signal LAVD measures — who is on the critical path — was never gaming-specific. Anything where tasks wait on each other has the same shape.

where it started

Gaming

SteamOS and handhelds: latency-critical, communication-heavy, and running on a battery.

where it went

Server fleets

Meta have explored scx_lavd as a default fleet scheduler — for the range of hardware and use cases where they do not need a specialised one.

reported via Phoronix
the next question

AI serving

Inference servers like vLLM and SGLang live or die on keeping the GPU busy. We are looking at whether the target should be GPU idle time, not CPU latency alone.

Two ways in, and both continue what you have already seen: categorise the workload, because the problem definition follows from what it is — then migrate selectively, per workload. Warmth already decides whether to migrate and turbulent DSQs decide where; next is deciding for whom.

謝謝COSCUP 2026Questions

Thank you

Read the source — main.bpf.c opens with a 180-line design document written by the person who built it.

Project

github.com/sched-ext/scx

Kernel docs

Documentation/scheduler/sched-ext.rst

scx_lavd author

Changwoo Min · Igalia
sponsored by Valve

Presented by Gavin Guo · Igalia O index · ← → navigate · T theme scx_lavd · ~12,100 lines · 904 commits