Linux Plumbers 2026sched_ext MC · Prague

Stickiness

Keeping tasks cache-warm —in work-conserving LAVD scheduler

slides · gavinguo.cc/lavd_sticky_lpc2026.html
Speaker: Gavin Guo · Igalia [email protected] 7 October 2026
§2Stickinessselect_cpu()

Work conservation picks the idle core

  • Work-conserving schedulers prefer running tasks on idle CPUs immediately, ensuring no processing capacity is wasted while work is waiting.
  • In lavd's select_cpu, when a task wakes, the scheduler does its best to find an idle core.
  • That ranks idle cores above cache-warm ones: more migrations, tasks turned cache-cold, L1/L2 caches and TLBs refilled again and again.
§3Stickinesswho pays

Some workloads pay for every migration

  • The work-conserving mechanism works well in many scenarios. But there are cache-sensitive workloads, such as edge routers using routing tables as in-memory KV stores.
  • A migration to a cold CPU has a global impact: the refill evicts other cache lines and pulls new ones in, making cache contention worse and pushing out the tail latency.
§4Outlinescx_lavd

Two kinds of stickiness

Intra-domain

Stay on the previous CPU, or take an idle one in the same domain?

Inter-domain

Stay in this LLC, or wake in a less loaded one?

  • queue load
  • latency time

This talk focuses on the first. If time permits, we also go through the second — inter‑LLC domain migration.

§5The problemselect_cpu()

Warm but taken, or cold but free?

T wakes. Its cache is on cpu A, which is busy running X. Cpu B is idle and holds nothing of T. Work conservation picks B every time.

T wakes its working set is on cpu A wait migrate cpu A previous CPU · busy L1 · L2 · TLB T's lines X's, evicting them running X residual ? per-CPU DSQ · queued ahead of T U V how much work? cpu B idle CPU · free now L1 · L2 · TLB Y's and Z's lines · none of T's running idle per-CPU DSQ empty compute domain · one LLC
the decision
wait for A if
wait(A) < refill(B)
wait(A)
X's residual, plus the work queued ahead of T
refill(B)
what T still has on A, fetched again — plus the per-CPU userspace state it leaves behind

Neither side is observable. Both are predictions — and while T waits, X keeps evicting its lines.

§6Existing approachesintra-domain

Three existing approaches to stickiness

SIS_CACHEChen Yu · Intel

keeps A idle for T's average sleep

→ slide 9

task warmthDavid Dai · Meta

waits for a warm A if the wait is short

→ slides 7–8

+ enhanced bycompletion timeChangwoo Min · Igalia

prices that wait in time, not tasks

→ slides 12–13

look‑aheadGavin Guo · Igalia

pulls a still-warm task from behind the head

→ slides 10–11

§7Task warmthDavid Dai · Meta · Jul 2026

Should a task wait for its old CPU?

Stay if the previous CPU is idle — or if the predicted wait, time until it stops + nr_queued × avg slice, fits a budget the task's heat stretches up to 2×. Otherwise migrate, and refill.

scx_lavd: Modulate CPU Stickness using Task Warmth — David Dai, Meta, July 2026 · scx #3714

§8Task warmth · resultsmemcached · memtier

Observation: the --warm-cpu-us curve

Wake-time warmth stick. 0 = lavd base. The win plateaus at 10–20 µs; 50 µs decays it.

light load (~25%)moderate (~55%)
p99 vs --warm-cpu-us 3 4 5 6 0 10 20 50 --warm-cpu-us (µs; 0 = off) p99 (ms) light load · --warm-cpu-us 0 · p99 4.02 ms light load · --warm-cpu-us 10 · p99 2.82 ms light load · --warm-cpu-us 20 · p99 3.07 ms light load · --warm-cpu-us 50 · p99 3.15 ms 4.02 2.82 3.07 3.15 moderate load · --warm-cpu-us 0 · p99 6.19 ms moderate load · --warm-cpu-us 10 · p99 5.09 ms moderate load · --warm-cpu-us 20 · p99 4.93 ms moderate load · --warm-cpu-us 50 · p99 5.36 ms 6.19 5.09 4.93 5.36 p99.9 vs --warm-cpu-us 4 6 8 10 0 10 20 50 --warm-cpu-us (µs; 0 = off) p99.9 (ms) light load · --warm-cpu-us 0 · p99.9 5.12 ms light load · --warm-cpu-us 10 · p99.9 3.93 ms light load · --warm-cpu-us 20 · p99.9 4.15 ms light load · --warm-cpu-us 50 · p99.9 4.25 ms 5.12 3.93 4.15 4.25 moderate load · --warm-cpu-us 0 · p99.9 9.94 ms moderate load · --warm-cpu-us 10 · p99.9 7.69 ms moderate load · --warm-cpu-us 20 · p99.9 7.57 ms moderate load · --warm-cpu-us 50 · p99.9 8.01 ms 9.94 7.69 7.57 8.01
§9SIS_CACHEChen Yu · Intel · Nov 2023

Keep the warm CPU idle for one short sleep

When a short sleeper leaves its CPU idle, the CPU stays cache-hot for the task's average sleep. Another wakee's idle search skips it and takes a cold CPU; the sleeper finds its previous CPU idle and takes it at once. A timer, not a lock — no owner is stored.

[PATCH v2 0/3] Introduce SIS_CACHE to choose previous CPU during task wakeup — Chen Yu, Intel, Nov 2023 · not merged

§10Look-ahead at dispatchGavin Guo · Igalia · Aug 2026

Look ahead for warm tasks in the domain queue

When a CPU drains the shared domain queue, it looks ahead a few tasks past the head and tries to find one still warm on this CPU, sequentially, in deadline-increasing order; when a warm task is found, it pulls it to the per-CPU queue and keeps every other task in the shared queue for any CPU.

scx_lavd: Add second-pass warm-task pull at domain DSQ dispatch — Gavin Guo, Igalia, Aug 2026 · not upstream

§11Warmth + look-ahead · results

Latency vs load — all configs

lavd basewarm+d2warm+d4
p99 latency vs load 2 4 6 8 10 12 light ~25% moderate ~55% heavy ~96% overcommit 4× p99 (ms) lavd base · light load · p99 4.02 ms lavd base · moderate load · p99 6.19 ms lavd base · heavy load · p99 11.35 ms lavd base · overcommit load · p99 12.17 ms warm+d2 · light load · p99 2.42 ms warm+d2 · moderate load · p99 4.25 ms warm+d2 · heavy load · p99 11.42 ms warm+d2 · overcommit load · p99 12.77 ms warm+d4 · light load · p99 2.58 ms warm+d4 · moderate load · p99 3.77 ms warm+d4 · heavy load · p99 11.37 ms warm+d4 · overcommit load · p99 12.94 ms p99.9 latency vs load 0 10 20 30 40 50 light ~25% moderate ~55% heavy ~96% overcommit 4× p99.9 (ms) lavd base · light load · p99.9 5.12 ms lavd base · moderate load · p99.9 9.94 ms lavd base · heavy load · p99.9 32.15 ms lavd base · overcommit load · p99.9 49.56 ms warm+d2 · light load · p99.9 3.72 ms warm+d2 · moderate load · p99.9 7.35 ms warm+d2 · heavy load · p99.9 33.59 ms warm+d2 · overcommit load · p99.9 49.71 ms warm+d4 · light load · p99.9 3.79 ms warm+d4 · moderate load · p99.9 7.15 ms warm+d4 · heavy load · p99.9 33.56 ms warm+d4 · overcommit load · p99.9 48.90 ms p50 latency vs load — the trade 0.10 0.15 0.20 0.25 light ~25% moderate ~55% heavy ~96% overcommit 4× p50 (ms) lavd base · light load · p50 0.077 ms lavd base · moderate load · p50 0.074 ms lavd base · heavy load · p50 0.079 ms lavd base · overcommit load · p50 0.079 ms warm+d2 · light load · p50 0.257 ms warm+d2 · moderate load · p50 0.130 ms warm+d2 · heavy load · p50 0.079 ms warm+d2 · overcommit load · p50 0.077 ms warm+d4 · light load · p50 0.220 ms warm+d4 · moderate load · p50 0.201 ms warm+d4 · heavy load · p50 0.081 ms warm+d4 · overcommit load · p50 0.079 ms
§12The waitscx_lavd PR #3806 · open

Counting the queue is not timing it

Changwoo Min's completion-time estimate (#3806) prices the wait in time, not tasks: the running task's remaining time plus each queued task's own service time, over the core's capacity. For task warmth, T now waits for its warm CPU behind short work and moves on when long work is really ahead.

today5699c7f08582

wait = (est_stopping − now) + nr_queued × slice_wall

per-CPU DSQ only · a count times an average

▼ Changwoo's enhancement: priced in time, not tasks

completion time enhancement80eabdc60198

wait = residual(running) + queued service ÷ capacity

local + per-CPU DSQ · each task's own service time

queued ahead of T nr_queued × slice queued service 20 µs 20 µs 10 ms 40 µs 4 ms 4 ms 10 ms 8 ms slice_wall = 5 ms, the default slice max, below saturation
§13The waitChangwoo Min · Igalia · PR #3806

The count is wrong in both directions

Completion time prices each queued task at its own service time, reads the local DSQ, and charges a running task that has outrun its average for the overrun.

Completion time: Changwoo Min, PR #3806 (open) · stickiness change 80eabdc60198, not yet split out

§14Discussionopen problems

Questions for the room

  • Warmth estimation
    • How long does L1/L2/TLB state really survive on a busy CPU?
    • Could the kernel expose a residency signal — PMU counters, cache occupancy?
    • Intel RDT and Arm MPAM count cache occupancy — per task?
    • Cheap enough to read at scheduling-decision frequency?
  • Wait or migrate?
    • What should a scheduler maintain to predict the wait honestly?
    • What is a migration worth, in time, for this task?
  • Scheduler ↔ userspace
    • Allocator per-CPU caches: the real victims of migration
    • Should they hint their locality needs to the scheduler?
    • Or should the scheduler expose warmth and migration state to them?
§15LPC 2026Questions

Questions?

project
github.com/sched-ext/scx
scheduler
scheds/rust/scx_lavd
speaker
Gavin Guo · Igalia
Join us

Igalia is a worker-owned cooperative — remote-friendly and consensus-run. We are hiring.

igalia.com/jobs
Linux Plumbers Conference 2026 · sched_ext MC