You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
CI perf: ci-ok — waits 9–18 min for an ubuntu-latest runner while ubuntu-22.04 starts in <1 min (merge-queue critical path, backlog hours) #1372
Window: 2026-10-09 16:16–20:16 UTC. Sample: job detail for 24 merge_group CI runs and 22 pull_requestCI runs, about 4,700 started jobs.
Queue wait (job created_at → started_at) by runner label and hour:
hour (UTC)
ubuntu-latest p50 / p90 (n)
ubuntu-22.04 p50 / p90 (n)
16
0.0 / 0.4 min (941)
0.0 / 0.0 (42)
17
1.1 / 4.7 (830)
0.1 / 0.4 (88)
18
7.8 / 18.3 (531)
0.9 / 1.5 (102)
19
16.6 / 18.9 (1,358)
0.1 / 1.2 (88)
20
10.7 / 12.3 (448)
–
Windows p90 was ≤2.8 min and macOS p90 ≤5.3 min in the same hours. Only ubuntu-latest is backed up. Jobs on ubuntu-22.04 (docker-base, coverage-docker) in the same runs started within about a minute.
ci-ok (runs on ubuntu-latest, takes ~5 s) waited for a runner in the 6 successful merge_group runs of the window:
The 8 merge_group runs created 19:32–19:48 (for example 37980898001) had every needed job done by +28 to +42 min, and ci-ok was still queued at 20:16. Across the sampled PR runs, ci-ok waited up to 17.3 min.
Where the time goes
In merge_group run 37971851954 (81 min, all jobs green), the critical path crosses 4 serial ubuntu-latest queue hops:
About 62 of the 81 minutes were queue wait and about 19 were run time. The ubuntu-22.04 coverage-docker legs in the same run waited 0.5–1.5 min.
Root cause
Between 18:00 and 19:00, 92 non-draft PR runs started (see the companion issue on main-sync pushes). Each one is about 350 Linux job-min, about 172 jobs.
That backlog sits in the ubuntu-latest queue. Merge-group jobs line up behind it at every needs: hop.
The ubuntu-22.04 label is not backlogged at the same moment. So the binding constraint is the ubuntu-latest queue, not the org-wide concurrency cap, which would delay both labels alike.
ci-ok is the last hop, a 5-second Python check. Wherever it waits, the wait is pure latency for every queued PR.
Proposed fix
ci-ok → runs-on: ubuntu-22.04 in .github/workflows/ci.yml (job at ~L2555). It runs only python3 over toJSON(needs), which 22.04 has. Keep name: ci-ok unchanged.
Optional, after measuring step 1: move the other short, cache-free preflight jobs (release-readiness, dispatch-tests) to ubuntu-22.04 as well. Leave clippy and e2e-build alone: they use Rust caches keyed on the runner, and e2e-build's binary is consumed by ubuntu-latest legs.
Merge-queue critical path: −9 to −18 min per merge_group run whenever ubuntu-latest is backlogged. That was every run from 18:14 to 20:16 today.
With 5–8 groups merging together, the same minutes come off enqueue → merge for every PR in the batch.
PR push → ci-ok: same effect, up to −17 min under backlog.
Zero effect, and zero cost, when the pool is idle. Job-minutes are unchanged.
Coverage and risk
No test moves or is skipped. ci-ok keeps its name and needs: list, so the required checks ci-ok/clippy are unchanged.
Risk: the ubuntu-22.04 image is older. ci-ok uses only the system python3 (3.10 on 22.04), and the script is plain json/sys.
Risk: if many jobs move to 22.04, that label may back up too. That is why step 1 moves only the 1 job per run that sits on the critical path.
Effort
S: a one-line change.
ROI
Saving ≈ 13 critical-path min (midpoint, under backlog) × confidence 0.6 (the label-split mechanism is observed, not documented by GitHub) ÷ effort 1 = 7.8
Measurement
Window: 2026-10-09 16:16–20:16 UTC. Sample: job detail for 24 merge_group
CIruns and 22pull_requestCIruns, about 4,700 started jobs.Queue wait (job
created_at→started_at) by runner label and hour:ubuntu-latestp50 / p90 (n)ubuntu-22.04p50 / p90 (n)Windows p90 was ≤2.8 min and macOS p90 ≤5.3 min in the same hours. Only
ubuntu-latestis backed up. Jobs onubuntu-22.04(docker-base,coverage-docker) in the same runs started within about a minute.ci-ok(runs onubuntu-latest, takes ~5 s) waited for a runner in the 6 successful merge_group runs of the window:ci-okstartci-okdone at (min from run creation)The 8 merge_group runs created 19:32–19:48 (for example 37980898001) had every needed job done by +28 to +42 min, and
ci-okwas stillqueuedat 20:16. Across the sampled PR runs,ci-okwaited up to 17.3 min.Where the time goes
In merge_group run 37971851954 (81 min, all jobs green), the critical path crosses 4 serial
ubuntu-latestqueue hops:clippy,release-readiness,dispatch-tests,lint-ecosystems)e2e-build/coveragee2e (ubuntu-latest, …)legsci-okAbout 62 of the 81 minutes were queue wait and about 19 were run time. The
ubuntu-22.04coverage-docker legs in the same run waited 0.5–1.5 min.Root cause
ubuntu-latestqueue. Merge-group jobs line up behind it at everyneeds:hop.ubuntu-22.04label is not backlogged at the same moment. So the binding constraint is theubuntu-latestqueue, not the org-wide concurrency cap, which would delay both labels alike.ci-okis the last hop, a 5-second Python check. Wherever it waits, the wait is pure latency for every queued PR.Proposed fix
ci-ok→runs-on: ubuntu-22.04in.github/workflows/ci.yml(job at ~L2555). It runs onlypython3overtoJSON(needs), which 22.04 has. Keepname: ci-okunchanged.release-readiness,dispatch-tests) toubuntu-22.04as well. Leaveclippyande2e-buildalone: they use Rust caches keyed on the runner, ande2e-build's binary is consumed byubuntu-latestlegs.Expected saving
ubuntu-latestis backlogged. That was every run from 18:14 to 20:16 today.ci-ok: same effect, up to −17 min under backlog.Coverage and risk
ci-okkeeps its name andneeds:list, so the required checksci-ok/clippyare unchanged.ubuntu-22.04image is older.ci-okuses only the systempython3(3.10 on 22.04), and the script is plainjson/sys.Effort
S: a one-line change.
ROI
Saving ≈ 13 critical-path min (midpoint, under backlog) × confidence 0.6 (the label-split mechanism is observed, not documented by GitHub) ÷ effort 1 = 7.8
Generated by Claude Code