Repository navigation
CI perf (settings): merge queue builds 5 entries at a time — a 13-PR burst took 100 min to drain on a 21-min CI run (~21 min per PR past the 5th) #1248
Description
Activity
- addedci-perfCI / merge-queue performance finding (profiler routine)CI / merge-queue performance finding (profiler routine)
on Oct 9, 2026 mikolalysenko commented
on Oct 9, 2026 CollaboratorAuthorMore actions[agent] Triaged p3. This is a merge-queue ruleset change (
max_entries_to_build), which agents may not apply, so it needs a maintainer decision. Labeledagent:needs-human.
Generated by Claude Code
mikolalysenko commented
on Oct 9, 2026 CollaboratorAuthorMore actionsMore evidence from the CI janitor, 2026-10-09 18:38Z: runner capacity is starving the merge queue.
- 154 runs queued and 95 in progress (235 after dedup). 187 of them are
pull_requestruns on current heads of 80 open PRs. 28 more are Copilotdynamicruns. - The 6 merge_group
CIruns created at 18:14Z and 18:20Z (pr-1284, 1286, 1288, 1247, 1290, 1274) were all stillqueuedat 18:38Z, 18 to 24 minutes after creation, before any of their test jobs started. - The main push
CIrun for 9ab72d4 (run 37970501466, created 18:02Z) was stillqueued36 minutes later. - When an entry was dequeued, rebuilt runs waited for runners all over again. Example: Follow the republished minimist canary patch #1302 was closed at about 18:05Z, the queue rebuilt at 18:14Z, and the
Merge queue janitorsweep for the 18:02Z group itself waited for a runner until 18:17Z.
I cancelled only clearly wasted runs: 4 on superseded PR SHAs and 5 duplicate runs on the same SHA. That is negligible next to the backlog. This is a capacity problem rather than a reliability one. Options for the owner: merge_group priority over PR runs (for example, a separate runner group or larger runners for
merge_group), or limiting PR pushes per agent. #1362 (Depot Linux runners) is the open PR in this direction.
Generated by Claude Code
- 154 runs queued and 95 in progress (235 after dedup). 187 of them are
mikolalysenko commented
on Oct 9, 2026 CollaboratorAuthorMore actionsCI janitor evidence, 2026-10-09 ~19:40 UTC:
- 121 queued + 90 in-progress Actions runs. The oldest queued PR run (CI on
agent/fix-vendored-reresolved-version-drift) was created 18:23, so it has waited about 75 min. - Merge-group wall time under this backlog: the four groups created 18:14 (Skip the vlt compat matrix on ci.yml-only changes #1284, Retry CI toolchain installs past rustup download blips #1286, Wire vendored nuget.config through formats::nuget (#594) #1288, Fix bun.lockb shared bundled pin being unmanageable (#1243) #1247) finished between 19:31 and 19:35, about 77–81 min each. The two groups created 18:20 (Fix gem checks judging unused system gem homes (#1098, #1109) #1290, Fix yarn classic empty-range lock keys (#1271) #1274) finished at 19:38, about 78 min each. Every CI job in those groups succeeded. The time was spent waiting for runners, not running tests.
- Three groups (Fix scan human/JSON exit and prune forks (#1062) #1298, Fix hosted requirements rewrite of user direct refs (#542) #1333, Document vendor --force as variant-probe bypass (#923) #1346) were created between 19:32 and 19:37 and are still all queued.
- The only merge_group CI failures in the last 24h were one semantic conflict (15:20, clippy compile error on Fix bun.lockb shared bundled pin being unmanageable (#1243) #1247's group, which cascaded to six entries) and one Windows gem e2e flake on 2026-10-08 (since fixed by Fix Windows flake in gem global-gemfile refusal e2e #1169). The queue is not losing entries to flakes. It is starved of runners because PR pushes compete with it.
I cancelled 3 runs to relieve the backlog: 2 on a superseded PR head and 1 on a draft PR.
Generated by Claude Code
- 121 queued + 90 in-progress Actions runs. The oldest queued PR run (CI on
mikolalysenko commented
on Oct 9, 2026 CollaboratorAuthorMore actionsCI profiler, 2026-10-09 20:16 UTC.
Settings changed since this issue was filed. Ruleset 24668460 now has
max_entries_to_build8 (was 5) andcheck_response_timeout_minutes90 (was 120).max_entries_to_mergeis still 5. The new values can't be judged on today's data:- Last enqueue → merge:
- 111–112 min for Skip the vlt compat matrix on ci.yml-only changes #1284, Retry CI toolchain installs past rustup download blips #1286, Wire vendored nuget.config through formats::nuget (#594) #1288 and Fix bun.lockb shared bundled pin being unmanageable (#1243) #1247 (enqueued 17:43–17:44, merged 19:35).
- 78 min for Fix yarn classic empty-range lock keys (#1271) #1274 and Fix gem checks judging unused system gem homes (#1098, #1109) #1290 (enqueued 18:20).
- Earlier today it was 76 / 94 min p50 / p90.
- Merge-group run time: the six groups that passed took 77–81 min, against ~22 min unloaded. About 62 of those minutes were queue wait for
ubuntu-latestrunners, not test time.ubuntu-22.04jobs in the same runs waited under 1.5 min. - Fixes in flight: CI perf: ci-ok — waits 9–18 min for an
ubuntu-latestrunner whileubuntu-22.04starts in <1 min (merge-queue critical path, backlog hours) #1372 (ci-okwaits 9–18 min for a runner, S fix) and CI perf: PR CI — 23% of full PR runs re-test unchanged diffs after a merge-main push (~20–30k Linux job-min/day, feeds the ubuntu-latest backlog) #1373 (merge-main pushes drive 23% of full PR runs). Depot runners (Run Linux CI jobs on Depot runners #1362) would remove the backlog itself.
Building 8 entries at once adds about 3 × 310 Linux job-min to the same backlog in a burst. Worth re-checking once #1362 or #1373 lands.
Generated by Claude Code
- Last enqueue → merge:
Filed by the CI profiler routine (observer only). This is a ruleset change, so a human has to decide; the implementer routine can't apply it.
Measurement
Window: 2026-10-08 08:17 → 2026-10-09 08:17 UTC. Sources: ruleset 24668460, 157
CImerge_group runs, job detail for 52 of them, and REST timelines of the 14 most recently merged PRs.Current merge_queue rule:
max_entries_to_build: 5,max_entries_to_merge: 5,min_entries_to_merge: 1,grouping_strategy: ALLGREEN,check_response_timeout_minutes: 120.CIsuccess, created → done, p50 / p90What happened in the 06:19–08:01 burst. 13 PRs were enqueued between 06:19 and 06:52: #1215, #1209, #1208, #1217, #1193, #1218, #1036, #1041, #1226, #1230, #1229, #1228, plus #1007. They were tested in waves of at most 5
CIruns, and each wave started only after the previous one finished:test (windows-latest, 1)failed in 37892957346 and evicted #1215; the other 4 were rebuilds (37892958452 etc.)coverage-docker (sbt)failure in 37894279963)So the queue's throughput is 5 PRs per ~21 min (≈14/h). The 6th–13th PRs of a burst wait one or two full waves before their own run even starts. Of the 64-min p50, about 21 min is the PR's own run and ~40 min is waiting behind earlier waves.
Where the time goes
test (windows-latest, 2)(Run tests 15.0 min p50 + Build 2.2),coverage-docker (sbt)→coverage-merge(CI perf: coverage-docker (sbt) — 5.6-min image rebuild warms 3 unused toolchains; merge-queue critical path (~2.5 min/merge-group run) #1225), and the Gradlee2e_redirect_gradle_build … gradle_hosted_blegs (13.4 min p50, CI perf: Gradle e2e shards unbalanced — b-p leg is the merge-queue critical path (~3 min/merge-group run) #1171) all finish at 20–22 min. Queue wait for those jobs is ≤0.1 min (Linux/Windows) and 0.2 min p50 (macOS), so the waves are not runner-starved. The limit is the build depth.Root cause
max_entries_to_build: 5caps the speculative pipeline at 5 entries. With ALLGREEN groups of up to 5 and every wave taking about the same ~21 min, the queue advances in lock-step waves. During the agent bursts that now produce most merges, every PR past the 5th waits at least one extra full wave.Proposed fix (ruleset 24668460, merge_queue rule)
max_entries_to_build5 → 8, and keepmax_entries_to_mergeat 5 (or raise it to 8 too, so a fully green window lands in one merge).grouping_strategy, timeouts and the required checks (ci-ok,clippy) unchanged.Expected saving
Coverage and risk
Effort
S (one ruleset field, reversible in seconds).
ROI
About 7: (≈21 critical-path min − ≈7 weighted for +5k job-min) × confidence 0.5 ÷ effort 1. Same scale as the dashboard.