Skip to content

parallel-checkout: default to four workers on Windows - #6469

Open
azchohfi wants to merge 1 commit into
git-for-windows:mainfrom
azchohfi:parallel-checkout-windows-defaults
Open

azchohfi wants to merge 1 commit into
git-for-windows:mainfrom
azchohfi:parallel-checkout-windows-defaults

Conversation

@azchohfi

@azchohfi azchohfi commented Oct 8, 2026

Copy link
Copy Markdown

Git checks out files one at a time unless checkout.workers says otherwise. On Windows,
replacing a file in the worktree (unlink, create, write, close) costs about 250 microseconds
even on a Dev Drive (ReFS), and several times that on NTFS with real-time antivirus scanning.
That cost overlaps well across processes, so parallel checkout makes a 2,000-file checkout
28-38% faster (see below). Since 373d334 (entry: flush fscache after creating directories and
writing files, #6250) and d7b8fab (parallel-checkout: limit worker count to what poll() can
wait on, #6395), parallel checkout works with the FSCache and with any worker count.

Starting a worker is not free on Windows either, so the upstream threshold of 100 files is too
low. With four workers and that threshold, switching between branches that differ in 100 files
took 33ms longer than the 138ms of a sequential checkout.

Default to four workers on Windows, and use them only for checkouts of at least 500 files. Both
settings keep working as before; this only changes what Git does when they are not set.

Switching between branches of a 40,000-file repository that differ in N files, Windows 11 on a
32-thread desktop with an NVMe drive. Medians; time saved is the median paired difference with
its 95% bootstrap interval, over 12-15 interleaved runs:

N files Dev Drive (ReFS): before saved NTFS, Defender real-time: before saved
250 145ms (sequential) 529ms (sequential)
500 187ms 0ms [-36, 7] 671ms 123ms [95, 148]
1000 287ms 40ms [27, 66] 1024ms 282ms [219, 308]
2000 461ms 129ms [61, 180] 1668ms 633ms [603, 645]

git worktree add of all 40,000 files went from 7.07s to 4.51s (2.56s saved [2.00, 3.13]) on
the Dev Drive.

On a real repository, nodejs/node at v26.10.0 (51,849 files), with the same two values passed with
-c to 2.55.0.windows.5 (medians of 5 interleaved passes; saved, with paired 95% intervals):

NTFS, Defender real-time Dev Drive (ReFS)
git worktree add 32.2s → 12.8s (18.6s [14.4, 26.8]) 9.2s → 3.6s (5.6s [5.1, 10.4])
checkout v24.0.0 → v26.0.0 (21,361 files written) 16.9s → 10.4s (6.2s [4.7, 12.9]) 6.05s → 2.86s (3.2s [2.8, 4.9])
local git clone 53.1s → 32.7s (19.6s [18.9, 45.2]) 5.9s saved, noisy [-7.6, 9.7]
checkout v26.9.0 → v26.10.0 (781 files) 0.76s → 0.53s (0.25s [0.07, 0.37]) no measurable change
checkout v25.8.0 → v25.8.1 (195 files, below the threshold) no change no change

Below the threshold the files are still written sequentially, and those checkouts were not
slower (250 files: 10ms faster [4, 16] on the Dev Drive, 24ms faster [5, 63] on NTFS). Eight
workers were no better than four at 2,000 files, and checkout.workers=0 (one per logical CPU,
32 here) was no faster than sequential checkout.

The defaults are compile-time macros, overridable per platform, in the same way as
FSYNC_METHOD_DEFAULT and PROTECT_NTFS_DEFAULT, so other platforms are unchanged. A compiled
default covers every Git for Windows distribution in one place. Shipping it as configuration,
the way core.fscache is shipped, would mean adding it separately to the installer, the MSI,
PortableGit and MinGit, and to any system gitconfig that a tool bundling Git writes for itself.

Add a test that checks the default worker count on both sides of the threshold.

Notes for reviewers

Git checks out files one at a time unless `checkout.workers` says
otherwise. On Windows, replacing a file in the worktree (unlink, create,
write, close) costs about 250 microseconds even on a Dev Drive (ReFS),
and several times that on NTFS with real-time antivirus scanning. That
cost overlaps well across processes, so parallel checkout makes a
2,000-file checkout 28-38% faster (see below). Since 373d334 (entry:
flush fscache after creating directories and writing files) and
d7b8fab (parallel-checkout: limit worker count to what poll() can
wait on), parallel checkout works with the FSCache and with any worker
count.

Starting a worker is not free on Windows either, so the upstream
threshold of 100 files is too low. With four workers and that
threshold, switching between branches that differ in 100 files took
33ms longer than the 138ms of a sequential checkout.

Default to four workers on Windows, and use them only for checkouts of
at least 500 files. Both settings keep working as before; this only
changes what Git does when they are not set.

Switching between branches of a 40,000-file repository that differ in
N files, Windows 11 on a 32-thread desktop with an NVMe drive. Medians;
time saved is the median paired difference with its 95% bootstrap
interval, over 12-15 interleaved runs:

                   Dev Drive (ReFS)          NTFS, Defender real-time
  N files       before   saved               before   saved
    250          145ms   (sequential)         529ms   (sequential)
    500          187ms     0ms [-36, 7]       671ms   123ms [95, 148]
   1000          287ms    40ms [27, 66]      1024ms   282ms [219, 308]
   2000          461ms   129ms [61, 180]     1668ms   633ms [603, 645]

`git worktree add` of all 40,000 files went from 7.07s to 4.51s
(2.56s saved [2.00, 3.13]) on the Dev Drive.

Below the threshold the files are still written sequentially, and
those checkouts were not slower (250 files: 10ms faster [4, 16] on the
Dev Drive, 24ms faster [5, 63] on NTFS). Eight workers were no better
than four at 2,000 files, and `checkout.workers=0` (one per logical
CPU, 32 here) was no faster than sequential checkout.

The defaults are compile-time macros, overridable per platform, in the
same way as FSYNC_METHOD_DEFAULT and PROTECT_NTFS_DEFAULT, so other
platforms are unchanged. A compiled default covers every Git for
Windows distribution in one place. Shipping it as configuration, the
way core.fscache is shipped, would mean adding it separately to the
installer, the MSI, PortableGit and MinGit, and to any system gitconfig
that a tool bundling Git writes for itself.

Add a test that checks the default worker count on both sides of the
threshold.

Signed-off-by: Alexandre Zollinger Chohfi <alzollin@microsoft.com>
@dscho

dscho commented Oct 9, 2026

Copy link
Copy Markdown
Member

To be honest, I'm a little bit concerned about this patch because of the unintended consequences. Yes, it accelerates the checkout, but also it increases the memory requirements.

For context, when I check out the Git for Windows SDK in CI, I don't use four workers, even if the runner agents have two CPU cores available. No. I use 56 workers. Yes. Fifty-six.

In my tests this struck the best balance between performance and robustness. The checkout was even faster when I increased that number to, say, 64 workers. The problem, however, was that frequently these checkouts would run out of memory. And that is the exact problem here.

Unfortunately, due to its attribute handling, it is not possible to make Git's checkout logic multi-threaded. Ask me how I know. Therefore, the only chance we have is to make it multi-process. That, however, multiplies the memory requirements. It's somewhat eased by the fact that at least the packfiles are memory-mapped, so those memory pages can be shared between processes. In theory. I did not actually verify that they are shared, but in theory they could be. However, the packfiles are packfiles, but unpacked objects are unpacked objects, and they occupy a lot of memory. And I mean a lot.

Therefore, if you increase the default number of workers, you immediately also put a potentially enormous strain on the memory requirements of checkouts, in particular of large repositories, the very scenario with which this PR is supposed to help.

@qub1n

qub1n commented Oct 9, 2026 •

Copy link
Copy Markdown

Actually this goal of increasing performance is related with the following pull request, maybe it could solve your problem.

#6467

I was even testing parallel cloning on ReFS with multiple workers and it should not be that memory demanding because those workers mainly wait for system call of cloning to be finished and not holding entire buffer for copying. Buffer is allocated only for last unallocated block of file.

So I will retest how the implementation in this branch works with parallel worktree add in NTFS/RFS old and new implementation.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants