The GitHub Actions workflow has two microVM test layers on Azure-hosted
self-hosted KVM, MSHV, and WHP virtual machines. Each backend has a pool of
runners labeled by operating system, backend, and virtual-machine; see
Job graph and runner capacity for their
sizes. Jobs target the shared backend labels so any available matching runner
can execute them. This allows the backend lanes to execute concurrently without
binding a workload to a specific host. openvmm-vmm-tests downloads the NVX
guest artifacts and uses the Linux-direct kernel and Alpine initramfs to
exercise OpenVMM's Linux MP-table lifecycle, TTRPC, and snapshot contracts on
each backend. openvmm-unit-tests runs the OpenVMM unit and documentation
tests once per operating system, on an MSHV runner and a WHP runner. The tests
that need /dev/kvm or /dev/mshv are ignored, so the KVM and MSHV runners
used to run the same 4,820 tests, and the MSHV pool has the most runners.
Failed openvmm-vmm-tests jobs upload Petri's test_results directory,
including guest and VMM logs, screenshots, and watchdog inspection data.
These seven-day artifacts are named
openvmm-vmm-tests-<os>-<backend>-<run-id>-<run-attempt>, so a successful rerun
does not replace the failed attempt's diagnostics. Linux collects them from
openvmm/target/vmm_tests/test_results; Windows uses
<runner-temp>/<backend>/test_results.
The nvx-microvm-tests-{kvm,mshv,whp} jobs consume the NVX Linux kernel and
the NVX Linux kernel plus the selected Alpine or Ubuntu initramfs and exercises
Linux, SMP, virtio, sandbox, and snapshot behavior through the public OpenVMM
CLI. Alpine-control-only scenarios remain explicit and are rejected for the
Ubuntu initramfs. Each job also boots the Azure Linux initramfs through its
one-vCPU smoke set, under the same time ABI checks; the debug-kernel jobs skip
it, as they skip the Ubuntu tests. Failure logs from the NVX layer are uploaded
per backend.
Every harness launch, in the tests and the benchmarks, scans the OpenVMM
console for the guest's time ABI output, and the tests
scan every virtio console they read over TCP the same way, to its end. An
NVX-TIME-ABI-VIOLATION event or a failed NVX-TIME-ABI conformance line
fails the scenario at once with the guest's code and detail. Guests keep the
console quiet, because every console byte costs a port exit. Only the initial
clock step precedes the guest's boot marker; the other boot checks finish
afterwards and still power off with status 193 on failure, possibly in the
middle of a scenario. A guest prints one NVX-TIME-ABI line per recorded
check phase (boot, then capture and restore after a restore) and a runtime
line only when /sbin/nvx-time status asks, after waiting up to 30 s for
pending checks. The test runners ask after every cold boot whose shell is on
a console the harness reads, the OpenVMM console or a virtio console over
TCP, before any other input, and wait for the query to exit, so the console's
echo of later input cannot split its lines. Cold boots with no shell there are
only scanned: one-shot workloads, and the managed lifecycle, where init starts
the managed agent instead of a shell. A failing boot check still powers them
off with status 193, which fails the run. The query must
exit 0 with a passing NVX-TIME-ABI boot line, which reports ABI version 1,
generation 0, a plausible TSC rate, and the backend's LAPIC rate; a missing
line means the guest image or OpenVMM does not implement the ABI, a check
still pending after the guest's 30 s wait fails, and a report-only guest
(NVX-TIME-REPORT) is never accepted. A guest image whose nvx-time has no
status subcommand fails with that reason. The runtime line only records
the wall-clock discipline's state. Benchmarks never ask, so their measured
intervals stay quiet. An OpenVMM exit
status of 193, 194, or 195 is reported as the guest's time ABI conformance,
runtime-violation, or restore-repair power-off, together with the event that
preceded it, instead of as a generic exit status.
The smp, smp-snapshot, and restore-processors scenarios run the guest
warp probe, /sbin/nvx-time-probe warp --bound-ns 1000, over every pair of
online CPUs after each boot or restore, in two rounds with all vCPUs halted
for 1 s between them, so that a host without an invariant TSC corrects the
guest TSC as idle host CPUs wake (#265). Every pair in every round must stay
within the time ABI's 1 µs
cross-vCPU skew bound for both the
backward TSC step and the ping-pong offset; a stalled pair or an inconclusive
measurement also fails. smp boots each requested processor
count. smp-snapshot captures one snapshot per requested count and restores
it once, and the first snapshot a second time to prove that a restore leaves it
reusable. restore-processors captures one boot-online CPU with capacity 8 and
probes the CPUs that each restore target activated. After the 1/2/4/8-CPU
restores, it restores the same snapshot once without --restore-processors.
After the warp probe, every smp-snapshot and restore-processors restore
runs /sbin/nvx-time status, which waits for the restore's deferred checks;
it must exit 0 with a passing NVX-TIME-ABI ... phase=restore line that
reports the backend's LAPIC rate, the restored CPU count, and generation 1:
every scenario captures a cold-booted guest, and OpenVMM cannot capture a
restored one. The query also prints the source's boot and capture lines,
which keep the CPU count their checks covered, so only the restore line must
match the restored CPU count.
Every restore runs with OpenVMM lifecycle profiling and must report exactly one
startup.vp_thread_bind record. Its startup.vp_bind_* records must show that
an explicit MSHV target binds exactly VPs 0..N-1, while untargeted MSHV
restores and all KVM and WHP restores bind the full capacity.
Captures do not wait for a clocksource: the time ABI registers tsc at
device_initcall on every backend, so the transitional tsc-early window that
once let the clocksource watchdog compare tsc-early with jiffies across a
restore (#253) never reaches the guest's userspace.
A restore fails as soon as its guest prints NVX-RESTORE-PROCESSORS-FAIL or
NVX-WARP-PROBE-FAIL, rather than waiting for the phase timeout; a failed
probe also powers the guest off with status 97. Restore-processor logs record
OpenVMM's time ABI rates declared event, which reports the identity MSR
route, the TSC synchronization method, the native and declared TSC rates, and
the rate deviation from the snapshot.
The restore-downtime scenario covers the time ABI's long-downtime case. It
captures four snapshots, at 1 and 8 vCPUs, each with and without
rcupdate.rcu_expedited=1, then waits 30 s, longer than the guest's 21 s RCU
stall timeout, and restores each one. Every restored guest must run the warp
probe, report a passing restore line through nvx-time status before the
harness stages its check, report /sys/kernel/rcu_stall_count as 0 two
seconds later, and show an uptime of at least 30 s, which proves that
monotonic time advanced by the downtime. The captures share one downtime
window, so the scenario adds about a minute per backend.
The time-abi-conformance scenario boots the largest requested vCPU count and
runs the guest's exhaustive CI check, /sbin/nvx-time exhaustive, which the
boot check leaves to CI. On every online CPU it checks every leaf
0x40000006..=0x400000ff and every base 0x40000100..=0x4000ff00 for another
hypervisor signature, every C3 MSR, the write rules of
HV_X64_MSR_TSC_INVARIANT_CONTROL, writes to the read-only identity MSRs, and
reads of IA32_TSC_ADJUST and IA32_TSC_DEADLINE (checks X1 to X6). The
harness requires a passing NVX-TIME-ABI-EXHAUSTIVE line for every check on
every CPU, a summary with status=ok, the requested CPU count, and no
failures, and exit status 0; a failure lists each failing check with the
guest's detail. The guest command fits on one console line, so the console's
echo of it ends before the check prints.
The cpu-profile-fallback scenario covers --cpu-profile auto's fallback to a
host profile on a CPU that no built-in profile serves, on runners that
built-in profiles serve. It cold boots one vCPU with OpenVMM's
host-cpu-unknown test hook, which makes auto treat the host CPU as
unknown, and requires exactly one NVX-CPU-PROFILE-FALLBACK: warning that
names a host profile with its digest and links the CPU profile request form,
a cold boot that declares that profile, and passing boot checks. It then
captures the guest and restores it with auto and no hook: the restore must
take the snapshot's host profile without another warning, and pass the warp
probe and its restore checks.
The snapshot-core scenario first sends a snapshot request to a guest that
OpenVMM launched without a snapshot destination. OpenVMM releases the request,
so it returns in the source, which must continue exactly once. Before the
request, the guest's snapshot agent saved and overrode the stall detectors'
settings, and it must restore them when the request returns
(nvx-time cancel-capture). Afterwards, nvx-time status must exit 0 with a
passing boot line at generation 0 and no restore line, and
/sys/module/rcupdate/parameters/rcu_cpu_stall_suppress must read 0.
At the end of a passing run, test-microvm prints one NVX-TIME-ABI-EVIDENCE:
line and adds it to the GitHub job summary, because CI keeps the guest logs only
for failed jobs: the number of warp probe runs with their worst
max_abs_offset_ns and max_backward_ns, the count and elapsed_us range of
the newest check that each nvx-time status query reported (boot after a cold
boot, restore after a restore, and capture after snapshot-core's released
request), the exhaustive check's summary, and the
restore-downtime stall counts. Where the guest reports a check's CPU time
(cpu_us), the line also gives its range and how many checks exceed the
backend's CPU-time budget
for their phase, which the spec sets from guest measurements. The spec exempts
the first capture after a rolled-back capture in the same VM process; CI never
makes one, because each VM process sends OpenVMM at most one snapshot request
and a failed capture fails its scenario.
CHECK_CPU_BUDGET_US in scripts/nvx_tools/time_abi.py holds the budgets, one
per backend and phase. elapsed_us is wall time, including waits behind the
workload, and has no budget. CI gates on neither: the performance gate is the
A/B comparison outside CI.
The nvx-microvm-debug-{kvm,mshv,whp} jobs run test-microvm --debug-kernel
on the CI debug kernel (build/vmlinux-debug, built from
kernel/config-microvm-debug), whose soft-lockup and hung-task detectors
production kernels leave out. It selects the same-host restore scenarios
smp, smp-snapshot, restore-processors, restore-downtime, and
snapshot-tiers. The jobs skip every other test step, because those boot the
production kernel, which the nvx-microvm-tests-{kvm,mshv,whp} jobs already
cover: the public managed execution configuration, the Ubuntu and Azure Linux
guest tests, the Ubuntu sandbox layer and live-share smoke tests, the
caller-owned live-share test, and the aci_edge_sandboxes lifecycle test. Any
RCU stall, soft lockup, or hung task makes the guest's time ABI watcher power
off with status 194, which fails the run. Every guest those scenarios restore
asks nvx-time status before it exits
(snapshot-tiers after its tier assertions), which waits for the restore's
deferred checks and must report a passing restore line, so no restore leaves
its checks pending. The one exception is snapshot-tiers' gate-timeout check,
whose guest OpenVMM stops at the restore gate on purpose. The harness refuses
a kernel whose vmlinux-debug.config lacks the detectors, because the guest's
C11 check passes vacuously without them. To bound the cost, pull requests
run the debug kernel on KVM only and dev pushes run it on every backend;
each job takes about five minutes on its own runner, in parallel with the
other microVM jobs. The jobs gate the required status check, the development
release, and performance persistence. The GitHub-hosted debug-kernel job
builds the debug kernel beside the shared artifacts job
(build-guest-artifacts with guest-images: "false") and caches it under its
own key, so a kernel rebuild delays only the debug jobs, and a failed debug
kernel build fails the required status check and blocks the release.
Every job that uses the validate-runner action first reports the Linux
runner's kernel, CPU, clocksource, and TSC flags. It then qualifies the runner
for the time ABI with nvx.py doctor --checks H1 H2 H4 --ci-schedule (see
Host qualification): the backend, the CPU fingerprint
and generation, and the TSC rate stability on its short schedule. The microVM
and platform jobs run it after downloading OpenVMM and add OpenVMM's CPU
profile check (H2) and preflight (H3), and keep the fingerprint as an artifact
when the profile check fails; the other jobs pass --no-openvmm. This takes a
few seconds. It reports the CPU generation, the built-in CPU profile that
auto selects, and the measured TSC rate in the log and the job summary, and
fails the job with a stable code when the runner is not qualified, for example
E_PROFILE_HOST_UNKNOWN on an unknown CPU generation, where auto would fall
back to a host profile, which CI never uses. The platform jobs also
run the guest warp probe (H6) on CI's schedule, which adds about 25 s, because
their benchmarks boot guests without the microVM scenarios' probes. The doctor
gates only on measured properties, alike on every backend: it records the host
OS's invariant-TSC flags and clocksource as evidence, and the guest warp probe
measures the skew that a host without an invariant TSC causes. Before the time
ABI, never-restored guests on such an MSHV runner VM, an Ice Lake 8370C, hit
cross-vCPU TSC warps when an idle host CPU woke (#211, #265), which is why the
probe schedules include idle gaps. Under the time ABI, its guests' warp probes
measure at most 36 ns, so validate-runner only notes a missing
nonstop_tsc, and the Linux runner setup script only warns. The OpenVMM
vmm-tests run no warp probe: they boot OpenVMM's own test guests, whose
verdicts don't depend on the skew bound. Runner labels do not encode the
generation; per-PR CI captures and restores on one runner, so generations
never mix.
The warp probe replaced the restore-tsc-sync scenario, its test-only
clearcpuid=tsc_adjust kernel option, the guest's scan of the kernel log for
TSC warp and instability messages, and the fresh-boot TSC control that
classified those failures (#211, #265). Under the time ABI the guest's TSC is
tsc_reliable, so Linux skips its CPU-online warp check and never logs the
messages that guard looked for; the probe measures the 1 µs bound on every CPU
pair instead.
The console-exit scenario delays host console reads for two seconds after
snapshot restore to exercise output backpressure. For each requested processor
count it requires byte-exact delivery of a 64 KiB payload and the final marker,
and preserves guest exit statuses 0 and 37. This checks both device and host-relay
draining without adding sleeps to the measured benchmark workloads.
The harness waits for the output reader's EOF notification even after the
process exits, so delayed final output chunks cannot create a false failure.
Shared guest artifacts are built with Docker on a GitHub-hosted Ubuntu runner.
The kernel, Alpine initramfs, Ubuntu initramfs, and Ubuntu EROFS layer use
separate cache keys. Ubuntu keys include the Canonical archive pin,
supplemental package lock, common guest sources, shared download and guest
descriptor modules, converter implementation, and Dockerfile. Artifact upload
retains the Alpine filenames and adds the distinct Ubuntu filenames. Each
backend also boots the Ubuntu initramfs and runs
/sbin/nvx-sandbox-smoke from the Ubuntu EROFS layer as UID/GID 65534 over a
fresh ext4 scratch copy, with a 32 MiB --memory-max and an 8-process
--pids-max. The probe verifies the
sandbox security profile: the
fixed identity, empty capability sets and no_new_privs, private mount, PID,
and UTS namespaces, the private device view and read-only sysfs, the
container cgroup, the overlay root over read-only image layers with its upper
directory on scratch, and that the memory controller kills an allocation of
twice the limit and the pids controller stops the workload at eight
processes. The live-share runs below repeat the profile checks. Each
backend also runs the public managed-exec-config and sandbox-lifecycle
scenarios of test-microvm, which drive nvx.py sandbox provision, start,
exec, stop, and deprovision as subprocesses (see
Run): request-scoped working directories and
environments, timeouts, nonzero exits, stdout and stderr, bounded output and
outcome reports, warm-guest state across requests and restarts, identity
refusal, fail-closed lifecycle transitions, a crashed OpenVMM's stale runtime
state, and teardown that leaves no OpenVMM process, control endpoint,
capability, or runtime record while the layer and scratch artifacts survive.
Their bounded evidence is part of the backend's failure logs. The same
entrypoint then verifies a live virtio-fs
share inside the container: a read-write /workspace share with a denied
subdirectory must round-trip guest writes to the host, and a read-only
/opt/hostedtoolcache share must reject writes and symbolic links. In the
read-write share, the guest also exercises the symbolic-link primitives that
package managers use: it creates an npm-style relative .bin link to an
executable and runs it through the link, renames, replaces, and removes links,
and creates absolute, dangling, outside-the-share, and denied-path links. The
host then checks that every target is preserved exactly (on Windows, by
decoding each WSL-style reparse point, which Windows does not follow) and that
no link read or modified host data outside the share or in the denied path. A
managed sandbox then repeats
the read-write check through provision, start, exec, and stop, and must
report a successful outcome with a cleanly unmounted scratch filesystem, which
shows that stop unmounted the share and overlay first. Each backend then
attaches a read-write /workspace share and a read-only
/opt/hostedtoolcache share to one sandbox at the same time, each with its
own denied subdirectory. The smoke script runs the read-write and read-only
checks on the two shares and verifies that a link in the read-write share
cannot write into the read-only share; the host requires the guest's writes in
the workspace, an unchanged tool cache, and unchanged denied files, both for a
one-shot run and for a managed sandbox whose format-4 configuration must
persist both shares and whose stop must leave a cleanly unmounted scratch
filesystem. Each backend then attaches a read-write /workspace share whose
access policy denies nvx-denied, allows
nvx-denied/nvx-allowed, and makes only nvx-writable writable, with every
host directory and file writable by every identity so that only OpenVMM can
refuse a write. The smoke script's policy mode requires the workload's
writes to succeed only in nvx-writable and to fail with EROFS everywhere
else, a hard link of a read-only file into nvx-writable to fail, and
nvx-denied to list only nvx-allowed, whose file stays readable while
nvx-denied/secret stays hidden. The host requires the guest's writes in
nvx-writable and no other change to the share, both for a one-shot run and
for a managed sandbox whose format-5 configuration must persist the denied,
allowed, and writable paths and whose stop must leave a cleanly unmounted
scratch filesystem. The default filesystem-policy scenario of test-microvm
covers the same policy as guest root, whose permission checks mode bits never
stop: every mutation outside the writable paths, through a second mount of
the tag as well, fails with EROFS, a hard link fails with EXDEV, a denied
path nested inside an allowed path stays hidden, invalid policies are rejected
before boot, and a snapshot restore requires the same policy and keeps a
writable handle open across the capture. On Linux, a
world-writable share then runs with --mount-owner caller. The runners grant
OpenVMM neither CAP_SETUID nor CAP_SETGID, so it can neither assume the
65534 workload identity nor drop the backend group through which it opens
/dev/kvm or /dev/mshv. The smoke script's eperm mode therefore
requires reading and listing the share to fail with EPERM and writes to
fail, and the host requires the share to be unchanged, both for a one-shot
run and for a managed sandbox whose configuration must persist the mode. The filesystem-owner scenario of
test-microvm likewise requires guest root and a foreign guest identity to
fail with EPERM on the runners, and a root-owned export to be rejected
before boot; on Windows, --mount-owner caller must be rejected before boot.
The OpenVMM unit tests check the host identity that each request runs as,
including guest root squashed to an export owner other than OpenVMM. Where
OpenVMM can drop its supplementary groups, the scenario instead requires
files that guest root creates to belong to the share owner rather than to
root, and squashed root to fail to chown a file, give it another of OpenVMM's
groups, or create a device node. When the test process also passes
CAP_SETUID and CAP_SETGID to OpenVMM, as on a privileged developer host,
the scenario requires the foreign identity to own the files it creates.
Linux/KVM runs the
broader Ubuntu SMP, managed lifecycle, network snapshot, blockless snapshot,
and workload-identity set.
OpenVMM release executables and provenance are built once by the independently
addressable build-openvmm-linux-gnu, build-openvmm-linux-musl, and
build-openvmm-windows-msvc producer jobs. KVM workloads and MSHV microVM tests
consume the GNU artifact, MSHV platform workloads consume the musl artifact,
and WHP workloads consume the Windows MSVC artifact. Each workload can start
after its compatible OpenVMM producer and the shared guest-artifact job finish,
without waiting for unrelated OpenVMM targets.
All three producers call the same Python build workflow, passing the validated
runner backend explicitly through build-openvmm --backend. The backend is
carried in OpenVmmBuildConfig; the build workflow maps KVM, MSHV, or WHP to
GNU, musl, or MSVC without probing runtime devices. CI therefore retains its
musl build for MSHV without maintaining a separate shell build path.
The kernel, debug kernel, and initramfs cache keys include
docker/Dockerfile, which pins the build image and its
toolchain and in which every guest artifact builds, and
build_config.py and
build_constants.py, so a toolchain,
shared build configuration, or constant change invalidates cached guest
artifacts and their provenance. The Ubuntu distro layer shares the Ubuntu input
hash.
The producer handoff uses one-day workflow artifacts rather than caches. Each
consumer downloads both the normalized executable and its build provenance,
then restores executable permissions on Linux. Once the required artifacts are
ready, benchmarks run in parallel with the NVX test layer and use any available
runner in the matching backend pool. All three use virtual-machine performance
series and the constrained eight-CPU affinity policy. Development releases and
performance baseline updates still require every applicable test and benchmark
lane to pass.
The microVM correctness jobs gate every use of benchmark results, so the
benchmark action runs no correctness scenario of its own (#286). Required status check requires each nvx-microvm-tests-* job that the change schedules
to succeed. Publish development release and Persist performance baseline run only on dev pushes in which every microVM test job succeeded or
was skipped. The pull-request Performance regression gate reads only the
platform jobs' results and publishes nothing. The counting LAPIC that the
benchmarks depend on is covered by the smp scenario of those jobs.
smp-lapic repeats smp with the counting-LAPIC facts asserted, for explicit
local use only; no CI job runs it. Benchmarks
lists the jobs that consume, gate, publish, or persist the results. The workflow uses the read-only OpenVMM deploy key stored in the
OPENVMM_DEPLOY_KEY Actions secret to fetch the private submodule at its pinned
commit. Shared guest binaries and development release packages move through
short-lived workflow artifacts alongside the OpenVMM handoff and benchmark
results. Caches only accelerate reproducible build inputs and outputs; consumers
do not depend on them as a handoff.
Pull requests gate regressions against recent matching-platform history, and
successful pushes to dev append their p50 values under data/. Metadata-only
performance jobs use GitHub-hosted Ubuntu runners. Provisioning instructions
are in the runner bootstrap guide.
Persistent runners accept pushes and same-repository pull requests only. Fork pull requests run the GitHub-hosted validation jobs but do not execute code on the Azure runner fleet. A maintainer must stage an external contribution on a trusted repository branch before running the backend matrices.
A dev push or a pull request that is ready for review starts these jobs at
once: quality, aci-edge-sandboxes, nvx-cli-linux, and nvx-cli-windows
on GitHub-hosted runners, and openvmm-changes, which classifies the change
(see Change classification). When it finishes,
artifacts and debug-kernel restore or build the guest artifacts on
GitHub-hosted runners, the three build-openvmm-* producers restore or build
OpenVMM on a KVM, an MSHV, and a WHP runner, and openvmm-unit-tests starts on
an MSHV and a WHP runner. openvmm-vmm-tests waits for artifacts. Each
microVM test and platform job waits for artifacts, its OpenVMM producer, and
the NVX CLI job of its operating system, and the debug-kernel jobs also wait
for debug-kernel. Performance regression gate, Publish development release, Persist performance baseline, and Required status check wait for
the jobs whose results they check. Required handoffs between these jobs use
same-run workflow artifacts; caches only skip rebuilding their inputs.
The NVX CLI tests do not depend on the backend, so they run once per operating
system in nvx-cli-linux and nvx-cli-windows rather than in every
hypervisor job. Each checks out OpenVMM, whose pinned CPU profiles a test
reads, and uses Python 3.12, the release that the self-hosted runners provide;
the Windows job also installs the pinned Rust release with the minimal profile
for the curl shim test. Fork pull requests lack the OpenVMM deploy key, so the
CLI jobs run, like the hypervisor jobs, only for pushes and same-repository
pull requests. Because the microVM and platform jobs of each operating system
wait for its CLI job, a broken CLI does not occupy the hypervisor runners.
Required status check, the development release, and performance persistence
require both CLI jobs.
The WHP pool has five runners, the KVM pool six, and the MSHV pool nine. Jobs wait seconds for a runner when one run is in flight, but the WHP jobs are the longest: the OpenVMM VMM and unit tests compile OpenVMM on Windows for about 15 and 13 minutes before testing. A full run occupies the WHP pool for most of its duration, so overlapping pull requests queue there first. A change that adds a job or a shard needs runners to run it in parallel; otherwise it only queues.
openvmm-changes lists the files that changed between the pull request's base
and head, or between a push's previous and new commits, and
nvx.py classify-ci-changes selects the suites from their paths:
run-workloadsis false only when every changed file is documentation, underdoc/or ending in.md. It enables the guest artifacts, the OpenVMM producers, and the platform benchmarks.run-testsis true when a test input changed: the OpenVMM pin, the guests and their package locks, the CLI and its modules, the workflows and actions that build and run the tests, the runner setup scripts, the Rust pin, or theaci_edge_sandboxescrate. It enables the NVX microVM tests.run-openvmm-unit-testsandrun-openvmm-vmm-testsenable the OpenVMM suites. A push sets both torun-tests. A pull request sets each only when one of that suite's own inputs changed. The unit tests depend on the OpenVMM pin, the Rust pin, the runner setup scripts,ci.yml, the actions that their job uses,nvx.py, and the modules thatci.pyimports. The VMM tests also boot the kernel and Alpine initramfs, so they add every input of those artifacts' cache keys and thebuild-guest-artifactsaction.
A pull request that changes neither OpenVMM suite's inputs, such as one that
changes only the microVM tests, the benchmarks, or the Ubuntu or Azure Linux
guests, therefore skips the two jobs that occupy the WHP runners longest,
while the dev push that merges it still runs them. The detector runs the
full matrix instead when the comparison base or a commit is unavailable, when
the diff or the classifier fails, or when the classifier reports a suite
twice or not at all, because a missing output would skip its jobs.
Required status check passes the same outputs to check-required-ci, which
expects every scheduled job to succeed and every other job to be skipped.
Tests in scripts/test_nvx_tools.py derive each suite's inputs from the
actions that its jobs use, the modules that ci.py imports, and the guest
cache keys, and fail when a pattern misses one.
python3 scripts/nvx.py doctor --backend <kvm|mshv|whp> qualifies a host for
the time ABI. It runs checks H1 to H7
in order, prints one NVX-DOCTOR: check=<id> status=<pass|fail> detail="..."
line per check, and exits with status 1 if any check fails. --checks selects
a subset, and --summary appends a Markdown table with the CPU generation,
profile, rates, and skew metrics to a file such as $GITHUB_STEP_SUMMARY. H4
and H6 run on the spec's long qualification schedules, which take about three
minutes; --ci-schedule selects CI's short ones. A failure that matches a time
ABI failure code starts its detail with the code in brackets, for example
[E_PROFILE_HOST_UNKNOWN]. A check fails, and never passes, when it can't
read a fact it gates on: missing or unrecognized tool output fails it, and an
error in one check fails only that check. Only evidence fields, such as the
microcode and the invariant-TSC flags, may read unknown.
| Check | Implementation |
|---|---|
| H1 | /dev/kvm or /dev/mshv is readable and writable, and a KVM host has no /dev/mshv; on Windows, WHvGetCapability reports a hypervisor |
| H2 | Vendor, family, model, stepping, microcode, and OS build from /proc/cpuinfo or the Windows registry. Then openvmm --hypervisor <backend> --cpu-fingerprint <path> writes the host's CPU fingerprint (by default nvx-cpu-fingerprint-<backend>.json in the probe directory; --cpu-fingerprint overrides it), checks it against the pinned profile that auto selects from OpenVMM's catalog, which shares one profile per generation across backends, and prints one NVX-CPU-PROFILE: line. The check never falls back to a host profile, so a CPU on which auto falls back fails H2. H2 reports the generation and the profile from that line, so a profile that an OpenVMM pin adds qualifies its hosts without an NVX change. H2 requires exit status 0, status=pass, the same backend, and a catalog profile of the reported generation, reports the profile and surface digests, and otherwise fails with OpenVMM's code, for example [E_PROFILE_HOST_UNKNOWN] on a CPU that no profile serves or [E_PROFILE_UNSUPPORTED] naming every unsupported CPUID bit. On MSHV and WHP, OpenVMM checks the entries outside the profile on a probe partition configured from the profile, as a cold boot does. With --no-openvmm, H2 maps the host with NVX's copy of the catalog instead: skylake-sp (6/85, steppings 0 to 4, intel.skylake-sp.v1), icelake-sp (6/106, intel.icelake-sp.v1), emeraldrapids (6/207, intel.emeraldrapids.v1), alderlake (6/151 and 6/154, intel.alderlake.v1), milan (AMD 25/1, amd.milan.v1), genoa (AMD 25/17, amd.genoa.v1), or turin (AMD 26/2, amd.turin.v1). Any other CPU, including Cascade Lake and Cooper Lake, fails with E_PROFILE_HOST_UNKNOWN. A unit test keeps the copy equal to the profiles of the OpenVMM submodule's pinned revision, which it reads from the gitlink's commit, so the NVX CLI jobs validate the CLI after they check out OpenVMM. The host OS's invariant-TSC flags (constant_tsc nonstop_tsc, or the CPUID bit on Windows) are recorded as evidence and never fail the check, because they don't decide what a guest observes: Azure WHP hosts show the CPUID bit but cannot offer invariant TSC to partitions, and their guests measure tens of nanoseconds of skew |
| H3 | openvmm --x-time-abi-verify builds the partition and runs the time ABI preflight without running the guest. Its NVX-TIME-ABI-VERIFY: line must report status=ok for the backend, plausible declared and native TSC rates, the backend's LAPIC rate, and a cpu_profile that is a catalog profile, <vendor>.<generation>.v<revision> of any generation but a host profile's host, which rejects the host profile that auto falls back to on a CPU that no built-in profile serves, and, when H2 runs too, a revision of the profile H2 names. A failed preflight reports OpenVMM's code, for example [E_TSC_SYNC_UNSUPPORTED] |
| H4 | Samples of the TSC against the host's monotonic clocks, with sleeps between them so the host's CPUs idle: 13 samples 10 s apart, or 3 samples 1 s apart with --ci-schedule. Each sample reads the TSC between two reads of a clock, keeping the tightest of 64 brackets; its uncertainty is half the bracket plus half the clock's resolution. A clock that returns the same value to consecutive reads is coarser than one read, so the probe takes its smallest step as its resolution: Hyper-V's reference TSC page advances the Linux clocks in 100 ns steps although clock_getres reports 1 ns. The interval stability is judged against a clock that time synchronization never steers, CLOCK_MONOTONIC_RAW on Linux and QueryPerformanceCounter on Windows: every interval between consecutive samples must be conclusive within 0.25 ppm, and the interval rates must agree within 1 ppm. chrony's frequency updates move CLOCK_MONOTONIC's rate by up to several ppm between seconds on the Azure runners, which says nothing about the TSC. The rate over the whole window is measured against the disciplined clock, CLOCK_MONOTONIC on Linux, and must lie within 100 ppm of the rate H3 reports when H3 runs in the same invocation. A Linux host's clocksource is recorded as evidence |
| H5 | Pinned-thread ping-pong rounds over every pair of host CPUs; max_abs_offset_ns is at most 1,000, the measurement is conclusive, and no pair stalls |
| H6 | A microVM with the largest supported vCPU count up to 8 reports a valid NVX-TIME-ABI boot line through nvx-time status and runs nvx-time-probe warp over every CPU pair five times, with all vCPUs halted for 0.1, 1, 5, and 1 s between the runs, so that a host without an invariant TSC corrects the guest TSC as idle host CPUs wake (#265); then a 1-vCPU microVM runs it once. Every run stays within 1,000 ns. --ci-schedule runs CI's schedule instead: two runs 1 s apart |
| H7 | adjtimex reports no STA_UNSYNC and no TIME_ERROR on Linux; w32tm /query /status names a synchronized source on Windows. When H7 can't read either, its detail starts with cannot verify host UTC synchronization: adjtimex fails or reports a clock state Linux doesn't define, or w32tm fails or prints no English Source and Leap Indicator lines, as in a Windows display language other than English |
H2, H4, and H5 use a dependency-free host probe,
host_time_probe.rs, which the
doctor builds with rustc once per source version into
$RUNNER_TOOL_CACHE/nvx-host-time-probe (or build/host-time-probe outside
CI). H2's profile check, H3, and H6 need the OpenVMM binary, and H3 and H6
also need the guest artifacts; --openvmm, --kernel, and --initrd
override their default build paths. --no-openvmm qualifies a host without
OpenVMM: H2 then checks the CPU identity and generation only, and H3 and H6
cannot run.
The doctor gates on measured properties, alike on every backend: the guest
warp probe at 1 µs with its idle gaps (H6), the TSC rate stability (H4), and
the CPU profile (H2 and H3). The host OS's invariant-TSC flags and clocksource
are evidence only, in the doctor and in CI. In CI, validate-runner runs the
cheap host-level checks H1, H2, and H4 in every job that uses it, on the short
--ci-schedule. The microVM and platform jobs run it after downloading
OpenVMM and the guest artifacts and add H2's CPU profile check and H3, so that
H4 also compares the measured rate with the backend's native rate; the other
jobs pass --no-openvmm. The guest warp probe needs a time ABI boot: the
platform jobs add H6 on CI's schedule before their benchmarks, and the microVM
boot and restore scenarios run CI's warp schedule after every boot and restore
and assert its verdict. H5 and H7 remain available for interactive
qualification.
rust-toolchain.toml pins the exact Rust release
that builds OpenVMM, its test guests, the aci_edge_sandboxes crate, and the
host tools that NVX compiles, such as the doctor's time probe and the Windows
curl shim. Rustup applies the file to every cargo and rustc run in the
checkout, including in the openvmm submodule, and test-openvmm also exports
the release as RUSTUP_TOOLCHAIN because flowey can run Cargo outside the
checkout. The pin is an exact release rather than stable, so runners
provisioned at different times build with the same compiler (#402). Rust 1.99.0
cannot link OpenVMM's guest_test_uefi, failing with undefined symbol: wcslen
(rust-lang/rust#163614).
Self-hosted runners keep the release in a trusted RUSTUP_HOME that jobs
cannot modify, so the runner setup scripts
install it with its OpenVMM targets. validate-runner disables rustup's
automatic installation and fails a job before it builds anything when its runner
lacks the release, naming the setup script to rerun. The OpenVMM binary cache
key includes a hash of the file, so a new release rebuilds OpenVMM, and a change
to the file runs the OpenVMM test matrix.
To move the pin, first confirm that guest_test_uefi links with the new
release:
rustup toolchain install RELEASE --profile minimal --target x86_64-unknown-uefi
cd openvmm
cargo +RELEASE build -p guest_test_uefi --target x86_64-unknown-uefi --profile releaseThen change the release in rust-toolchain.toml,
OpenVMMBuildConstants.RUST_TOOLCHAIN, the three setup scripts in
scripts/setup, and the Specula runner setup together; a test in
scripts/test_nvx_tools.py fails while they differ. Update every self-hosted
runner as the setup guide describes before the change merges.
The required aci-edge-sandboxes job checks the aci_edge_sandboxes crate on
GitHub-hosted Ubuntu and Windows runners through the
check-aci-edge-sandboxes action. The action
runs rustfmt, Clippy with warnings denied, and rustdoc, all on the Rust
release that rust-toolchain.toml pins. It also runs the unit, mock, and fake-OpenVMM
integration tests, checks the declared minimum Rust version, and cross-checks
the macOS build that MXC compiles. The fake-OpenVMM tests drive the real
OpenVMM backend through its control protocol without a hypervisor.
The action also validates bundled artifact staging with a fixture package.
The development release job depends on aci-edge-sandboxes and requires its combined
Linux/Windows result to be successful; failed, cancelled, or skipped crate
checks cannot publish a release.
Each nvx-microvm-tests-{kvm,mshv,whp} job then runs
nvx.py test-aci-edge-sandboxes on its self-hosted runner; the debug-kernel
jobs skip it, because it boots the production kernel. This command drives a
complete provision, start, exec, stop, start, and deprovision cycle of the
Alpine guest with the crate's OpenVMM backend. It also checks cancellation,
that guest state lasts only until a stop, that a start terminates the VM of an
interrupted earlier start, host path mappings, egress rules, and per-execution
environments and working directories. On failure
it keeps the OpenVMM log under build/test-results/aci-edge-sandboxes-<backend>, which is
uploaded with the other microVM logs. Changes under aci_edge_sandboxes/ therefore trigger
the NVX microVM test matrix.
The separate
adversarial.yml workflow runs
Copilot-driven campaigns only on trusted manual dispatches or schedules from
dev. It is not part of pull-request CI. The workflow's dedicated
nvx-adversarial-controller runner must already have an authenticated Copilot
CLI and an administrator-owned executor wrapper named by the
NVX_ADVERSARIAL_EXECUTOR repository variable. The workflow does not install
Copilot or initiate login.
The wrapper provisions a distinct disposable KVM, MSHV, or WHP target with no production or GitHub credentials and forwards only the typed executor protocol. Existing persistent microVM and performance runners are not valid adversarial targets. Loss of the target heartbeat, a policy oracle, or a teardown/post-campaign boot failure fails the job and requires quarantine and reimage.
Normal Actions artifacts contain only the guest-text-free public summary, catalogued case identifiers, and replay manifest. The external provisioner must collect controller transcripts and complete target logs into access-controlled security storage. See Copilot-driven adversarial testing for the architecture and operational contract.
copilot-setup-steps.yml
prepares the GitHub-hosted runner of Copilot cloud agent sessions, including
sessions started from the GitHub web interface. The session's 59-minute limit
includes these steps, so they only install pinned tools, restore caches, and
stage downloads; they never compile NVX or OpenVMM. The workflow initializes
the public OpenVMM submodule without a deploy key, grants access to the
runner's /dev/kvm, installs the Python development tools into a virtual
environment on the image's Python 3, which validate-nvx also uses, and
installs Rust and cargo-nextest at the versions that
rust-toolchain.toml, the crate manifest, and the Linux runner bootstrap
pin. It restores the shared guest artifacts
through the restore-only input of
build-guest-artifacts
and the KVM OpenVMM binary that build-openvmm caches for the pinned revision,
so agents can run microVM tests without rebuilding either. The guest artifact
restore computes its keys with the same action that saves them. The binary
restore repeats the key, paths, and gzip compression of the cache entry that
build-openvmm saves, and a test in scripts/test_nvx_tools.py fails when
they differ, because a mismatch would otherwise only cause a silent cache miss.
Because the agent firewall blocks cdn.kernel.org and cdimage.ubuntu.com,
even inside containers, the workflow also installs the native guest build
prerequisites, allows the unprivileged user namespaces that Alpine's apk uses
for package triggers, and stages the pinned Linux and Ubuntu Base archives, so
agents can rebuild guest artifacts with build-guest --native. When a runner
has a separate /mnt disk with more free space than /, it mounts the
workspace there. A failed step would make Copilot skip every later setup step,
so each step after checkout continues on error. The last step sets
NVX_COPILOT_SETUP to complete or incomplete, lists failed step IDs in
NVX_COPILOT_SETUP_FAILED, and fails the run when setup is incomplete.
Copilot always runs the version on dev, even for sessions based on other
branches, so changes reach agent sessions only after they merge. Pushes run it
as a normal workflow for validation when they change the workflow or a file
that its steps take versions, requirements, metadata, or code from: the
build-guest-artifacts action, the files that pin its Rust, cargo-nextest,
and shell linter versions, requirements-dev.txt, SOURCE-MANIFEST.json,
.gitmodules, the OpenVMM submodule pin, and the NVX CLI and its modules.
Copilot code review uses
copilot-code-review.yml
instead. Reviews do not build code, so it only checks out the repository and
does not install the build and test dependencies used by cloud agent sessions.