Commit 7ca3e2

2026-07-08 06:42:13 Ralf Lici: dco: add a page to describe the Linux kernel module CI on GitHub Actions
/dev/null .. DataChannelOffload/CI for DCO Linux.md
@@ 0,0 1,262 @@
+ # CI for DCO Linux
+
+ Distributing an out-of-tree kernel module means signing up to support a wide
+ variety of kernels. An enterprise distro freezes one for years and backports
+ fixes into it the whole time; a rolling release ships something new every week;
+ an LTS sits somewhere in between. Distros carry these very different kernels
+ and patch them on their own schedule, without asking anyone. A stable-tree
+ update or an enterprise backport can quietly change an assumption the module
+ relied on, and the first sign of it is a bug report.
+
+ So the OpenVPN kernel modules need more than a normal userspace CI job. A
+ useful test has to build against real distro kernel headers, boot that kernel,
+ load the module, and either run the module selftests or at least prove the
+ module still compiles on the supported platforms.
+
+ That creates an awkward constraint: free GitHub-hosted runners are only provided
+ as Ubuntu, macOS, or Windows machines, while the module has to work across
+ Debian, Ubuntu, Fedora, RHEL, and openSUSE kernels. The CI therefore cannot
+ simply "use a Fedora runner" or "use a RHEL runner". (Red Hat Enterprise Linux
+ runner images entered [public preview](https://github.blog/changelog/2026-06-25-red-hat-enterprise-linux-runner-images-are-now-in-public-preview/)
+ in mid-2026, but only on the larger paid runners.) Instead, the Ubuntu runner
+ becomes a build host that generates a target distro rootfs, boots it in a nested
+ VM, and runs the actual module workload inside that guest.
+
+ The result is [kmod-ci](https://github.com/mandelbitdev/kmod-ci), a reusable
+ GitHub Actions workflow for out-of-tree OpenVPN kernel modules. Caller
+ repositories such as
+ [`ovpn-backports`](https://github.com/OpenVPN/ovpn-backports) and
+ [`ovpn-dco`](https://github.com/OpenVPN/ovpn-dco) only provide the small
+ module-specific script to run inside the guest. The shared CI repository owns
+ the distro matrix, rootfs generation, virtme-ng boot logic, RHEL handling, and
+ scheduled-run optimization.
+
+ ## High-level flow
+
+ Each matrix job starts on a standard `ubuntu-24.04` GitHub runner. The runner
+ installs the host-side tools needed to build root filesystems and boot virtual
+ machines: `mmdebstrap`, `dnf`, `qemu`, `virtme-ng`, `virtiofsd`, and related
+ utilities.
+
+ The job then checks out two repositories:
+
+ - the caller repository: the module repository being tested, copied into the
+ guest as `/repo`.
+ - the shared CI repository: it provides the scripts that know how to generate
+ root filesystems and boot the guest.
+
+ For each distro target, the workflow builds a rootfs using that distro's normal
+ package repositories. Debian and Ubuntu are generated with `mmdebstrap`.
+ Fedora, RHEL, AlmaLinux, and openSUSE targets are generated with `dnf
+ --installroot`. The generated rootfs includes the distro kernel, kernel
+ headers, compiler toolchain, module build dependencies, and the small set of
+ userspace tools needed by the module tests.
+
+ For RPM-based targets, the shared CI also carries the small repository
+ definitions needed to bootstrap the installroot. They are based on the official
+ repo files shipped by each distribution, with only the adjustments needed to use
+ them from an Ubuntu GitHub runner.
+
+ Once the rootfs exists, the workflow boots it with virtme-ng. Inside the guest,
+ the caller-provided script runs as root. For `ovpn-backports`, that script
+ builds the module and runs the ovpn kselftests. For `ovpn-dco`, the current
+ useful signal is primarily whether the module still builds across the supported
+ kernel matrix.
+
+ ## Why boot the distro kernel?
+
+ The important part is that the module is not only compiled on Ubuntu. It is
+ compiled and loaded against the kernel shipped by the target distribution.
+
+ That matters because enterprise distributions routinely carry large backport
+ sets. A RHEL 9 kernel may report a 5.14 base version, but its APIs do not
+ behave exactly like upstream Linux 5.14. The same is true, to varying degrees,
+ for SUSE and other long-term distro kernels. Testing only upstream version
+ numbers would miss the real compatibility surface.
+
+ The CI therefore treats the distro kernel package as the source of truth. It
+ installs the kernel and matching headers from the distro repositories, boots
+ that kernel, and runs the module workload there.
+
+ ## Current distro matrix
+
+ The exact kernel release is not hardcoded in the workflow. It moves whenever
+ the distribution updates the corresponding kernel package. At the time of
+ writing, the full selftest matrix covers this kernel range:
+
+ | Target | Kernel family | Notes |
+ | --- | --- | --- |
+ | Debian 10 | 4.19 | old LTS coverage |
+ | Debian 11 | 5.10 | LTS coverage |
+ | Debian 12 | 6.1 | LTS coverage |
+ | Debian 13 | 6.12 | LTS coverage |
+ | Fedora 44 | 7.1 | fast-moving Fedora target |
+ | openSUSE Leap 15.6 | 6.4 | enterprise-style openSUSE target |
+ | openSUSE Leap 16.0 | 6.12 | enterprise-style openSUSE target |
+ | openSUSE Tumbleweed | 7.1 | rolling release |
+ | RHEL 8 | 4.18 | enterprise LTS kernel |
+ | RHEL 9 | 5.14 | enterprise LTS kernel |
+ | RHEL 10 | 6.12 | enterprise LTS kernel |
+ | Ubuntu 20.04 | 5.4 | LTS coverage |
+ | Ubuntu 22.04 | 5.15 | LTS coverage |
+ | Ubuntu 24.04 | 6.8 | LTS coverage |
+ | Ubuntu 25.10 | 6.17 | interim Ubuntu target |
+ | Ubuntu 26.04 | 7.0 | LTS coverage |
+
+ ## RHEL support
+
+ RHEL is the one target family that cannot be generated from public repository
+ files alone. The workflow uses Red Hat `subscription-manager` inside a
+ temporary [Red Hat Universal Base Image
+ (UBI)](https://catalog.redhat.com/en/software/base-images) container to
+ authenticate against Red Hat CDN and enable the required repositories.
+
+ AlmaLinux remains supported as a convenient EL-family target, and it was very
+ useful while the CI was being brought up because it made iteration quick and
+ did not require credentials. For this project, though, real RHEL is the better
+ signal. RHEL kernel updates are exactly the updates that often matter to users
+ running OpenVPN in enterprise environments, and clones only see those changes
+ after Red Hat has already shipped them.
+
+ The caller repository provides two GitHub Actions secrets:
+
+ - `RHEL_ORG_ID`
+ - `RHEL_ACTIVATION_KEY`
+
+ Those secrets are passed only to the RHEL rootfs generation step. The UBI
+ container registers with RHSM, enables BaseOS, AppStream, and CodeReady
+ Builder, generates the `redhat.repo` file, and uses that to build the
+ installroot.
+
+ The container is run privileged because the RPM rootfs generation path mounts
+ `/dev`, `/proc`, `/sys`, and `/run` into the installroot. Kernel package
+ scriptlets expect something closer to a normal running system, especially when
+ installing kernel packages and running dracut or kernel-install hooks.
+
+ The registration is cleaned up on exit. This avoids accumulating RHSM consumer
+ registrations on every matrix run, including failing runs. The generated rootfs
+ is also scrubbed defensively so RHSM identity and entitlement state are not
+ carried forward into the VM artifact.
+
+ ## virtme-ng and filesystem transports
+
+ The guest is booted with virtme-ng because it is much lighter than maintaining
+ full disk images for every distro. The workflow generates a directory rootfs,
+ asks virtme-ng to boot the target kernel, and exposes the caller repository to
+ the guest. It also aligns with [NIPA](https://github.com/linux-netdev/nipa),
+ the netdev CI system, which uses virtme-ng for kernel test boots as well.
+
+ Most targets use virtiofs, which is fast enough for repeated module builds and
+ selftests. Older kernels can be more difficult. Debian 10, in particular, uses
+ a 4.19 kernel and has to fall back to the legacy 9p path in this setup. That
+ target is noticeably slower, but still useful because it covers the oldest
+ kernel generation the module is expected to support.
+
+ The other important performance detail is KVM. GitHub's Linux runners currently
+ expose `/dev/kvm`, so nested virtualization works and the guest boots are quick.
+ The guest runner script still probes for `/dev/kvm` before booting. If it is
+ missing, it adds `--disable-kvm` and continues with software emulation rather
+ than failing outright. That fallback is a safety net: GitHub documents nested
+ virtualization on hosted runners as
+ ["experimental and done at your own risk," with no guarantees on stability or performance](https://docs.github.com/en/actions/concepts/runners/github-hosted-runners#runner-images).
+ Today it works; the CI probes anyway.
+
+ RHEL kernels also forced work in this area because they do not provide 9p at
+ all, and virtme-ng historically mounted some helper exports over 9p. That
+ work became
+ [`virtme-ng` PR #472](https://github.com/arighi/virtme-ng/pull/472), which
+ removes the remaining hardcoded 9p helper exports and uses virtiofs instead
+ when virtiofs is available.
+
+ One related issue is still open at the time of writing:
+ [`virtme-ng` issue #475](https://github.com/arighi/virtme-ng/issues/475).
+ Some older kernels fail external-rootfs virtiofs boots when `virtiofsd` is
+ started with `--posix-acl`. The CI still carries the corresponding workaround
+ until that behavior is handled upstream.
+
+ ## Scheduled runs
+
+ The workflow can be triggered manually, but it is also designed to run nightly.
+
+ That is useful because the module repository may not change every day, while
+ distro kernels do. A RHEL or openSUSE kernel update can break the module even
+ if no OpenVPN code changed. Scheduled CI catches that kind of regression soon
+ after the distro publishes the update.
+
+ To keep scheduled runs from wasting time, the workflow records a small success
+ marker keyed by the caller repository revision, the distro target, and the
+ kernel release installed in the generated rootfs. On a scheduled run, the
+ rootfs is still generated first, because that is the reliable way to know which
+ kernel the distro currently serves. If the same repo revision has already
+ passed with that same target kernel, the expensive guest workload is skipped.
+
+ This keeps the behavior simple and correct. It does not try to predict
+ repository state with a separate metadata parser. The package manager remains
+ the source of truth.
+
+ ## Runtime
+
+ With the `ovpn-backports` selftest workload, most targets complete in roughly
+ seven to eleven minutes. Debian 10 is the expected outlier: because it uses the
+ older 9p path, it takes about fifteen minutes, and therefore determines the
+ total runtime of the full parallel matrix.
+
+ The compile-only `ovpn-dco` workload is shorter because it does not run the
+ kselftests. The Ubuntu and Debian jobs complete in about two minutes, and the
+ RHEL jobs complete in roughly three to five minutes.
+
+ ## Shared workflow model
+
+ The useful boundary is that the shared CI owns the infrastructure, while each
+ module repository owns only the command that should run inside the guest.
+
+ That keeps the caller workflows small. A repository can opt into the matrix by
+ calling the reusable workflow and passing:
+
+ - the cache prefix,
+ - the guest script path,
+ - optionally a preparation command,
+ - optionally a custom distro list.
+
+ The same mechanism can support both `ovpn-backports` and `ovpn-dco` without
+ duplicating the rootfs and virtme-ng machinery in each repository.
+
+ It also keeps distro-specific quirks contained. Debian archive handling, Ubuntu
+ keyrings, RHEL subscription setup, openSUSE repository files, old-kernel 9p
+ fallback, and RPM scriptlet behavior are all handled in one place.
+
+ ## Tradeoffs
+
+ This is not as fast as a native runner per distribution would be, but those
+ runners do not exist on GitHub-hosted Actions. It is also not as hermetic as
+ prebuilt qcow images, but avoiding checked-in or externally hosted VM images
+ makes the CI self-contained and easier to reproduce.
+
+ The rootfs is generated every time before the scheduled cache decision. That is
+ intentionally a little wasteful. The alternative would be duplicating
+ package-manager logic just to predict whether the kernel changed, and that
+ would be more fragile than letting apt or dnf answer the question directly.
+
+ Some distro rough edges remain. Debian 10 is slow because it falls back to 9p.
+ openSUSE has required explicit key handling and has shown transient failures
+ caused by repo unreachability. RHEL requires credentials and a temporary RHSM
+ registration.
+
+ ## Current status
+
+ The system now runs nightly across the default matrix and gives the OpenVPN
+ kernel modules a real distro-kernel compatibility signal on ordinary
+ GitHub-hosted runners.
+
+ Tested modules:
+ - ovpn: https://github.com/OpenVPN/ovpn-backports/actions/workflows/vng-selftests.yml
+ - ovpn-dco-v2: https://github.com/OpenVPN/ovpn-dco/actions/workflows/vng-build.yml
+
+ ## Roadmap
+
+ The next useful additions are arm64 coverage and upstreaming the remaining
+ virtme-ng fix so the CI can stop depending on a fork. Deliberately not on the
+ roadmap is an ever-longer distro list. This CI is testing kernel code rather
+ than userspace integration, so the important coverage is the range of kernels
+ people actually run. That set changes naturally as distributions move to new
+ kernel versions.
0 1 2 3 4 5 6 7 8 9 0 1 2 3 4 5 6 7 8 9 0 1 2 3 4 5 6 7 8 9 0 1 2 3 4 5 6 7 8 9