Commit 7ca3e2
2026-07-08 06:42:13 Ralf Lici: dco: add a page to describe the Linux kernel module CI on GitHub Actions| /dev/null .. DataChannelOffload/CI for DCO Linux.md | |
| @@ 0,0 1,262 @@ | |
| + | # CI for DCO Linux |
| + | |
| + | Distributing an out-of-tree kernel module means signing up to support a wide |
| + | variety of kernels. An enterprise distro freezes one for years and backports |
| + | fixes into it the whole time; a rolling release ships something new every week; |
| + | an LTS sits somewhere in between. Distros carry these very different kernels |
| + | and patch them on their own schedule, without asking anyone. A stable-tree |
| + | update or an enterprise backport can quietly change an assumption the module |
| + | relied on, and the first sign of it is a bug report. |
| + | |
| + | So the OpenVPN kernel modules need more than a normal userspace CI job. A |
| + | useful test has to build against real distro kernel headers, boot that kernel, |
| + | load the module, and either run the module selftests or at least prove the |
| + | module still compiles on the supported platforms. |
| + | |
| + | That creates an awkward constraint: free GitHub-hosted runners are only provided |
| + | as Ubuntu, macOS, or Windows machines, while the module has to work across |
| + | Debian, Ubuntu, Fedora, RHEL, and openSUSE kernels. The CI therefore cannot |
| + | simply "use a Fedora runner" or "use a RHEL runner". (Red Hat Enterprise Linux |
| + | runner images entered [public preview](https://github.blog/changelog/2026-06-25-red-hat-enterprise-linux-runner-images-are-now-in-public-preview/) |
| + | in mid-2026, but only on the larger paid runners.) Instead, the Ubuntu runner |
| + | becomes a build host that generates a target distro rootfs, boots it in a nested |
| + | VM, and runs the actual module workload inside that guest. |
| + | |
| + | The result is [kmod-ci](https://github.com/mandelbitdev/kmod-ci), a reusable |
| + | GitHub Actions workflow for out-of-tree OpenVPN kernel modules. Caller |
| + | repositories such as |
| + | [`ovpn-backports`](https://github.com/OpenVPN/ovpn-backports) and |
| + | [`ovpn-dco`](https://github.com/OpenVPN/ovpn-dco) only provide the small |
| + | module-specific script to run inside the guest. The shared CI repository owns |
| + | the distro matrix, rootfs generation, virtme-ng boot logic, RHEL handling, and |
| + | scheduled-run optimization. |
| + | |
| + | ## High-level flow |
| + | |
| + | Each matrix job starts on a standard `ubuntu-24.04` GitHub runner. The runner |
| + | installs the host-side tools needed to build root filesystems and boot virtual |
| + | machines: `mmdebstrap`, `dnf`, `qemu`, `virtme-ng`, `virtiofsd`, and related |
| + | utilities. |
| + | |
| + | The job then checks out two repositories: |
| + | |
| + | - the caller repository: the module repository being tested, copied into the |
| + | guest as `/repo`. |
| + | - the shared CI repository: it provides the scripts that know how to generate |
| + | root filesystems and boot the guest. |
| + | |
| + | For each distro target, the workflow builds a rootfs using that distro's normal |
| + | package repositories. Debian and Ubuntu are generated with `mmdebstrap`. |
| + | Fedora, RHEL, AlmaLinux, and openSUSE targets are generated with `dnf |
| + | --installroot`. The generated rootfs includes the distro kernel, kernel |
| + | headers, compiler toolchain, module build dependencies, and the small set of |
| + | userspace tools needed by the module tests. |
| + | |
| + | For RPM-based targets, the shared CI also carries the small repository |
| + | definitions needed to bootstrap the installroot. They are based on the official |
| + | repo files shipped by each distribution, with only the adjustments needed to use |
| + | them from an Ubuntu GitHub runner. |
| + | |
| + | Once the rootfs exists, the workflow boots it with virtme-ng. Inside the guest, |
| + | the caller-provided script runs as root. For `ovpn-backports`, that script |
| + | builds the module and runs the ovpn kselftests. For `ovpn-dco`, the current |
| + | useful signal is primarily whether the module still builds across the supported |
| + | kernel matrix. |
| + | |
| + | ## Why boot the distro kernel? |
| + | |
| + | The important part is that the module is not only compiled on Ubuntu. It is |
| + | compiled and loaded against the kernel shipped by the target distribution. |
| + | |
| + | That matters because enterprise distributions routinely carry large backport |
| + | sets. A RHEL 9 kernel may report a 5.14 base version, but its APIs do not |
| + | behave exactly like upstream Linux 5.14. The same is true, to varying degrees, |
| + | for SUSE and other long-term distro kernels. Testing only upstream version |
| + | numbers would miss the real compatibility surface. |
| + | |
| + | The CI therefore treats the distro kernel package as the source of truth. It |
| + | installs the kernel and matching headers from the distro repositories, boots |
| + | that kernel, and runs the module workload there. |
| + | |
| + | ## Current distro matrix |
| + | |
| + | The exact kernel release is not hardcoded in the workflow. It moves whenever |
| + | the distribution updates the corresponding kernel package. At the time of |
| + | writing, the full selftest matrix covers this kernel range: |
| + | |
| + | | Target | Kernel family | Notes | |
| + | | --- | --- | --- | |
| + | | Debian 10 | 4.19 | old LTS coverage | |
| + | | Debian 11 | 5.10 | LTS coverage | |
| + | | Debian 12 | 6.1 | LTS coverage | |
| + | | Debian 13 | 6.12 | LTS coverage | |
| + | | Fedora 44 | 7.1 | fast-moving Fedora target | |
| + | | openSUSE Leap 15.6 | 6.4 | enterprise-style openSUSE target | |
| + | | openSUSE Leap 16.0 | 6.12 | enterprise-style openSUSE target | |
| + | | openSUSE Tumbleweed | 7.1 | rolling release | |
| + | | RHEL 8 | 4.18 | enterprise LTS kernel | |
| + | | RHEL 9 | 5.14 | enterprise LTS kernel | |
| + | | RHEL 10 | 6.12 | enterprise LTS kernel | |
| + | | Ubuntu 20.04 | 5.4 | LTS coverage | |
| + | | Ubuntu 22.04 | 5.15 | LTS coverage | |
| + | | Ubuntu 24.04 | 6.8 | LTS coverage | |
| + | | Ubuntu 25.10 | 6.17 | interim Ubuntu target | |
| + | | Ubuntu 26.04 | 7.0 | LTS coverage | |
| + | |
| + | ## RHEL support |
| + | |
| + | RHEL is the one target family that cannot be generated from public repository |
| + | files alone. The workflow uses Red Hat `subscription-manager` inside a |
| + | temporary [Red Hat Universal Base Image |
| + | (UBI)](https://catalog.redhat.com/en/software/base-images) container to |
| + | authenticate against Red Hat CDN and enable the required repositories. |
| + | |
| + | AlmaLinux remains supported as a convenient EL-family target, and it was very |
| + | useful while the CI was being brought up because it made iteration quick and |
| + | did not require credentials. For this project, though, real RHEL is the better |
| + | signal. RHEL kernel updates are exactly the updates that often matter to users |
| + | running OpenVPN in enterprise environments, and clones only see those changes |
| + | after Red Hat has already shipped them. |
| + | |
| + | The caller repository provides two GitHub Actions secrets: |
| + | |
| + | - `RHEL_ORG_ID` |
| + | - `RHEL_ACTIVATION_KEY` |
| + | |
| + | Those secrets are passed only to the RHEL rootfs generation step. The UBI |
| + | container registers with RHSM, enables BaseOS, AppStream, and CodeReady |
| + | Builder, generates the `redhat.repo` file, and uses that to build the |
| + | installroot. |
| + | |
| + | The container is run privileged because the RPM rootfs generation path mounts |
| + | `/dev`, `/proc`, `/sys`, and `/run` into the installroot. Kernel package |
| + | scriptlets expect something closer to a normal running system, especially when |
| + | installing kernel packages and running dracut or kernel-install hooks. |
| + | |
| + | The registration is cleaned up on exit. This avoids accumulating RHSM consumer |
| + | registrations on every matrix run, including failing runs. The generated rootfs |
| + | is also scrubbed defensively so RHSM identity and entitlement state are not |
| + | carried forward into the VM artifact. |
| + | |
| + | ## virtme-ng and filesystem transports |
| + | |
| + | The guest is booted with virtme-ng because it is much lighter than maintaining |
| + | full disk images for every distro. The workflow generates a directory rootfs, |
| + | asks virtme-ng to boot the target kernel, and exposes the caller repository to |
| + | the guest. It also aligns with [NIPA](https://github.com/linux-netdev/nipa), |
| + | the netdev CI system, which uses virtme-ng for kernel test boots as well. |
| + | |
| + | Most targets use virtiofs, which is fast enough for repeated module builds and |
| + | selftests. Older kernels can be more difficult. Debian 10, in particular, uses |
| + | a 4.19 kernel and has to fall back to the legacy 9p path in this setup. That |
| + | target is noticeably slower, but still useful because it covers the oldest |
| + | kernel generation the module is expected to support. |
| + | |
| + | The other important performance detail is KVM. GitHub's Linux runners currently |
| + | expose `/dev/kvm`, so nested virtualization works and the guest boots are quick. |
| + | The guest runner script still probes for `/dev/kvm` before booting. If it is |
| + | missing, it adds `--disable-kvm` and continues with software emulation rather |
| + | than failing outright. That fallback is a safety net: GitHub documents nested |
| + | virtualization on hosted runners as |
| + | ["experimental and done at your own risk," with no guarantees on stability or performance](https://docs.github.com/en/actions/concepts/runners/github-hosted-runners#runner-images). |
| + | Today it works; the CI probes anyway. |
| + | |
| + | RHEL kernels also forced work in this area because they do not provide 9p at |
| + | all, and virtme-ng historically mounted some helper exports over 9p. That |
| + | work became |
| + | [`virtme-ng` PR #472](https://github.com/arighi/virtme-ng/pull/472), which |
| + | removes the remaining hardcoded 9p helper exports and uses virtiofs instead |
| + | when virtiofs is available. |
| + | |
| + | One related issue is still open at the time of writing: |
| + | [`virtme-ng` issue #475](https://github.com/arighi/virtme-ng/issues/475). |
| + | Some older kernels fail external-rootfs virtiofs boots when `virtiofsd` is |
| + | started with `--posix-acl`. The CI still carries the corresponding workaround |
| + | until that behavior is handled upstream. |
| + | |
| + | ## Scheduled runs |
| + | |
| + | The workflow can be triggered manually, but it is also designed to run nightly. |
| + | |
| + | That is useful because the module repository may not change every day, while |
| + | distro kernels do. A RHEL or openSUSE kernel update can break the module even |
| + | if no OpenVPN code changed. Scheduled CI catches that kind of regression soon |
| + | after the distro publishes the update. |
| + | |
| + | To keep scheduled runs from wasting time, the workflow records a small success |
| + | marker keyed by the caller repository revision, the distro target, and the |
| + | kernel release installed in the generated rootfs. On a scheduled run, the |
| + | rootfs is still generated first, because that is the reliable way to know which |
| + | kernel the distro currently serves. If the same repo revision has already |
| + | passed with that same target kernel, the expensive guest workload is skipped. |
| + | |
| + | This keeps the behavior simple and correct. It does not try to predict |
| + | repository state with a separate metadata parser. The package manager remains |
| + | the source of truth. |
| + | |
| + | ## Runtime |
| + | |
| + | With the `ovpn-backports` selftest workload, most targets complete in roughly |
| + | seven to eleven minutes. Debian 10 is the expected outlier: because it uses the |
| + | older 9p path, it takes about fifteen minutes, and therefore determines the |
| + | total runtime of the full parallel matrix. |
| + | |
| + | The compile-only `ovpn-dco` workload is shorter because it does not run the |
| + | kselftests. The Ubuntu and Debian jobs complete in about two minutes, and the |
| + | RHEL jobs complete in roughly three to five minutes. |
| + | |
| + | ## Shared workflow model |
| + | |
| + | The useful boundary is that the shared CI owns the infrastructure, while each |
| + | module repository owns only the command that should run inside the guest. |
| + | |
| + | That keeps the caller workflows small. A repository can opt into the matrix by |
| + | calling the reusable workflow and passing: |
| + | |
| + | - the cache prefix, |
| + | - the guest script path, |
| + | - optionally a preparation command, |
| + | - optionally a custom distro list. |
| + | |
| + | The same mechanism can support both `ovpn-backports` and `ovpn-dco` without |
| + | duplicating the rootfs and virtme-ng machinery in each repository. |
| + | |
| + | It also keeps distro-specific quirks contained. Debian archive handling, Ubuntu |
| + | keyrings, RHEL subscription setup, openSUSE repository files, old-kernel 9p |
| + | fallback, and RPM scriptlet behavior are all handled in one place. |
| + | |
| + | ## Tradeoffs |
| + | |
| + | This is not as fast as a native runner per distribution would be, but those |
| + | runners do not exist on GitHub-hosted Actions. It is also not as hermetic as |
| + | prebuilt qcow images, but avoiding checked-in or externally hosted VM images |
| + | makes the CI self-contained and easier to reproduce. |
| + | |
| + | The rootfs is generated every time before the scheduled cache decision. That is |
| + | intentionally a little wasteful. The alternative would be duplicating |
| + | package-manager logic just to predict whether the kernel changed, and that |
| + | would be more fragile than letting apt or dnf answer the question directly. |
| + | |
| + | Some distro rough edges remain. Debian 10 is slow because it falls back to 9p. |
| + | openSUSE has required explicit key handling and has shown transient failures |
| + | caused by repo unreachability. RHEL requires credentials and a temporary RHSM |
| + | registration. |
| + | |
| + | ## Current status |
| + | |
| + | The system now runs nightly across the default matrix and gives the OpenVPN |
| + | kernel modules a real distro-kernel compatibility signal on ordinary |
| + | GitHub-hosted runners. |
| + | |
| + | Tested modules: |
| + | - ovpn: https://github.com/OpenVPN/ovpn-backports/actions/workflows/vng-selftests.yml |
| + | - ovpn-dco-v2: https://github.com/OpenVPN/ovpn-dco/actions/workflows/vng-build.yml |
| + | |
| + | ## Roadmap |
| + | |
| + | The next useful additions are arm64 coverage and upstreaming the remaining |
| + | virtme-ng fix so the CI can stop depending on a fork. Deliberately not on the |
| + | roadmap is an ever-longer distro list. This CI is testing kernel code rather |
| + | than userspace integration, so the important coverage is the range of kernels |
| + | people actually run. That set changes naturally as distributions move to new |
| + | kernel versions. |
