Blame

7ca3e2 Ralf Lici 2026-07-08 06:42:13 1
# CI for DCO Linux
2
3
Distributing an out-of-tree kernel module means signing up to support a wide
4
variety of kernels. An enterprise distro freezes one for years and backports
5
fixes into it the whole time; a rolling release ships something new every week;
6
an LTS sits somewhere in between. Distros carry these very different kernels
7
and patch them on their own schedule, without asking anyone. A stable-tree
8
update or an enterprise backport can quietly change an assumption the module
9
relied on, and the first sign of it is a bug report.
10
11
So the OpenVPN kernel modules need more than a normal userspace CI job. A
12
useful test has to build against real distro kernel headers, boot that kernel,
13
load the module, and either run the module selftests or at least prove the
14
module still compiles on the supported platforms.
15
16
That creates an awkward constraint: free GitHub-hosted runners are only provided
17
as Ubuntu, macOS, or Windows machines, while the module has to work across
18
Debian, Ubuntu, Fedora, RHEL, and openSUSE kernels. The CI therefore cannot
19
simply "use a Fedora runner" or "use a RHEL runner". (Red Hat Enterprise Linux
20
runner images entered [public preview](https://github.blog/changelog/2026-06-25-red-hat-enterprise-linux-runner-images-are-now-in-public-preview/)
21
in mid-2026, but only on the larger paid runners.) Instead, the Ubuntu runner
22
becomes a build host that generates a target distro rootfs, boots it in a nested
23
VM, and runs the actual module workload inside that guest.
24
25
The result is [kmod-ci](https://github.com/mandelbitdev/kmod-ci), a reusable
26
GitHub Actions workflow for out-of-tree OpenVPN kernel modules. Caller
27
repositories such as
28
[`ovpn-backports`](https://github.com/OpenVPN/ovpn-backports) and
29
[`ovpn-dco`](https://github.com/OpenVPN/ovpn-dco) only provide the small
30
module-specific script to run inside the guest. The shared CI repository owns
31
the distro matrix, rootfs generation, virtme-ng boot logic, RHEL handling, and
32
scheduled-run optimization.
33
34
## High-level flow
35
f35a29 Ralf Lici 2026-07-21 07:09:48 36
Each matrix job starts on a standard `ubuntu-24.04` or `ubuntu-24.04-arm64`
37
GitHub runner. The runner installs the host-side tools needed to build root
38
filesystems and boot virtual machines: `mmdebstrap`, `dnf`, `qemu`,
39
`virtme-ng`, `virtiofsd`, and related utilities.
7ca3e2 Ralf Lici 2026-07-08 06:42:13 40
41
The job then checks out two repositories:
42
43
- the caller repository: the module repository being tested, copied into the
44
guest as `/repo`.
45
- the shared CI repository: it provides the scripts that know how to generate
46
root filesystems and boot the guest.
47
48
For each distro target, the workflow builds a rootfs using that distro's normal
49
package repositories. Debian and Ubuntu are generated with `mmdebstrap`.
50
Fedora, RHEL, AlmaLinux, and openSUSE targets are generated with `dnf
51
--installroot`. The generated rootfs includes the distro kernel, kernel
52
headers, compiler toolchain, module build dependencies, and the small set of
53
userspace tools needed by the module tests.
54
55
For RPM-based targets, the shared CI also carries the small repository
56
definitions needed to bootstrap the installroot. They are based on the official
57
repo files shipped by each distribution, with only the adjustments needed to use
58
them from an Ubuntu GitHub runner.
59
60
Once the rootfs exists, the workflow boots it with virtme-ng. Inside the guest,
f35a29 Ralf Lici 2026-07-21 07:09:48 61
the caller-provided script runs as root. For `ovpn-backports` on x86_64, that
62
script builds the module and runs the ovpn kselftests. For `ovpn-dco` and
63
`ovpn-backports` on arm64, the current useful signal is primarily whether the
64
module still builds across the supported kernel matrix.
7ca3e2 Ralf Lici 2026-07-08 06:42:13 65
66
## Why boot the distro kernel?
67
68
The important part is that the module is not only compiled on Ubuntu. It is
69
compiled and loaded against the kernel shipped by the target distribution.
70
71
That matters because enterprise distributions routinely carry large backport
72
sets. A RHEL 9 kernel may report a 5.14 base version, but its APIs do not
73
behave exactly like upstream Linux 5.14. The same is true, to varying degrees,
74
for SUSE and other long-term distro kernels. Testing only upstream version
75
numbers would miss the real compatibility surface.
76
77
The CI therefore treats the distro kernel package as the source of truth. It
78
installs the kernel and matching headers from the distro repositories, boots
79
that kernel, and runs the module workload there.
80
81
## Current distro matrix
82
83
The exact kernel release is not hardcoded in the workflow. It moves whenever
84
the distribution updates the corresponding kernel package. At the time of
85
writing, the full selftest matrix covers this kernel range:
86
87
| Target | Kernel family | Notes |
88
| --- | --- | --- |
89
| Debian 10 | 4.19 | old LTS coverage |
90
| Debian 11 | 5.10 | LTS coverage |
91
| Debian 12 | 6.1 | LTS coverage |
92
| Debian 13 | 6.12 | LTS coverage |
93
| Fedora 44 | 7.1 | fast-moving Fedora target |
94
| openSUSE Leap 15.6 | 6.4 | enterprise-style openSUSE target |
95
| openSUSE Leap 16.0 | 6.12 | enterprise-style openSUSE target |
96
| openSUSE Tumbleweed | 7.1 | rolling release |
97
| RHEL 8 | 4.18 | enterprise LTS kernel |
98
| RHEL 9 | 5.14 | enterprise LTS kernel |
99
| RHEL 10 | 6.12 | enterprise LTS kernel |
100
| Ubuntu 20.04 | 5.4 | LTS coverage |
101
| Ubuntu 22.04 | 5.15 | LTS coverage |
102
| Ubuntu 24.04 | 6.8 | LTS coverage |
103
| Ubuntu 25.10 | 6.17 | interim Ubuntu target |
104
| Ubuntu 26.04 | 7.0 | LTS coverage |
105
106
## RHEL support
107
108
RHEL is the one target family that cannot be generated from public repository
109
files alone. The workflow uses Red Hat `subscription-manager` inside a
110
temporary [Red Hat Universal Base Image
111
(UBI)](https://catalog.redhat.com/en/software/base-images) container to
112
authenticate against Red Hat CDN and enable the required repositories.
113
114
AlmaLinux remains supported as a convenient EL-family target, and it was very
115
useful while the CI was being brought up because it made iteration quick and
116
did not require credentials. For this project, though, real RHEL is the better
117
signal. RHEL kernel updates are exactly the updates that often matter to users
118
running OpenVPN in enterprise environments, and clones only see those changes
119
after Red Hat has already shipped them.
120
121
The caller repository provides two GitHub Actions secrets:
122
123
- `RHEL_ORG_ID`
124
- `RHEL_ACTIVATION_KEY`
125
126
Those secrets are passed only to the RHEL rootfs generation step. The UBI
127
container registers with RHSM, enables BaseOS, AppStream, and CodeReady
128
Builder, generates the `redhat.repo` file, and uses that to build the
129
installroot.
130
131
The container is run privileged because the RPM rootfs generation path mounts
132
`/dev`, `/proc`, `/sys`, and `/run` into the installroot. Kernel package
133
scriptlets expect something closer to a normal running system, especially when
134
installing kernel packages and running dracut or kernel-install hooks.
135
136
The registration is cleaned up on exit. This avoids accumulating RHSM consumer
137
registrations on every matrix run, including failing runs. The generated rootfs
138
is also scrubbed defensively so RHSM identity and entitlement state are not
139
carried forward into the VM artifact.
140
141
## virtme-ng and filesystem transports
142
143
The guest is booted with virtme-ng because it is much lighter than maintaining
144
full disk images for every distro. The workflow generates a directory rootfs,
145
asks virtme-ng to boot the target kernel, and exposes the caller repository to
146
the guest. It also aligns with [NIPA](https://github.com/linux-netdev/nipa),
147
the netdev CI system, which uses virtme-ng for kernel test boots as well.
148
149
Most targets use virtiofs, which is fast enough for repeated module builds and
150
selftests. Older kernels can be more difficult. Debian 10, in particular, uses
151
a 4.19 kernel and has to fall back to the legacy 9p path in this setup. That
152
target is noticeably slower, but still useful because it covers the oldest
153
kernel generation the module is expected to support.
154
f35a29 Ralf Lici 2026-07-21 07:09:48 155
The other important performance detail is KVM. GitHub's x86_64 Linux runners currently
7ca3e2 Ralf Lici 2026-07-08 06:42:13 156
expose `/dev/kvm`, so nested virtualization works and the guest boots are quick.
157
The guest runner script still probes for `/dev/kvm` before booting. If it is
158
missing, it adds `--disable-kvm` and continues with software emulation rather
159
than failing outright. That fallback is a safety net: GitHub documents nested
160
virtualization on hosted runners as
161
["experimental and done at your own risk," with no guarantees on stability or performance](https://docs.github.com/en/actions/concepts/runners/github-hosted-runners#runner-images).
f35a29 Ralf Lici 2026-07-21 07:09:48 162
In fact, on arm64, KVM is not exposed so we resort to software emulation which
163
is not only way slower but also instable (for example it occasionally segfaults
164
while building or installing the module). For this reason arm64 targets are
165
kept minimal and manual-dispatched only.
7ca3e2 Ralf Lici 2026-07-08 06:42:13 166
167
RHEL kernels also forced work in this area because they do not provide 9p at
168
all, and virtme-ng historically mounted some helper exports over 9p. That
169
work became
170
[`virtme-ng` PR #472](https://github.com/arighi/virtme-ng/pull/472), which
171
removes the remaining hardcoded 9p helper exports and uses virtiofs instead
172
when virtiofs is available.
173
f35a29 Ralf Lici 2026-07-21 07:09:48 174
Addidionatlly, some older kernels failed external-rootfs virtiofs boots when
175
`virtiofsd` was started with `--posix-acl`. [`virtme-ng` PR
176
#482](https://github.com/arighi/virtme-ng/pull/482) added `--no-root-posix-acl`
177
to explicitly control this behavior.
178
179
Finally, newer arm64 distro kernels are no longer plain Image files from QEMU's
180
point of view. Fedora 44 and recent Ubuntu development releases ship EFI zboot
181
images, and some Ubuntu kernels add another signed PE wrapper around the
182
bootable payload. [`virtme-ng` PR
183
#483](https://github.com/arighi/virtme-ng/pull/483) added support for arm64
184
kernel images normalization before passing them to QEMU.
7ca3e2 Ralf Lici 2026-07-08 06:42:13 185
186
## Scheduled runs
187
f35a29 Ralf Lici 2026-07-21 07:09:48 188
The workflow can be triggered manually, but the x86_64 matrix is also designed
189
to run nightly.
7ca3e2 Ralf Lici 2026-07-08 06:42:13 190
191
That is useful because the module repository may not change every day, while
192
distro kernels do. A RHEL or openSUSE kernel update can break the module even
193
if no OpenVPN code changed. Scheduled CI catches that kind of regression soon
194
after the distro publishes the update.
195
196
To keep scheduled runs from wasting time, the workflow records a small success
197
marker keyed by the caller repository revision, the distro target, and the
198
kernel release installed in the generated rootfs. On a scheduled run, the
199
rootfs is still generated first, because that is the reliable way to know which
200
kernel the distro currently serves. If the same repo revision has already
201
passed with that same target kernel, the expensive guest workload is skipped.
202
203
This keeps the behavior simple and correct. It does not try to predict
204
repository state with a separate metadata parser. The package manager remains
205
the source of truth.
206
207
## Runtime
208
209
With the `ovpn-backports` selftest workload, most targets complete in roughly
210
seven to eleven minutes. Debian 10 is the expected outlier: because it uses the
211
older 9p path, it takes about fifteen minutes, and therefore determines the
212
total runtime of the full parallel matrix.
213
f35a29 Ralf Lici 2026-07-21 07:09:48 214
The x86_64 `ovpn-dco` workload is shorter because it does not run the
7ca3e2 Ralf Lici 2026-07-08 06:42:13 215
kselftests. The Ubuntu and Debian jobs complete in about two minutes, and the
216
RHEL jobs complete in roughly three to five minutes.
217
f35a29 Ralf Lici 2026-07-21 07:09:48 218
The arm64 matrix, on the other hand, can take up to twenty minutes despite
219
being compile-only, because of the limitations of software emulation.
220
7ca3e2 Ralf Lici 2026-07-08 06:42:13 221
## Shared workflow model
222
223
The useful boundary is that the shared CI owns the infrastructure, while each
224
module repository owns only the command that should run inside the guest.
225
226
That keeps the caller workflows small. A repository can opt into the matrix by
227
calling the reusable workflow and passing:
228
229
- the cache prefix,
230
- the guest script path,
231
- optionally a preparation command,
232
- optionally a custom distro list.
233
234
The same mechanism can support both `ovpn-backports` and `ovpn-dco` without
235
duplicating the rootfs and virtme-ng machinery in each repository.
236
237
It also keeps distro-specific quirks contained. Debian archive handling, Ubuntu
238
keyrings, RHEL subscription setup, openSUSE repository files, old-kernel 9p
239
fallback, and RPM scriptlet behavior are all handled in one place.
240
241
## Tradeoffs
242
243
This is not as fast as a native runner per distribution would be, but those
244
runners do not exist on GitHub-hosted Actions. It is also not as hermetic as
245
prebuilt qcow images, but avoiding checked-in or externally hosted VM images
246
makes the CI self-contained and easier to reproduce.
247
248
The rootfs is generated every time before the scheduled cache decision. That is
249
intentionally a little wasteful. The alternative would be duplicating
250
package-manager logic just to predict whether the kernel changed, and that
251
would be more fragile than letting apt or dnf answer the question directly.
252
253
Some distro rough edges remain. Debian 10 is slow because it falls back to 9p.
254
openSUSE has required explicit key handling and has shown transient failures
255
caused by repo unreachability. RHEL requires credentials and a temporary RHSM
256
registration.
257
258
## Current status
259
f35a29 Ralf Lici 2026-07-21 07:09:48 260
The system now runs nightly across the default x86_64 matrix and gives the
261
OpenVPN kernel modules a real distro-kernel compatibility signal on ordinary
7ca3e2 Ralf Lici 2026-07-08 06:42:13 262
GitHub-hosted runners.
263
264
Tested modules:
f35a29 Ralf Lici 2026-07-21 07:09:48 265
- ovpn: https://github.com/OpenVPN/ovpn-backports/actions
266
- ovpn-dco-v2: https://github.com/OpenVPN/ovpn-dco/actions