Commit 9b7f54

2025-06-30 10:20:07 cron2: TCP mode
MTU and Fragments.md ..
@@ 177,6 177,8 @@
The OpenVPN AS product defaults to `tun-mtu 1420` and pushes that to the OpenVPN Clients, because for that scenario it's more likely to work in most cases. The community OpenVPN defaults to `tun-mtu 1500` because this avoids MTU jumps in more complex scenarios.
+ Also, `tun-mtu` is not going to work if you want ethernet briding (tap mode) where you need the tap interface MTU to be the same as the LAN interface (no "packet too big" handling on the ethernet layer). This is somewhat of a niche use case, but it is something where OpenVPN can come in handy - and we need to be aware of the limitations.
+
### mssfix
If `tun-mtu` can not be changed, or does not work correctly due to ICMP packet too big getting lost, there is another feature in OpenVPN which makes it "work for most cases", `--mssfix`, for example setting `--mssfix 500 mtu`. What this does is to look at TCP SYN and SYN-ACK packets passing through OpenVPN, and modifying the "MSS value" in there
@@ 185,16 187,53 @@
11:20:34.626186 tun2 In IP6 (flowlabel 0xe3a26, hlim 64, next-header TCP (6) payload length: 40) fd00:abcd:194:2::1.53 > fd00:abcd:194:2::100c.41402: Flags [S.], cksum 0x5383 (correct), seq 1040071206, ack 1007237652, win 65535, options [mss 388,nop,wscale 6,sackOK,TS val 3142440466 ecr 1953824165], length 0
```
- (look for the "mss <nnn>" setting there).
+ (look for the `mss <nnn>` option there).
The outgoing SYN is seen "before OpenVPN could touch it", so the client normally asks for a MSS of 1440 - which translates to "I do not want to receive TCP packets with a TCP segment size >1440" - TCP segment size being the payload of a TCP packet, so with TCP header (20) + IPv6 header (40) added, this is "1500 byte inside packet size" - which makes sense: "dear machine on the other side, this is my interface MTU, and I *know* that I can not handle anything bigger". When OpenVPN sees this packet, and the configured `mssfix 500` setting, it will replace the option in the TCP SYN with `mss 388`, which translates to "on the inside packet, add TCP + IP header, and then add OpenVPN overhead, and the resulting UDP packet must not be larger than 500 bytes = `mssfix 500 mtu`.
We can not see the effect of `mssfix` on the client SYN (tcpdump happens before OpenVPN encapsulation, and OpenVPN encaps is where mssfix happens) - but we can see the effect on the SYN/ACK coming back, where the server specififies what it could handle. Both directions are reduced appropriately.
-
+
`mssfix` can be combined with `tun-mtu`, or used on its own - and it will nicely fix all fragment/packet size problems for TCP connections inside an OpenVPN tunnel. What it can not do is fix "non TCP" protocols (ping, UDP DNS, ...) because these have no signalling mechanism of this sort (except "ICMP packet too big").
-
+
The default in current OpenVPN versions (after 2.6.0) is `mssfix 1492 mtu` which translates to "make sure that no packet ever leaves the machine larger than 1492 bytes" - to take PPPoE connections with 1492 into account. How much inside MSS remains depends on IPv4, IPv6 (inside and outside), cipher overhead, etc. - but OpenVPN has learned to reliably do this math. Took us a while.
+ The big remaining problem is that `--mssfix` is not yet supported for the DCO implementations on Linux and FreeBSD (see https://community.openvpn.net/DataChannelOffload/Features). So if both ends use DCO and a tun-mtu of 1500, you end up with a TCP MSS of 1440, full-sized internal packets, and outside fragmentation. This can be worked around using FreeBSD `pf(4)` firewall rules or Linux `iptables/nftable` firewall rules, to achieve the same result - but OpenVPN will not install these rules for you. On Windows, the DCO driver is advanced enough to handle `--mssfix` ;-)
# OpenVPN over TCP
- (to be written)
+ So what happens with OpenVPN over TCP? In this case, OpenVPN is not sending outside "UDP packets" but there is one continuous "TCP session" between client and server, and "encapsulated OpenVPN packets" are just sent as a continous byte stream in that TCP session. To enable the receiver to recognize where one "encapsulated packet" ends and the next starts, each is prefixed with a 2-byte length field.
+ After OpenVPN hands the packet (small or large) to TCP, the outside TCP layer takes care of "producing packets", taking MSS between OpenVPN client and server into account (= a PPPoE router with MTU 1492 in between that manipulates MSS will thus affect OpenVPN client/server).
+
+ Here's a `ping -s 1542` again (TCP tunnel, so different inside IP)
+ ```
+ 12:03:55.998620 tun2 Out IP6 (flowlabel 0x56067, hlim 64, next-header ICMPv6 (58) payload length: 1460) fd00:abcd:194:1::100c > fd00:abcd:194:2::1: [icmp6 sum ok] ICMP6, echo request, id 15805, seq 1
+
+ 12:03:55.998688 enp3s0 Out IP (tos 0x0, ttl 64, id 12998, offset 0, flags [DF], proto TCP (6), length 1420)
+ 193.149.48.143.45782 > 199.102.77.82.51194: Flags [.], cksum 0x0c5c (incorrect -> 0xc2c4), seq 3035142372:3035143740, ack 3439945976, win 670, options [nop,nop,TS val 1060262734 ecr 3877527626], length 1368
+ 12:03:55.998690 enp3s0 Out IP (tos 0x0, ttl 64, id 12999, offset 0, flags [DF], proto TCP (6), length 210)
+ 193.149.48.143.45782 > 199.102.77.82.51194: Flags [P.], cksum 0x07a2 (incorrect -> 0xa274), seq 1368:1526, ack 1, win 670, options [nop,nop,TS val 1060262734 ecr 3877527626], length 158
+
+ 12:03:56.127660 enp3s0 In IP (tos 0x0, ttl 48, id 0, offset 0, flags [DF], proto TCP (6), length 52)
+ 199.102.77.82.51194 > 193.149.48.143.45782: Flags [.], cksum 0x612d (correct), seq 1, ack 1526, win 1033, options [nop,nop,TS val 3877538931 ecr 1060262734], length 0
+
+ 12:03:56.127661 enp3s0 In IP (tos 0x0, ttl 48, id 0, offset 0, flags [DF], proto TCP (6), length 1578)
+ 199.102.77.82.51194 > 193.149.48.143.45782: Flags [P.], cksum 0x0cfa (incorrect -> 0x3c8f), seq 1:1527, ack 1526, win 1035, options [nop,nop,TS val 3877538931 ecr 1060262734], length 1526
+
+ 12:03:56.127699 enp3s0 Out IP (tos 0x0, ttl 64, id 13000, offset 0, flags [DF], proto TCP (6), length 52)
+ 193.149.48.143.45782 > 199.102.77.82.51194: Flags [.], cksum 0x5c2b (correct), seq 1526, ack 1527, win 660, options [nop,nop,TS val 1060262863 ecr 3877538931], length 0
+
+ 12:03:56.127747 tun2 In IP6 (hlim 64, next-header ICMPv6 (58) payload length: 1460) fd00:abcd:194:2::1 > fd00:abcd:194:1::100c: [icmp6 sum ok] ICMP6, echo reply, id 15805, seq 1
+ ```
+
+ note that there are no fragments now, and in addition to the expected packets, we see "length 52" byte packets that consist of ACKs only (confirming delivery to the other side). Also, we see that the system is lying to me again - the "incoming ping reply" is displayed as "length 1578" and "checksum incorrect" - smart network cards can do "TCP segmentation offloading", so what you really see on the wire and what tcpdump is being presented may not be the same thing. So in doubt, do not believe anything, and tcpdump on a router in between...
+
+ Generally, TCP mode solves all outside fragmentation issues - but it adds overhead (instead of "one `recvmsg()` call to get a packet" we now need to look at the bytestream, the length, find where the packets are, and deal with half-received packets as well). Also, doing TCP inside a TCP tunnel can lead to most interesting performance issues if there is a bit of packet loss involved, and both layers start doing transmissions and congestion avoidance.
+
+ Also, TCP is not supported in FreeBSD DCO mode...
+
+
+ # Summary
+ The best aproach today is
+ - use UDP mode
+ - pick `tun-mtu` according to what makes sense in your use case (1420 or 1500)
+ - use `mssfix` (if using Linux or FreeBSD DCO, turn this on in your firewalls)
+ - ensure your packet path handles fragmentation, if possible ("at least on all firewalls under your control") for large non-TCP packets inside the VPN
0 1 2 3 4 5 6 7 8 9 0 1 2 3 4 5 6 7 8 9 0 1 2 3 4 5 6 7 8 9 0 1 2 3 4 5 6 7 8 9