Blame

78c13c cron2 2025-06-30 09:32:24 1
# MTU and Fragments
2
3
This page is an attempt to explain when and why OpenVPN is plagued by fragments, and what can be done about it.
4
It's not actually specific to OpenVPN, but to any tunnel technology - IPSEC, for example, has the same problems (and vendors mostly have same solutions).
5
It's also not specific to IPv4 or IPv6 - both protocols have fragments, and even if there is a common misunderstanding that "fragmentation is not allowed for IPv6" this only applies to routers (= systems that are tempted to fragment packets they have not created themselves). End systems can fragment all they want.
6
7
# Why fragments: OpenVPN over UDP
8
9
So, a short summary what is happening:
10
- there is a "tun" interface that has an interface MTU of 1500 bytes.
11
- A program like "ping" sends a packet that is 1500 byte large (because that's what the interface permits).
12
- OpenVPN adds some bytes to it (for encryption, authentication)
13
- OpenVPN sends this as an UDP packet, which adds 8 bytes for UDP + 20 or 40 bytes for the IPv4/IPv6 header
14
- the resulting packet is (no matter how many bytes are added) "larger than 1500 bytes"
15
- the system now wants to send this resulting UDP packet to the OpenVPN server, and under normal conditions this will be sent over an Ethernet interface with a MTU of 1500
f7ca5a cron2 2025-06-30 17:16:32 16
- "send a packet larger than the MTU of the egress interface" is a well-defined scenario, and so "IP Fragmentation" is used - that is, on the IP layer, the UDP packet OpenVPN has sent is split into 2 smaller IP packets - sometimes "half:half", sometimes "1500:rest", both is allowed
78c13c cron2 2025-06-30 09:32:24 17
- theoretically these packets could also hit a router with a PPPoE MTU of 1492 bytes next, splitting a 1500 byte fragment *again*
18
19
## tcpdump examples
20
21
Here is now this looks like in tcpdump (Linux `tcpdump -i any "host $inside or host $outside"`) - using IPv6 inside, IPv4 outside, so it's a bit easier to see "what goes where" - but it looks basically the same for IPv4-over-IPv4 or IPv6-over-IPv6
22
23
### small packet in the tunnel ("ping -s 64 fd00:abcd:194:2::1")
24
```
25
09:55:17.887555 tun2 Out IP6 fd00:abcd:194:2::100c > fd00:abcd:194:2::1: ICMP6, echo request, id 29334, seq 1, length 72
26
09:55:17.887663 enp3s0 Out IP 193.149.48.143.46220 > 199.102.77.82.51194: UDP, length 136
27
09:55:18.015810 enp3s0 In IP 199.102.77.82.51194 > 193.149.48.143.46220: UDP, length 136
28
09:55:18.015951 tun2 In IP6 fd00:abcd:194:2::1 > fd00:abcd:194:2::100c: ICMP6, echo reply, id 29334, seq 1, length 72
29
```
30
31
note how everything related to "size foo" always lies to you - so "ping -s 64" will send a ping packet with a 64 byte payload, and then there's the icmp header size of 8 bytes added to it - leading to a 72 byte ICMP payload "inside", but the actual packet is 112 bytes in size (40 byte IPv6 header, 8 byte ICMP header, 64 byte "ping" = 112) - tcpdump by default only prints the IP payload size ("72"), which is really non-helpful when trying to understand what happens.
32
33
`wireshark` does this in a better way, but the output is too long to paste in a nice form into a wiki page...
34
35
Let's do this again with `tcpdump -vv ...`
36
```
37
09:58:19.091512 tun2 Out IP6 (flowlabel 0xa3885, hlim 64, next-header ICMPv6 (58) payload length: 72) fd00:abcd:194:2::100c > fd00:abcd:194:2::1: [icmp6 sum ok] ICMP6, echo request, id 29771, seq 1
38
09:58:19.091629 enp3s0 Out IP (tos 0x0, ttl 64, id 49946, offset 0, flags [DF], proto UDP (17), length 164)
39
193.149.48.143.46220 > 199.102.77.82.51194: UDP, length 136
40
09:58:19.220412 enp3s0 In IP (tos 0x0, ttl 48, id 17738, offset 0, flags [none], proto UDP (17), length 164)
41
199.102.77.82.51194 > 193.149.48.143.46220: UDP, length 136
42
09:58:19.220522 tun2 In IP6 (hlim 64, next-header ICMPv6 (58) payload length: 72) fd00:abcd:194:2::1 > fd00:abcd:194:2::100c: [icmp6 sum ok] ICMP6, echo reply, id 29771, seq 1
43
```
28926b cron2 2025-06-30 17:20:09 44
... so there's more length fields here (though still not more useful for the inside ICMPv6 packet). For IPv4 UDP, one can now see that the "UDP, length 136" (second line) is "the UDP payload", to which +8 (UDP header) +20 (IPv4 header) are added to then result in `length 164` for the real packet being sent to the network (first line).
45
For IPv4, when running `tcpdump -vv`, the first line printed for each packet will show the overall size correctly.
78c13c cron2 2025-06-30 09:32:24 46
47
This example has small packets, so no fragmentation, and no complications:
48
- We see one(1) ping packet go "tun2 Out" (so this is something "we send").
49
- This gets turned into one(1) UDP packet, going "enp3s0 Out" - this is the LAN interface here, sending the packet to the OpenVPN server.
50
- one(1) UDP packet comes back "enp3s0 In" ("In" = "we receive", from the Internet)
51
- OpenVPN decapsulates the packet and gives it back to the tun interface, so the next is
52
- one(1) IPv6 ping packet coming "tun2 In" -> our ping reply
53
54
### large packet in the tunnel ("ping -s 1452")
55
28926b cron2 2025-06-30 17:20:09 56
With the math above, we now know that `ping -s 1452 v6host` will now create a 1500 byte IPv6 packet "inside the tunnel" (1452 + 8 byte ICMP + 40 byte IPv6 header = 1500)
78c13c cron2 2025-06-30 09:32:24 57
58
```
59
10:12:15.432260 tun2 Out IP6 (flowlabel 0xa3885, hlim 64, next-header ICMPv6 (58) payload length: 1460) fd00:abcd:194:2::100c > fd00:abcd:194:2::1: [icmp6 sum ok] ICMP6, echo request, id 31596, seq 1
60
10:12:15.432390 enp3s0 Out IP (tos 0x0, ttl 64, id 39726, offset 0, flags [+], proto UDP (17), length 1500)
61
193.149.48.143.46220 > 199.102.77.82.51194: UDP, length 1524
62
10:12:15.432404 enp3s0 Out IP (tos 0x0, ttl 64, id 39726, offset 1480, flags [none], proto UDP (17), length 72)
63
193.149.48.143 > 199.102.77.82: ip-proto-17
64
10:12:15.562396 enp3s0 In IP (tos 0x0, ttl 48, id 17921, offset 0, flags [+], proto UDP (17), length 1492)
65
199.102.77.82.51194 > 193.149.48.143.46220: UDP, length 1524
66
10:12:15.562396 enp3s0 In IP (tos 0x0, ttl 48, id 17921, offset 1472, flags [+], proto UDP (17), length 28)
67
199.102.77.82 > 193.149.48.143: ip-proto-17
68
10:12:15.562396 enp3s0 In IP (tos 0x0, ttl 48, id 17921, offset 1480, flags [none], proto UDP (17), length 72)
69
199.102.77.82 > 193.149.48.143: ip-proto-17
70
10:12:15.562510 tun2 In IP6 (hlim 64, next-header ICMPv6 (58) payload length: 1460) fd00:abcd:194:2::1 > fd00:abcd:194:2::100c: [icmp6 sum ok] ICMP6, echo reply, id 31596, seq 1
71
```
72
73
So this what we can observe
74
- one ICMPv6 ping packet "tun2 Out"
75
- *two* UDP packets going "enp3s0 Out" to the VPN server
28926b cron2 2025-06-30 17:20:09 76
- one is decoded as "UDP length 1524" (that is the UDP payload!), and "length 1500" on the wire - so parts of the packet are missing
77
- the second one is decoded as "length 72" on the wire, and "ip-proto-17" inside - this is the rest
78c13c cron2 2025-06-30 09:32:24 78
- when calculating "how many fragments and how big?" it needs to be taken into account that each fragment has its own IP header, so the "on the wire" sum of both fragments is larger than the "UDP payload 1524 + 1x UDP header + 1x IPv4 header" (1552) would lead you to expect
79
- *three* UDP packets coming in "enp3s0 In" from the Internet
80
- one 1492 byte packet
81
- one 28 byte packet
82
- so it seems there was a router with PPPoE and 1492 MTU on the way to me, splitting the expected "first packet is 1500 byte on the wire" into "1492 + rest" - 8 byte remaining + IP header = 28 byte on the wire
83
- one 72 byte packet
84
- the linux kernel dutifully reassembles all these fragments into one (big) UDP packet, handed to OpenVPN (= so OpenVPN never sees these fragments)
85
- after decapsulation in OpenVPN, we see
86
- *one* ICMPv6 packet "tun2 In", which is our "echo reply", and by convention it has the same size as the "echo request"
87
88
### huge packet in the tunnel ("ping -s 3000")
89
90
Part of the client side stress test suite (t_client) is pinging remote IPs with -s 3000, which creates a lot of packets inside and outside - this is important to test for MTU mismatches on the tun interface (so if the client has a MTU of 1500 and the server has a MTU of 1400, the idea of a "full size" packet is different, and this fairly reliably uncovers bugs). Also, this means "a fragmented packet is handed to the IP stack to be fragmented *again*, which uncovered a bunch of bugs in the Linux DCO implementation...
91
92
```
93
10:24:29.450403 tun2 Out IP6 (flowlabel 0xa3885, hlim 64, next-header Fragment (44) payload length: 1456) fd00:abcd:194:2::100c > fd00:abcd:194:2::1: frag (0x5793f87d:0|1448) ICMP6, echo request, id 658, seq 1
94
10:24:29.450407 tun2 Out IP6 (flowlabel 0xa3885, hlim 64, next-header Fragment (44) payload length: 1456) fd00:abcd:194:2::100c > fd00:abcd:194:2::1: frag (0x5793f87d:1448|1448)
95
10:24:29.450408 tun2 Out IP6 (flowlabel 0xa3885, hlim 64, next-header Fragment (44) payload length: 120) fd00:abcd:194:2::100c > fd00:abcd:194:2::1: frag (0x5793f87d:2896|112)
96
97
10:24:29.450458 enp3s0 Out IP (tos 0x0, ttl 64, id 9295, offset 0, flags [+], proto UDP (17), length 1500)
98
193.149.48.143.46220 > 199.102.77.82.51194: UDP, length 1520
99
10:24:29.450462 enp3s0 Out IP (tos 0x0, ttl 64, id 9295, offset 1480, flags [none], proto UDP (17), length 68)
100
193.149.48.143 > 199.102.77.82: ip-proto-17
101
10:24:29.450470 enp3s0 Out IP (tos 0x0, ttl 64, id 9296, offset 0, flags [+], proto UDP (17), length 1500)
102
193.149.48.143.46220 > 199.102.77.82.51194: UDP, length 1520
103
10:24:29.450471 enp3s0 Out IP (tos 0x0, ttl 64, id 9296, offset 1480, flags [none], proto UDP (17), length 68)
104
193.149.48.143 > 199.102.77.82: ip-proto-17
105
10:24:29.450477 enp3s0 Out IP (tos 0x0, ttl 64, id 9297, offset 0, flags [DF], proto UDP (17), length 212)
106
193.149.48.143.46220 > 199.102.77.82.51194: [udp sum ok] UDP, length 184
107
108
10:24:29.580873 enp3s0 In IP (tos 0x0, ttl 47, id 18128, offset 0, flags [none], proto UDP (17), length 212)
109
199.102.77.82.51194 > 193.149.48.143.46220: [udp sum ok] UDP, length 184
110
10:24:29.580932 enp3s0 In IP (tos 0x0, ttl 48, id 18126, offset 0, flags [+], proto UDP (17), length 1492)
111
199.102.77.82.51194 > 193.149.48.143.46220: UDP, length 1520
112
10:24:29.580932 enp3s0 In IP (tos 0x0, ttl 48, id 18126, offset 1472, flags [+], proto UDP (17), length 28)
113
199.102.77.82 > 193.149.48.143: ip-proto-17
114
10:24:29.580932 enp3s0 In IP (tos 0x0, ttl 48, id 18126, offset 1480, flags [none], proto UDP (17), length 68)
115
199.102.77.82 > 193.149.48.143: ip-proto-17
116
117
10:24:29.581000 tun2 In IP6 (hlim 64, next-header Fragment (44) payload length: 120) fd00:abcd:194:2::1 > fd00:abcd:194:2::100c: frag (0x5e590dc5:2896|112)
118
10:24:29.581032 tun2 In IP6 (hlim 64, next-header Fragment (44) payload length: 1456) fd00:abcd:194:2::1 > fd00:abcd:194:2::100c: frag (0x5e590dc5:0|1448) ICMP6, echo reply, id 658, seq 1
119
10:24:29.581626 enp3s0 In IP (tos 0x0, ttl 48, id 18127, offset 0, flags [+], proto UDP (17), length 1492)
120
121
199.102.77.82.51194 > 193.149.48.143.46220: UDP, length 1520
122
10:24:29.581627 enp3s0 In IP (tos 0x0, ttl 48, id 18127, offset 1472, flags [+], proto UDP (17), length 28)
123
199.102.77.82 > 193.149.48.143: ip-proto-17
124
10:24:29.581627 enp3s0 In IP (tos 0x0, ttl 48, id 18127, offset 1480, flags [none], proto UDP (17), length 68)
125
199.102.77.82 > 193.149.48.143: ip-proto-17
126
127
10:24:29.581686 tun2 In IP6 (hlim 64, next-header Fragment (44) payload length: 1456) fd00:abcd:194:2::1 > fd00:abcd:194:2::100c: frag (0x5e590dc5:1448|1448)
128
```
129
I'm not going to explain these in detail - but looking closely it's possible to figure out what is happening, and why ;-)
130
131
One important thing to see in this dump is that we are seeing multiple different "layers of fragments"
132
- "inside fragments" - this is user data that is too big for the OpenVPN tun interface, and is fragmented *before* going into OpenVPN (so it's "fragments inside of the OpenVPN tunnel")
133
- "outside fragments" - these are packets leaving my machine, going to the OpenVPN server ("fragments are seen outside my system")
134
135
## why is this a problem?
136
137
In principle, this works, as can be seen from the tcpdump examples. But sometimes it doesn't
138
139
- there are NAT routers that choke on fragments, and NAT only some parts and drop others, or produce broken headers
140
- there are firewalls that consider fragments to be "AN ATTACK! MUST THROW AWAY!"
141
- there are firewalls that drop everything that is not explicitly allowed, and users tend to forget about fragments -> drop
142
- there are ISPs that rate-limt fragmented packets, because "over the wide Internet" they tend to be mostly attack traffic these days (reflection attacks against UDP based services that can be prompted to send huge replies to a 3rd party victim address)
143
- fragments create extra load on the receiving end (because first all the fragments need to be reassembled before handing them as "one big UDP packet" to the receiving application - here, OpenVPN) - so this limits throughput and wastes energy
144
145
## how to avoid fragments?
146
There are a number of approaches, none of which are 100% satisfying
147
148
### --compress
34fb88 cron2 2025-06-30 17:23:12 149
This is not as magic as we hope it to be - for some packets ("ping packet filled with 0") it will do a good job in making the resulting OpenVPN packet actually *smaller* than the incoming ICMP packet (avoiding outside fragmentation). For the packets normally seen on a VPN, which are like "compressed pictures downloaded by a web browser", compression will actually *add* a bit of overhead, so it is not fixing the fragmentation problem - and it adds its own problems.
150
Do not use `--compress` these days.
78c13c cron2 2025-06-30 09:32:24 151
152
### --fragment
153
This is a switch to OpenVPN (`--fragment 1300`) which adds a *middle fragment* layer to the already-complicated picture.
154
With this, OpenVPN will never send a packet with a larger payload than 1300 (+ UDP + IPv4/IPv6, so 1328/1348-ish overall packet size).
155
If a larger packet is coming in via the tun, fragmentation happens in the OpenVPN protocol, and no "outside fragments" are ever observed.
156
157
This does work, but it is not supported if using kernel offloading (DCO), because it would require adding all the `--fragment` handling code (+reassembly) to the kernel layer - and the intent of the kernel implementations is "implement only the most common use cases, to keep the code size and complexity low".
158
The benefit of this is that there is no dealing with "outside fragments", so all the firewall/rate-limiting reasons for dropping are no longer a problem - but the extra overhead for fragment reassembly on the receiving end is still valid (reassembly in openvpn userland instead of kernel IP handler, but the same principle - energy, memory, CPU).
159
34fb88 cron2 2025-06-30 17:23:12 160
### --tun-mtu
78c13c cron2 2025-06-30 09:32:24 161
Running OpenVPN with `--tun-mtu 1400` (e.g.) will create a "tun" interface with an interface MTU of less-than 1500 bytes - in this case, 1400 bytes. So no inside packets larger than 1400 byte can happen, because the IP stack *before* OpenVPN takes care of this (and `ping -s 1452` would actually see 2 inside packets in this scenario, one 1400 byte fragment and one with the rest).
162
163
The actual overhead "how many bytes will be added by OpenVPN to the inside packet?" depends on a number of factors - cipher and auth hash used, IPv4 or IPv6 transport on the outside, but as a rule of thumb it's something like 24+8+40 for "AES-GCM, UDP, IPv6" = 72 byte.
164
165
This said, if you configure your inside MTU to 1400 byte, and the typical overhead is no more than 72 bytes, UDP packets generated by OpenVPN will never be larger than 1492/1500 byte, and you will never see outside fragmentation. This is generally good...
166
167
The drawback in this scenario are MTU jumps, so in a scenario like this
168
```
169
Client PC --(LAN/1500)--> OpenVPN Router --(tun/1400)--> OpenVPN Server --(LAN/1500)--> Server PC
170
```
171
172
the end systems ("Client PC" and "Server PC") will not be aware that there is a 1400 piece in the middle, and will attempt to send packets on what they think "a full size packet is allowed to be" = 1500 byte. The OpenVPN boxes can now either create "internal fragment" packets on behalf of Client/Server PC (which is allowed for IPv4 and explicitly disallowed for IPv6), or send back "ICMP packet too big" packets to the sender, telling the sending side that "here's an MTU jump, please do not send packets larger than 1400 byte". This "mostly" works, except when it does not, for example because a stupid firewall in the way throws away all ICMP packets ("this is evil hacker tool stuff").
173
174
Where this scenario works well is in the "roadwarrior" case, where you do not route 3rd parties, but only have clients connected to a server
175
```
176
Client PC with OpenVPN --(tun/1400) --> OpenVPN Server --(LAN/1500) --> Server PC
177
```
178
because in this case the Client knows "my maximum allowed packet size is 1400", and will never send more - and for TCP sessions to "Server PC" it will actually signal (via TCP MSS option) that it is not willing to receive TCP packets that would be larger than 1400 bytes overall. So for the most common case ("TCP") this will work perfectly.
179
180
The OpenVPN AS product defaults to `tun-mtu 1420` and pushes that to the OpenVPN Clients, because for that scenario it's more likely to work in most cases. The community OpenVPN defaults to `tun-mtu 1500` because this avoids MTU jumps in more complex scenarios.
181
9b7f54 cron2 2025-06-30 10:20:07 182
Also, `tun-mtu` is not going to work if you want ethernet briding (tap mode) where you need the tap interface MTU to be the same as the LAN interface (no "packet too big" handling on the ethernet layer). This is somewhat of a niche use case, but it is something where OpenVPN can come in handy - and we need to be aware of the limitations.
183
34fb88 cron2 2025-06-30 17:23:12 184
### --mssfix
78c13c cron2 2025-06-30 09:32:24 185
If `tun-mtu` can not be changed, or does not work correctly due to ICMP packet too big getting lost, there is another feature in OpenVPN which makes it "work for most cases", `--mssfix`, for example setting `--mssfix 500 mtu`. What this does is to look at TCP SYN and SYN-ACK packets passing through OpenVPN, and modifying the "MSS value" in there
186
187
```
188
11:20:34.497750 tun2 Out IP6 (flowlabel 0xce5b3, hlim 64, next-header TCP (6) payload length: 40) fd00:abcd:194:2::100c.41402 > fd00:abcd:194:2::1.53: Flags [S], cksum 0x5bf7 (correct), seq 1007237651, win 64800, options [mss 1440,sackOK,TS val 1953824165 ecr 0,nop,wscale 7], length 0
189
11:20:34.626186 tun2 In IP6 (flowlabel 0xe3a26, hlim 64, next-header TCP (6) payload length: 40) fd00:abcd:194:2::1.53 > fd00:abcd:194:2::100c.41402: Flags [S.], cksum 0x5383 (correct), seq 1040071206, ack 1007237652, win 65535, options [mss 388,nop,wscale 6,sackOK,TS val 3142440466 ecr 1953824165], length 0
190
```
191
9b7f54 cron2 2025-06-30 10:20:07 192
(look for the `mss <nnn>` option there).
78c13c cron2 2025-06-30 09:32:24 193
194
The outgoing SYN is seen "before OpenVPN could touch it", so the client normally asks for a MSS of 1440 - which translates to "I do not want to receive TCP packets with a TCP segment size >1440" - TCP segment size being the payload of a TCP packet, so with TCP header (20) + IPv6 header (40) added, this is "1500 byte inside packet size" - which makes sense: "dear machine on the other side, this is my interface MTU, and I *know* that I can not handle anything bigger". When OpenVPN sees this packet, and the configured `mssfix 500` setting, it will replace the option in the TCP SYN with `mss 388`, which translates to "on the inside packet, add TCP + IP header, and then add OpenVPN overhead, and the resulting UDP packet must not be larger than 500 bytes = `mssfix 500 mtu`.
195
We can not see the effect of `mssfix` on the client SYN (tcpdump happens before OpenVPN encapsulation, and OpenVPN encaps is where mssfix happens) - but we can see the effect on the SYN/ACK coming back, where the server specififies what it could handle. Both directions are reduced appropriately.
9b7f54 cron2 2025-06-30 10:20:07 196
78c13c cron2 2025-06-30 09:32:24 197
`mssfix` can be combined with `tun-mtu`, or used on its own - and it will nicely fix all fragment/packet size problems for TCP connections inside an OpenVPN tunnel. What it can not do is fix "non TCP" protocols (ping, UDP DNS, ...) because these have no signalling mechanism of this sort (except "ICMP packet too big").
9b7f54 cron2 2025-06-30 10:20:07 198
78c13c cron2 2025-06-30 09:32:24 199
The default in current OpenVPN versions (after 2.6.0) is `mssfix 1492 mtu` which translates to "make sure that no packet ever leaves the machine larger than 1492 bytes" - to take PPPoE connections with 1492 into account. How much inside MSS remains depends on IPv4, IPv6 (inside and outside), cipher overhead, etc. - but OpenVPN has learned to reliably do this math. Took us a while.
200
9b7f54 cron2 2025-06-30 10:20:07 201
The big remaining problem is that `--mssfix` is not yet supported for the DCO implementations on Linux and FreeBSD (see https://community.openvpn.net/DataChannelOffload/Features). So if both ends use DCO and a tun-mtu of 1500, you end up with a TCP MSS of 1440, full-sized internal packets, and outside fragmentation. This can be worked around using FreeBSD `pf(4)` firewall rules or Linux `iptables/nftable` firewall rules, to achieve the same result - but OpenVPN will not install these rules for you. On Windows, the DCO driver is advanced enough to handle `--mssfix` ;-)
78c13c cron2 2025-06-30 09:32:24 202
203
# OpenVPN over TCP
204
9b7f54 cron2 2025-06-30 10:20:07 205
So what happens with OpenVPN over TCP? In this case, OpenVPN is not sending outside "UDP packets" but there is one continuous "TCP session" between client and server, and "encapsulated OpenVPN packets" are just sent as a continous byte stream in that TCP session. To enable the receiver to recognize where one "encapsulated packet" ends and the next starts, each is prefixed with a 2-byte length field.
206
After OpenVPN hands the packet (small or large) to TCP, the outside TCP layer takes care of "producing packets", taking MSS between OpenVPN client and server into account (= a PPPoE router with MTU 1492 in between that manipulates MSS will thus affect OpenVPN client/server).
207
208
Here's a `ping -s 1542` again (TCP tunnel, so different inside IP)
209
```
210
12:03:55.998620 tun2 Out IP6 (flowlabel 0x56067, hlim 64, next-header ICMPv6 (58) payload length: 1460) fd00:abcd:194:1::100c > fd00:abcd:194:2::1: [icmp6 sum ok] ICMP6, echo request, id 15805, seq 1
211
212
12:03:55.998688 enp3s0 Out IP (tos 0x0, ttl 64, id 12998, offset 0, flags [DF], proto TCP (6), length 1420)
213
193.149.48.143.45782 > 199.102.77.82.51194: Flags [.], cksum 0x0c5c (incorrect -> 0xc2c4), seq 3035142372:3035143740, ack 3439945976, win 670, options [nop,nop,TS val 1060262734 ecr 3877527626], length 1368
214
12:03:55.998690 enp3s0 Out IP (tos 0x0, ttl 64, id 12999, offset 0, flags [DF], proto TCP (6), length 210)
215
193.149.48.143.45782 > 199.102.77.82.51194: Flags [P.], cksum 0x07a2 (incorrect -> 0xa274), seq 1368:1526, ack 1, win 670, options [nop,nop,TS val 1060262734 ecr 3877527626], length 158
216
217
12:03:56.127660 enp3s0 In IP (tos 0x0, ttl 48, id 0, offset 0, flags [DF], proto TCP (6), length 52)
218
199.102.77.82.51194 > 193.149.48.143.45782: Flags [.], cksum 0x612d (correct), seq 1, ack 1526, win 1033, options [nop,nop,TS val 3877538931 ecr 1060262734], length 0
219
220
12:03:56.127661 enp3s0 In IP (tos 0x0, ttl 48, id 0, offset 0, flags [DF], proto TCP (6), length 1578)
221
199.102.77.82.51194 > 193.149.48.143.45782: Flags [P.], cksum 0x0cfa (incorrect -> 0x3c8f), seq 1:1527, ack 1526, win 1035, options [nop,nop,TS val 3877538931 ecr 1060262734], length 1526
222
223
12:03:56.127699 enp3s0 Out IP (tos 0x0, ttl 64, id 13000, offset 0, flags [DF], proto TCP (6), length 52)
224
193.149.48.143.45782 > 199.102.77.82.51194: Flags [.], cksum 0x5c2b (correct), seq 1526, ack 1527, win 660, options [nop,nop,TS val 1060262863 ecr 3877538931], length 0
225
226
12:03:56.127747 tun2 In IP6 (hlim 64, next-header ICMPv6 (58) payload length: 1460) fd00:abcd:194:2::1 > fd00:abcd:194:1::100c: [icmp6 sum ok] ICMP6, echo reply, id 15805, seq 1
227
```
228
229
note that there are no fragments now, and in addition to the expected packets, we see "length 52" byte packets that consist of ACKs only (confirming delivery to the other side). Also, we see that the system is lying to me again - the "incoming ping reply" is displayed as "length 1578" and "checksum incorrect" - smart network cards can do "TCP segmentation offloading", so what you really see on the wire and what tcpdump is being presented may not be the same thing. So in doubt, do not believe anything, and tcpdump on a router in between...
230
231
Generally, TCP mode solves all outside fragmentation issues - but it adds overhead (instead of "one `recvmsg()` call to get a packet" we now need to look at the bytestream, the length, find where the packets are, and deal with half-received packets as well). Also, doing TCP inside a TCP tunnel can lead to most interesting performance issues if there is a bit of packet loss involved, and both layers start doing transmissions and congestion avoidance.
232
233
Also, TCP is not supported in FreeBSD DCO mode...
234
235
236
# Summary
237
The best aproach today is
238
- use UDP mode
239
- pick `tun-mtu` according to what makes sense in your use case (1420 or 1500)
240
- use `mssfix` (if using Linux or FreeBSD DCO, turn this on in your firewalls)
241
- ensure your packet path handles fragmentation, if possible ("at least on all firewalls under your control") for large non-TCP packets inside the VPN