What this lab does
Lab 01 found that an entire tailnet was relayed and could not say why. It ended by naming the measurement that would settle it: capture packets on both ends of one pair at the same time, and see whether the probes each side sends ever arrive at the other. This lab is that measurement.
It is still read only. A packet capture observes; it changes no configuration and breaks no path. It does need root on both machines, which is the only new requirement over lab 01.
The reason this specific measurement is worth the setup is that it collapses a large space of theories into one of four answers, and it does so with evidence rather than inference. Everything before this point in an investigation like this is narrowing. This is the step that decides.
How to read the result before you run it
Decide in advance what each outcome means, because it stops you from rationalizing whatever you happen to see. Two hosts, A and B. Each capture answers one question: did packets leave, and did packets arrive.
Run it
Step 1: pick the pair and learn both public addresses
Pick one pair, not several. A capture of many peers at once produces a file you will not read.
You need the address each host presents to the internet, because that is what the other side aims at and therefore what you filter on.
tailscale netcheck | grep IPv4
ssh lab-vm-1 'tailscale netcheck | grep IPv4'
For the pair under test, the local machine reported 198.51.100.14 and the remote reported 203.0.113.9.
Step 2: start both captures at the same time
Both captures must overlap, or you are comparing two different moments and proving nothing. Start them, give them a few seconds of headroom, then trigger the traffic.
Remote, in the background:
ssh lab-vm-1 'sudo timeout 20 tcpdump -ni any "udp and host 198.51.100.14" \
-w /tmp/lab02-remote.pcap' &
Local:
sudo timeout 20 tcpdump -ni en0 "udp and host 203.0.113.9" \
-w /tmp/lab02-local.pcap &
Note what the filters say. Each side captures all UDP to or from the other side’s public address. That is deliberately wider than port 41641, because if the far side is answering from an unexpected port you want to see it rather than filter it out. A filter that only matches what you expect can only ever confirm you.
Step 3: force a fresh attempt
sleep 4
tailscale ping --c 10 --until-direct=false lab-vm-1
pong from lab-vm-1 (100.64.0.41) via DERP(nyc) in 56ms
pong from lab-vm-1 (100.64.0.41) via DERP(nyc) in 52ms
pong from lab-vm-1 (100.64.0.41) via DERP(nyc) in 95ms
Relayed throughout, as expected. The pongs are not the data. The captures are.
Step 4: read the local capture
sudo tcpdump -nr /tmp/lab02-local.pcap
16:30:24.701719 IP 10.0.0.213.41641 > 203.0.113.9.4844: UDP, length 124
16:30:25.489252 IP 10.0.0.213.41641 > 203.0.113.9.4844: UDP, length 124
16:30:25.779617 IP 10.0.0.213.41641 > 203.0.113.9.4844: UDP, length 124
16:30:26.837448 IP 10.0.0.213.41641 > 203.0.113.9.4844: UDP, length 124
16:30:27.895454 IP 10.0.0.213.41641 > 203.0.113.9.4844: UDP, length 124
16:30:28.952941 IP 10.0.0.213.41641 > 203.0.113.9.4844: UDP, length 124
Thirteen packets in total, and the source of every single one is the local machine. Confirm that rather than eyeballing it:
sudo tcpdump -nr /tmp/lab02-local.pcap | awk '{print $3}' | cut -d. -f1-4 | sort | uniq -c
13 10.0.0.213
One source address, thirteen packets, zero inbound. The local host is trying, repeatedly, and nothing is coming back.
Step 5: read the remote capture
ssh lab-vm-1 'sudo tcpdump -nr /tmp/lab02-remote.pcap'
16:30:25.516855 eth0 Out IP 10.42.0.42.41641 > 198.51.100.14.3541: UDP, length 124
16:30:28.510728 eth0 Out IP 10.42.0.42.41641 > 198.51.100.14.3541: UDP, length 124
16:30:31.553775 eth0 Out IP 10.42.0.42.41641 > 198.51.100.14.3541: UDP, length 124
16:30:34.512342 eth0 Out IP 10.42.0.42.41641 > 198.51.100.14.3541: UDP, length 124
16:30:37.521684 eth0 Out IP 10.42.0.42.41641 > 198.51.100.14.3541: UDP, length 124
16:30:40.516328 eth0 Out IP 10.42.0.42.41641 > 198.51.100.14.3541: UDP, length 124
ssh lab-vm-1 'sudo tcpdump -nr /tmp/lab02-remote.pcap | awk "{print \$3}" | sort | uniq -c'
6 Out
Six packets, every one of them outbound, zero inbound. The remote host is also trying, and also hearing nothing.
What the capture proved
Against the table written before the run, this is the bottom left cell. Both sides send, neither receives.
That single result eliminates a list of suspects at once, and each elimination is evidence backed rather than assumed:
- Not the local daemon. It sent thirteen probes. It is attempting a direct path.
- Not the remote daemon. It sent six. Also attempting.
- Not endpoint discovery or the coordination plane. Each side is aiming at a specific mapped port on the other, which means STUN worked and the candidates were exchanged.
- Not a host firewall. Lab 01 had already shown no firewall on the remote host and the daemon listening on the documented port. The capture confirms it from the other direction: packets leave the host cleanly.
- Not a one directional filter. An asymmetric drop would show one side receiving. Neither does.
What remains is the only thing left: the packets are dropped between the two hosts, in both directions, by something neither host controls.
The thing nobody had noticed
Look again at the source address in the remote capture. It is 10.42.0.42, on a /16. That is a private address, and it is the VPS’s own interface address:
ssh lab-vm-1 'ip -4 addr show eth0 | grep inet'
inet 10.42.0.42/16 brd 10.42.255.255 scope global eth0
The cloud host does not have a public address attached to it. It sits behind its provider’s NAT, and the 203.0.113.9 it reports through netcheck is a mapping on that shared NAT, not an address it owns.
That matters because it changes the shape of the problem. This is not one NAT with a host behind it. It is two NATs facing each other, and the far one belongs to a hosting provider and is not configurable by the person running the VPS. Opening a port on the VPS cannot help, because the packets are not reaching the VPS to be filtered.
What to do about it
The honest answer is that no host side change fixes this pair, and recognizing that is the value of the measurement. Options in order of how much control they give you:
- Accept the relay. It works, it is the designed fallback, and for control plane traffic like SSH sessions and web dashboards the latency cost is unremarkable. Know that it is happening and that relay bandwidth limits apply.
- Run a relay you control. A peer relay is a node in your own tailnet that relays for peers when direct paths fail, which moves the hop off shared infrastructure and onto a machine whose capacity and location you choose. This is the designed answer for exactly this situation, and it requires that at least one node in your tailnet is actually reachable.
- Change the underlay. Move a host to a provider or plan that assigns a routable address, or obtain port forwarding on a NAT you control. This is the only route to a genuinely direct path here, and it is a procurement decision rather than a configuration one.
What you should not do is keep tuning host firewalls. The capture already proved that is not where the packets die.
Reproduce this on your own pair
The technique generalizes to any two hosts that will not connect directly:
- Get both public addresses from netcheck.
- Start overlapping captures on both ends filtered to the other side’s public address, wider than the port you expect.
- Force an attempt with
tailscale ping --until-direct=false. - Tally sent versus received on each side, with
awkrather than by eye. - Read the answer off the table at the top of this page, and check the far host’s own interface address while you are in there. A private address on a cloud host is a finding by itself.
Clean up the capture files when you are done. They contain your own network’s addressing.
Cross references
Module 03 explains hole punching and what each netcheck field does and does not tell you, which is the theory this lab tests against reality. Module 11 covers the symptom driven version of the same investigation. Lab 01 is the census that produced the question this lab answers. The relay stuck drill is the constructed version of this scenario, and comparing its tidy resolution with this one is instructive: the drill resolves, and the real network did not.