What this lab does
It asks one question of a real tailnet: for every peer, is the connection direct or relayed, and why. Nothing is changed, nothing is broken, and no configuration is touched. Every command here is read only, which makes this the right first lab and the correct first move on a network you did not build.
It is worth doing even when nothing appears wrong. The tailnet that produced the output below looked completely healthy from the outside. Every machine was reachable, every service worked, nobody had filed a complaint. The census found that only same LAN peers were taking direct paths and every other connection in the fleet was riding a relay.
Prerequisites
- Tailscale running on the machine you are sitting at, and at least a handful of peers online.
- Nothing else. No sudo, no configuration changes, no downtime.
Run it
Step 1: count the paths
The status output in JSON is the fastest way to see the whole fleet at once. A peer that is Active with a CurAddr has a direct path; Active with an empty CurAddr is relayed; everything else is idle, meaning no session is currently established either way.
tailscale status --json | python3 -c "
import json,sys
d=json.load(sys.stdin)
peers=d.get('Peer') or {}
direct=relay=idle=0
for p in peers.values():
if not p.get('Online'): continue
if p.get('Active'):
if p.get('CurAddr'): direct+=1
else: relay+=1
else: idle+=1
print('direct=%d relay=%d idle=%d' % (direct,relay,idle))"
On the tailnet under test, with 47 peers and 25 online:
direct=3 relay=0 idle=22
That looks excellent, and it is misleading. Only three sessions were active at that moment, and all three happened to be machines on the same physical LAN. The twenty two idle peers had no session at all, so they contributed nothing to the count. An idle peer is not evidence of anything. To learn what path a peer would take, you have to make one.
Step 2: establish a session and watch the path get chosen
tailscale ping is the instrument. By default it stops as soon as it gets a direct pong, which is convenient for a quick check and useless for watching the sequence. Turn that off so it keeps going and shows you the whole story.
tailscale ping --c 6 --until-direct=false cloud-1
A connection that behaves the way the documentation describes looks like this: the first pongs come back through a relay while both sides work on a direct path, then the path upgrades and latency drops.
What came back instead, every time, for every remote machine:
pong from cloud-1 (100.64.0.29) via DERP(nyc) in 64ms
pong from cloud-1 (100.64.0.29) via DERP(nyc) in 57ms
pong from cloud-1 (100.64.0.29) via DERP(nyc) in 51ms
pong from cloud-1 (100.64.0.29) via DERP(nyc) in 52ms
pong from cloud-1 (100.64.0.29) via DERP(nyc) in 51ms
pong from cloud-1 (100.64.0.29) via DERP(nyc) in 52ms
Six pongs, no upgrade. Extending the window to fifteen pings changed nothing. The status line for that peer settled at active; relay "nyc" and stayed there.
Meanwhile a peer on the same physical LAN:
pong from node-b (100.64.0.38) via 10.0.0.200:41641 in 36ms
Direct immediately, over the local address, never touching a relay. That contrast is the finding: the tailnet was not failing to establish direct paths in general, it was failing to establish them to anything that was not already on the same wire.
Step 3: prove it is systemic, not one bad peer
One relayed peer is a peer problem. Every remote peer relaying is a network property. Check several, across different providers, so a single provider’s behavior cannot explain it.
for h in cloud-1 lab-vm-1 node-b; do
printf "%-12s " "$h"
tailscale ping --c 3 --until-direct=false "$h" 2>&1 | tail -1
done
Three different cloud hosts, two different hosting providers, all relayed through the same region. One LAN host, direct. The pattern held without exception.
Step 4: read the NAT fingerprint on both ends
netcheck is the report that says whether direct paths are even possible from where you are standing. Run it locally, then run it on the far end, because a direct path needs both.
tailscale netcheck
Local machine:
* UDP: true
* IPv4: yes, 198.51.100.14:62517
* MappingVariesByDestIP: false
* PortMapping:
* Nearest DERP: Dallas
The far end, a cloud host at a different provider:
* UDP: true
* IPv4: yes, 203.0.113.9:37737
* MappingVariesByDestIP: false
* PortMapping:
* Nearest DERP: New York City
Read those two reports carefully, because they are the reason this lab is interesting. UDP: true on both. MappingVariesByDestIP: false on both, which is the easy NAT case: the mapping a host gets is the same regardless of who it is talking to, which is the condition hole punching is designed for. Two easy NATs with working UDP is the configuration that should produce a direct path.
It did not.
Step 5: eliminate your own machine as the variable
The obvious suspect is the machine you are sitting at. Eliminate it by taking yourself out of the path entirely: ask two remote peers to talk to each other, and see what they choose.
ssh lab-vm-1 'tailscale ping --c 5 --until-direct=false cloud-1 | tail -3'
pong from cloud-1 (100.64.0.29) via DERP(nyc) in 2ms
pong from cloud-1 (100.64.0.29) via DERP(nyc) in 2ms
pong from cloud-1 (100.64.0.29) via DERP(nyc) in 2ms
Two remote hosts, neither of them the machine running the census, still relayed. Note the 2ms: both sit near the same relay region, so the relay is fast enough that no user would ever complain. It is still a relay.
That result rules out the local machine as the sole cause and turns a suspicion into a fleet property.
Step 6: check the far end is not simply firewalled
Before theorizing, check the boring explanation. On a host you control:
sudo ufw status
ss -lunp | grep 41641
Status: inactive
UNCONN 0 0 0.0.0.0:41641 0.0.0.0:*
No host firewall, and the daemon is listening on all addresses on the documented port. The host is not blocking anything, which pushes the cause upstream of the host, to the provider network.
What the census found
Stated honestly, including the part that is still open:
- Every connection in the fleet that was not between two machines on the same LAN was relayed, and stayed relayed.
- This held across two unrelated hosting providers and held between two remote hosts with the local machine removed from the path.
- Both sampled endpoints report
UDP: trueand endpoint independent mapping, which is the configuration where direct paths are expected to work. - Neither sampled host runs a firewall that would explain it, and the daemon listens on the documented port.
- Therefore the cause is upstream of the hosts, in one or more provider networks, and identifying it per pair needs evidence this lab does not collect: a packet capture on both ends of a single pair, taken at the same time, to see whether probes leave one side and arrive at the other.
What that means operationally is concrete: every byte between the workstation and every cloud host in this fleet crosses a relay. It works, it has always worked, and it is subject to relay bandwidth limits and an extra network leg. Nobody had noticed, because nothing was broken.
What to do with the result
- If your fleet is mostly relayed and you did not know, that alone is worth the twenty minutes. Latency and throughput to those hosts are worse than they need to be, and a relay outage becomes a fleet outage rather than a degradation.
- Where the far end is a cloud host whose provider filters inbound UDP, a self hosted relay inside your own infrastructure is the designed answer, and it moves the relay hop from a shared service to hardware you control.
- Where a host genuinely can accept inbound UDP, opening the documented port and confirming with a fresh census is the cheapest possible fix.
- Re-run this census after any network change. It is read only, it takes minutes, and it is the difference between believing your tailnet is direct and knowing.
Cross references
Module 03 explains the mechanism this lab measures, including what each netcheck line implies about the paths a node can form. Module 11 turns these same commands into a symptom driven playbook. The relay stuck drill works a constructed version of exactly this finding, and reading it next is the natural follow on.