Scan 365 Pronetwork troubleshooting, reviewed

Guide · PingPlotter

Find the hop where packet loss really starts — and when it's only ICMP deprioritisation

A step-by-step method for locating the hop where packet loss begins, telling real loss from router rate limiting, and building evidence an ISP will accept.

The ticket says “calls keep dropping” and somebody has already pasted a traceroute into it with a scary 40% next to hop 6. Before you escalate to the carrier, you need to answer two questions: does the loss actually reach the destination, and if so, which hop is the first one where it shows up and never goes away? This guide walks through that method using PingPlotter, with mtr as the command-line equivalent. The goal is not a pretty graph; it’s a piece of evidence that survives the first reply from an ISP’s support desk.

The one rule that decides everything

Loss that appears at a hop but does not continue to the final destination is almost never real loss. Routers forward transit traffic in hardware, but they answer ICMP TTL-exceeded messages (the replies traceroute-style tools depend on) on their control plane, which is often rate-limited or given low priority. The mtr man page says this directly: some routers give ICMP ECHO lower priority, so their reported reliability is “significantly lower than the actual reliability”.

So the pattern you are looking for is:

  • Loss starts at hop N,
  • every hop after N shows roughly the same or higher loss,
  • the final destination shows it too.

Hop N (or the link just before it) is where the problem begins. Anything else is noise.

Step 1: Pick the right target and the right vantage point

  1. Choose a target that the users actually talk to — the SIP provider’s edge, the SaaS endpoint, the branch firewall. Tracing to a random public DNS resolver proves little about the application path.
  2. Run the trace from a host on the affected segment, ideally wired. A laptop on congested Wi-Fi will put loss at hop 1 and hide everything upstream.
  3. If you can, run a second trace from a host that is not affected. Two traces that diverge at a specific hop are far more persuasive than one.

Step 2: Start a long-running trace

In PingPlotter, enter the target, leave the interval at the default 2.5 seconds for anything longer than an hour, and start tracing. For short reproductions (a user who can trigger the problem on demand), 1 second is fine — note that the Free edition limits continuous monitoring to 10 minutes and doesn’t offer the 1-second sample rate, so plan around that.

With mtr, the equivalent is a report run with enough cycles to be meaningful:

# 600 probes at 1 s intervals ≈ 10 minutes, wide hostnames, show IPs and ASNs
mtr -r -w -b -z -c 600 sip.example-provider.net > mtr-sip-$(date +%F-%H%M).txt

Ten probes, which is what a quick mtr -r gives you by default, is not enough samples to say anything about 1–2% loss. Aim for several hundred.

Step 3: Read the table from the destination backwards

Don’t start at hop 1. Start at the last row.

  1. Destination shows 0% loss? The path is delivering packets. Any intermediate loss is control-plane rate limiting. Stop here and look elsewhere (application, codec, jitter buffer, Wi-Fi).
  2. Destination shows loss? Walk upward until you find the first hop where loss is at or near the destination’s figure and every hop below it agrees. That’s your starting point.
  3. Check latency the same way. A jump in average RTT that persists to the destination marks a congested or long link; a jump that shows only on one hop and disappears is, again, slow ICMP generation.

A quick illustration of the difference (figures are hypothetical):

Hop  Host                 Loss%   Avg ms
 4   core1.isp.net         0.0     8.1
 5   agg3.isp.net         38.0    41.7    <- loss here...
 6   edge2.isp.net         0.0     9.4    <- ...but not after: ignore hop 5
 7   sip.provider.net      0.0     9.9
Hop  Host                 Loss%   Avg ms
 4   core1.isp.net         0.0     8.1
 5   agg3.isp.net          4.1    31.2    <- starts here
 6   edge2.isp.net         4.4    32.0
 7   sip.provider.net      4.3    32.5    <- and reaches the target: real

Step 4: Correlate with time

A single summary table hides the most useful information: when. PingPlotter’s timeline graph shows whether loss is constant (often a physical-layer or duplex problem), clustered at business hours (congestion), or tied to a specific minute every hour (a scheduled job, a backup, a flapping link). Mark the times users reported problems and see if they line up.

Step 5: Try a different protocol before you blame the path

Some networks treat ICMP differently from real application traffic. If ICMP looks clean but users still complain, or ICMP looks awful but the application is fine, repeat the trace with TCP or UDP probes. PingPlotter exposes this in Engine Options; in mtr:

# TCP SYN probes to the HTTPS port
sudo mtr -r -w -T -P 443 -c 300 app.example.com

# UDP probes to a SIP port
sudo mtr -r -w -u -P 5060 -c 300 sip.example-provider.net

If loss disappears with TCP probes, you may be looking at an ICMP policer rather than a lossy link.

Step 6: Package the evidence

An ISP ticket should contain:

  • source public IP and the target,
  • start/end time with time zone,
  • the full hop table with sample count,
  • a timeline image covering the incident window,
  • ideally a second trace from another vantage point or the reverse direction.

PingPlotter can save the session as a .pp2 file, export CSV, save images and (in paid editions) produce shareable reports. mtr’s -j (JSON) or -C (CSV) output drops cleanly into a ticket or a spreadsheet. Remember that the path back to you may differ from the path out; ask the provider for a reverse trace if the loss starts at their border.

Common mistakes

  • Escalating on a mid-path spike. The single most common false alarm. If the destination is clean, the path is clean.
  • Too few samples. Ten or twenty probes can’t separate 1% loss from luck.
  • Tracing over Wi-Fi. Radio loss at hop 1 masks everything else. Use a cable, or trace from the wired gateway as well.
  • Ignoring your own edge. Loss starting at hop 1 or 2 is your LAN, firewall, or uplink. Check interface error counters and duplex before calling anyone.
  • Using sub-second intervals on someone else’s routers. Aggressive intervals trigger more rate limiting and make results worse, not better.
  • Forgetting the return path. A problem in the reverse direction still shows up as loss at the destination, but the hop list won’t show where. A trace from the far end, or a monitoring agent there, closes the gap.

Where to go next

If the problem is intermittent and you need weeks of history, a single trace session isn’t the right tool; look at continuous monitors such as SmokePing or Obkio, compared in SmokePing vs Obkio. If you’re deciding between the two tools used above, read PingPlotter vs mtr, and the full PingPlotter review and mtr review. The wider field is in Latency & Path Analysis. When loss is ruled out and the complaint is “slow”, measure capacity instead with iPerf3 throughput testing.

Get any of these tools only from the vendor or project site — see where to get the tools safely.

Tool used in this guide