Getting packets off the wire is only half the work
The first thing most people get wrong about Performing Packet Capture And Traffic Analysis is that they think the capture itself is the hard part. It isn't. tcpdump will grab every packet you throw at it. The actual difficulty sits in making sense of what you captured once you have a 4-gigabyte pcap file open at 2 AM and you're still looking for a single TLS handshake that may or may not have dropped. I still use tcpdump for quick captures on Linux, but I reached a point years ago where I just wrap it in a small bash script that names the file with a timestamp, sets a size limit, and rotates old captures automatically. Something like tcpdump -i eth0 -w /tmp/capture_$(date +%Y%m%d_%H%M).pcap -G 300 -W 4 --snapshot-length 256. That gives you five-minute rotated files capped at 256 bytes per packet for quick triage, then you switch to full captures when you actually need to dig in. The rotation keeps your disk from filling up, and the snapshot length keeps the initial look-through fast without dropping the IP headers you actually need. On Windows, I tend toward Wireshark with a capture filter set before I hit the big green button. Capturing without a filter on a busy link is how you end up with a file so large you can't open it in Wireshark without swapping it to disk twice. A display filter like tcp port 443 and ip.addr == 10.0.0.0/8 keeps the file manageable from the start.
One thing nobody tells you about capture filters versus display filters: capture filters are applied by libpcap before the packet ever reaches your screen, so they actually reduce file size. Display filters only hide packets after they are already written to disk. I learned this the hard way during a production issue where I was capturing on a 1 Gbps uplink with no capture filter and wondering why my system kept dropping packets from the capture itself. The ring buffer was overflowing because the NIC was feeding packets faster than I could write them. Adding a capture filter cut the file size by roughly eighty percent and eliminated the kernel-level drop count I was seeing in the interface statistics. When you open a capture, the first column to check is the Ioana or drop count on the interface. If that number is climbing during your capture, your data is incomplete and any timing analysis you do is going to be wrong. A drop count in the low hundreds on a quiet lab link is fine. A drop count in the thousands on a production link means you are missing pieces of conversations and you should not build a conclusion on what you captured. For traffic analysis, the TShark command-line tool is faster than clicking through Wireshark for most things. tshark -r capture.pcap -Y "http.request" -T fields -e ip.src -e ip.dst -e http.host pulls exactly the fields you want without opening a GUI. I run that kind of filter first to answer whether a given host even talked to a destination before I bother opening the full file. It usually takes three to five seconds on a multi-gigabyte capture where opening Wireshark would take thirty seconds and then freeze while it parses everything.
The counter-intuitive part that catches people out: TCP retransmissions are not always retransmissions. Sometimes they are just duplicates caused by packet reordering, which is normal on any path with multiple routes or ECMP. If you immediately mark every retransmit as a problem you will spend hours chasing ghosts. Use the tcp.analysis.duplicate_ack and tcp.analysis.out_of_order flags in Wireshark to separate true retransmits from reorder events. True retransmits show up with a sequence number you have already seen. Out-of-order packets show up with a sequence number higher than the last in-sequence packet you received. They look similar in the timeline but they mean completely different things for your diagnosis. Another thing beginners miss is that TLS encrypted traffic is not un-analyzable just because you cannot read the payload. The handshake itself leaks a lot. SNI in the ClientHello tells you the domain being contacted. Certificate Subject Alternative Names tell you who owns the endpoint. JA3 fingerprints can identify the TLS client library. I once spent two days troubleshooting a suspicious outbound connection only to realize the destination was a legitimate SaaS provider because I never looked at the SNI value in the ClientHello and instead got stuck trying to decrypt traffic that I had no key for. The answer was in the clear text before the encryption started. Here is a scenario I ran into that does not appear in any tutorial. I was capturing on a host behind a NAT gateway where the internal IPs were not the ones showing up in my capture. The packets on the wire had already been source-NATted by the router, so ip.addr filters against the internal subnet returned zero results and I wasted about forty minutes convinced the traffic was not happening at all. The workaround was to capture on the internal interface instead, or if that was not possible, to add a static route mirror or span port on the switch before the NAT boundary. Capturing after NAT means you can only see the translated IP, which is useless if you are trying to correlate back to a specific internal host. I switched to port mirroring on the upstream switch and saw the pre-NAT addresses immediately.
Get the Full Details
There are real limitations to this approach that people gloss over. Packet capture on a busy link will always miss bursts that exceed your write speed, regardless of how much RAM you have. Ring buffers help, but they drop packets under sustained load. If you need guaranteed completeness, you need a dedicated tap or hardware-based capture device, not just a faster disk. Also, capturing on a virtual machine adds another layer of uncertainty. Hypervisor scheduling can delay packet delivery to the VM by milliseconds, which distorts timing analysis and makes round-trip time calculations unreliable. I stopped using VMs for timing-critical captures years ago and just plug a laptop into a span port or use a hardware tap instead. If your goal is understanding what an application is doing and you have access to the host, application-level logging will almost always give you more useful information than packet capture. A capture tells you that a TCP handshake happened. It does not tell you why the application decided to send that request. I use packet capture when I need to see network-level behavior that the application does not log, like DNS resolution failures, certificate errors visible in the TLS handshake, or unexpected connections to infrastructure I did not expect. For everything else, logs are faster and less noisy. The common tools are tcpdump, Wireshark, tshark, and for larger scale work,Zeek or Suricata for logging. Zeek turns pcap into structured logs that you can query without reloading the original capture. It costs CPU and memory but it pays for itself if you are doing this regularly. For a one-off investigation, Wireshark and tshark are enough.
I keep a few saved capture filters in a text file on every machine I work on. They are not sexy but they save time when you are already stressed. Something like tcp port 443 or tcp port 80 or udp port 53 or tcp port 853 for a basic web and DNS capture. not arp and not rarp if you just want to ignore noise. host 192.168.1.50 and not port 22 if you are isolating a specific machine but do not care about your own SSH session. Writing these down once means you are not typing them from memory under pressure. The actual workflow I use now is very mechanical. Set the capture filter. Start the capture. Let it run for the window you need. Stop it. Check the drop count. Run a tshark summary to see protocol distribution. Open Wireshark only if the summary points to something worth digging into. This keeps the GUI open for actual analysis instead of wasting it parsing a capture you already know is clean. That is it. It is not exciting. It works if you respect the limits of what a packet capture can show you and what it cannot.