Digital voice makes you wait. Here's where the delay comes from, why all the linking out there makes it worse, and the one habit that turns a stumbling QSO into a clean one.
You key up, say your callsign, and… a beat passes before anyone answers. Then two people come back at once and step on each other. If digital voice has ever felt a half-step behind, you weren't imagining it. There is a real, built-in delay between your microphone and the other operator's speaker — and unlike a hiss or a dropout, it isn't a fault you can fix. It's the nature of the mode. But once you understand it, you can operate with it instead of tripping over it.
Where the delay comes from
Getting your voice from your radio to a listener across the internet is a chain, and every link in the chain adds a little time:
The vocoder. To crush your voice down to a few thousand bits per second, the codec can't work one instant at a time — it has to gather a small frame of audio before it has enough to compress. That framing means the first bits don't even leave your radio until a chunk of speech has been collected. It happens again in reverse at the far end, decoding back to audio.
The radio keying up. When you press PTT, the transmitter needs a moment to ramp up and send its sync and header so the receiver can lock on before your voice rides in behind it.
The trip across the internet. Your hotspot or repeater packetizes the stream and sends it to a server or reflector, which fans it back out to everyone connected. That's a round trip over the public internet — longer the farther away the server, and never perfectly steady.
The jitter buffer. Because the internet delivers packets unevenly, every node deliberately holds a tiny buffer to smooth them back into gap-free audio. It's a good trade — a little more delay in exchange for speech that doesn't stutter.
Each of these is small on its own. Stacked up, even a clean, direct path adds a noticeable fraction of a second between you and the person you're talking to.
VERIFY: if you want to quote an end-to-end figure (commonly a few hundred milliseconds on a typical hotspot → reflector path), measure your own path and cite it, rather than a number from memory.
Why all the linking makes it worse
Here is the part that actually bites, and it's the reason this matters more every year: the digital-voice world is linked together like never before. Talkgroups are bridged to other talkgroups, reflectors to other reflectors, and whole modes to other modes. Every one of those links adds another full set of the delays above.
A bridge to another network is another internet hop and another jitter buffer.
A transcode — crossing from one mode's vocoder to another — is a complete decode-and-re-encode, so it piles a whole extra vocoder delay onto the path. (This is the same vocoder story behind transcoding and the Rosetta Stone.)
The big everything-linked-to-everything talkgroups can chain several of these back to back.
The subtle killer: on a heavily linked system the delay isn't just longer — it's uneven. Different listeners are reached by different-length paths, so two people hearing “the same” transmission may hear it half a second apart. That's exactly why the most-bridged, multi-mode talkgroups feel the laggiest and why doubling runs rampant on them.
What the delay does to a QSO
First-word clipping. Start talking the instant you hit PTT and your first syllable never makes it — the stream hasn't established and the vocoder is still filling its first frame. “…is is W6ABC” is the sound of someone who didn't wait.
Getting doubled. When you unkey and hear silence, the channel isn't actually clear yet. Your tail and the path delay mean the “silence” you're hearing is stale. Two operators drop into that gap, and they collide.
The broken rhythm. Quick, snappy back-and-forth — the kind that works on a local FM repeater — falls apart when every exchange carries a half-second tax.
How to operate with it
You can't delete the delay. You can operate so it never causes a problem — and it comes down to a few habits:
Pause after you key. Press PTT, count a quiet “one,” then speak. That single habit saves your first word every time. If you take nothing else from this page, take this.
Leave a real gap between overs. A full second or two of silence before you hand it back, and again before you come back. It lets the delay drain, lets the network catch up, and — most important — leaves a door open for someone to break in.
The more linked the system, the bigger the gap. On a bridged, multi-mode talkgroup, exaggerate the pause. More hops mean more delay, which means more room for a collision.
Listen a little longer than feels necessary before you transmit. What you're hearing already happened; give it a beat to be sure the channel is truly clear.
Slow the rhythm down. Digital voice rewards a slightly more deliberate cadence than analog FM. It isn't a race.
Identify early, not just at the end. On a delayed or bridged path, someone may join mid-over; leading with your callsign tells them who's talking without waiting for the tail.
The takeaway
The delay isn't a defect to be fixed — it's the honest price of squeezing your voice through a codec and sending it around the world over the internet. You'll never tune it out entirely. But the moment you know it's there, a one-second pause does the rest: it turns a doubled-up, half-clipped scramble into a calm, courteous contact. And on today's heavily linked systems, that pause isn't just politeness — it's the whole difference between being heard and being stepped on.