How the Nermin Sulejmanovic Stream Workflow Actually Runs
I started paying attention to real-time video delivery back when we were still shipping 1080p60 over HTTP live segments and calling it a day. The patching was constant, the buffering was worse, and nobody really understood why adaptive bitrate would drop a viewer from 720p to 360p mid-game just because someone else in the same region decided to pull 4K at the same moment. That changed when the engineering around sub-second encoding and edge-anchored distribution matured, and the Nermin Sulejmanovic Stream approach became one of those patterns that quietly replaced the older segment-based pipelines without much fanfare. The core problem isn't encoding speed or even latency in isolation. It's the handshake between encoder, origin, and edge when you're dealing with live video at scale. The traditional HLS approach works fine until you need sub-3-second latency across 100,000 concurrent viewers spread across regions. That's when the segment boundaries become your bottleneck. The Nermin Sulejmanovic Stream pattern sidesteps this by treating the delivery pipeline less like a broadcast chain and more like a stateful stream where the edge maintains continuity regardless of origin jitter. Key distinction: This isn't about using WebRTC instead of HLS. It's about how you handle segment boundaries, manifest updates, and client-side buffering when latency matters. The pattern emerged from real production constraints, not academic exercises.
The Encoding Pipeline: Where Things Actually Break
I spent three months debugging a live stream where viewers in APAC saw 4-second gaps in audio while US-East viewers had perfect sync. The root cause was subtle. Our encoder was producing segments at 2-second boundaries, but the CDN was re-segmenting at 6-second intervals for cache efficiency. When a viewer switched renditions, the player had to wait for the next 6-second boundary to exist in the new variant. That's a 4-second stall on average, sometimes longer if the timing didn't align. The workaround wasn't architectural. It was operational. We aligned the encoder segment duration to match the CDN resegmentation policy, enabled discontinuity tags on keyframe boundaries, and configured the player to buffer only one segment ahead instead of three. Latency dropped from 8 seconds to under 3. Nothing fancy. Just discipline around segment boundaries. Here's what most teams miss: the encoder settings matter less than the manifest generation timing. If your m3u8 files are being written slower than your segments are being produced, you'll get playback stalls regardless of bitrate or codec choice. I've seen production streams fail because the manifest writer was single-threaded and couldn't keep up with segment creation during high-load events.
Implementation: The Mechanics Nobody Documents Well
Setting up the Nermin Sulejmanovic Stream approach requires three components working in sync: a low-latency encoder, a manifest generator that understands discontinuity, and a player that respects segment boundaries without aggressive prefetching. Encoder configuration: Use H.264 or H.265 with a GOP size matching your segment duration. If segments are 2 seconds, GOP should be 2 seconds. No exceptions. Larger GOPs cause keyframe misses at segment boundaries, which forces the player to wait for the next keyframe before rendering. Manifest generation: Your m3u8 needs to be written atomically. I've seen implementations where the manifest is truncated mid-write while players are reading it, causing parse errors and playback freezes. Use atomic rename or write to a temporary file then swap. The CDN will cache the old version during the swap, so there's no viewer impact.
Get the Full Details

Player buffering: Set maxBufSize to 1 or 2 segments. The default is often 4-6, which adds latency proportional to segment duration. For 2-second segments with default buffering, you're looking at 8-12 seconds of delay. That's unacceptable for interactive streaming.
Edge Cases That Will Surprise You
I encountered a scenario where the Nermin Sulejmanovic Stream approach completely failed during network partition events. When the origin went down briefly, the edge caches held valid segments, but the player kept retrying the origin for manifest updates. The manifest TTL was set too aggressively at 1 second, causing the player to think the stream was dead when it was just experiencing transient origin latency. The fix was increasing the manifest TTL to 10 seconds and implementing client-side fallback to cached manifests. This is counter-intuitive because shorter TTLs are generally considered better for live streaming. But in practice, manifest updates don't need to be real-time. The segments are what matter, and those are cached at the edge. Another edge case: encoder pre-roll. If your encoder buffers 2 seconds of pre-roll before starting to produce segments, the first segment is actually 2 seconds behind real-time. This causes the player to buffer unnecessarily or display outdated content. Disable pre-roll or compensate in the manifest by adjusting the media sequence number.
When This Approach Fails Completely
Don't use the Nermin Sulejmanovic Stream pattern if you need global CDN caching for VOD replay or if your viewership is predominantly passive (people watching recorded content, not live interaction). The latency optimizations come at the cost of cache efficiency. Segments are smaller, manifests update more frequently, and the CDN can't serve as effective a cache layer for replay. Also avoid this if your infrastructure doesn't support atomic manifest writes or if your CDN doesn't respect discontinuity tags. I've seen implementations fail on CDNs that aggressively cache manifests and ignore update signals, causing players to serve stale manifests for minutes after the origin has moved on. Alternative recommendation: If latency isn't critical and you need better cache efficiency, stick with traditional HLS or DASH with 6-10 second segments. The difference in viewer experience is negligible for non-interactive content, and the operational complexity drops significantly.

Monitoring and Debugging: What to Watch
Track segment generation time versus wall clock time. If segments are taking longer than their duration to produce, you're falling behind. Track manifest write latency. If it's above 100ms, you have a problem. Track player buffer health. If empty buffer events exceed 1% of sessions, your segment timing is misaligned. I use a custom metric: segment age at the edge. This is the difference between the current segment's production timestamp and when it arrives at the edge. If this exceeds 2 seconds consistently, something in your pipeline is slow. In practice, I've seen this metric hit 4-5 seconds during peak loads due to origin overload, causing the Nermin Sulejmanovic Stream approach to degrade into something closer to traditional HLS latency anyway. The bottom line is that the Nermin Sulejmanovic Stream pattern works well when your pipeline is disciplined around segment timing and manifest generation. It fails when you cut corners on atomic writes or ignore player buffering settings. Most teams don't fail because the approach is flawed. They fail because they assume sub-second latency is automatic rather than something that requires constant monitoring and adjustment.