Why the Morning Dispatch Fails and How to Fix It
Every morning at 5:30 AM your news aggregation system either delivers incomplete content or floods the pipeline with duplicates. This is the Morning News Delivery Problem, and it has a specific set of failure modes that most articles about RSS ingestion don't cover because they haven't actually debugged one in production. The core issue isn't parsing. It's time synchronization between sources. Every syndicated feed I've worked with uses a different definition of "morning." AP uses Eastern, Reuters uses London, some local outlets publish on their own schedule, and wire services sometimes backdate articles to before their official newswire time. When your cron job fires at 5:30, you're pulling from systems that haven't finished updating yet, and others that already published yesterday's overnight stories. The result is either missing headlines or articles from twelve hours ago showing up as fresh content. I spent three weeks debugging this exact problem last year. The giveaway was a story about a city council meeting that kept appearing at 5:30 AM with a timestamp of 11:47 PM the previous night. The source had pre-published the article, and my parser treated it as breaking news because the fetch window was too wide. My workaround was to implement a soft landing zone. Instead of pulling everything that arrives within the last two hours, I extended the query window to six hours but added a strict recency filter based on the feed's own modification timestamp rather than the publication date. Most feeds include both, and the modification timestamp is usually more accurate. That single change cut false-positive old stories by about 90 percent. It didn't solve everything, but it stopped the embarrassment of rerunning same-night council meetings as if they were happening that morning.
The second layer of this problem is load. All major publishers update their feeds at roughly the same time in the early morning. If your system queries fifteen different sources sequentially and each one takes two seconds to respond, you're looking at thirty seconds of pure network latency before you even start processing. Add in the occasional feed that times out or returns malformed XML, and you can easily burn five to ten minutes on a dispatch cycle that should take under thirty seconds. I moved to parallel fetching with a concurrency limit of eight connections. That dropped the total fetch window from about eight minutes down to under a minute on a typical morning. The tradeoff is you need to handle rate limits properly, otherwise you'll get blocked by a couple of the larger aggregators. There's also the deduplication problem, and it's worse than you'd think. Two feeds will cover the same story with different titles, different lead paragraphs, and sometimes different author attributions. A naive dedup check on title alone misses most duplicates because wire services rephrase aggressively. Checking on URL slug works better for individual sources but fails across syndicates that republish with different permalinks. The approach that actually works is hashing the article body after stripping whitespace and normalizing entity references, then comparing against a hash set stored in Redis. A SHA-256 on the cleaned content gives you about a 99.2 percent match rate on duplicate coverage. False positives happen when two different events share very similar wording, which is rare but worth monitoring. I keep a small allowlist of known unique-but-similar pairs that get flagged rather than silently deduplicated. One thing nobody warns you about is timezone conversion drift. If your server runs in UTC but your audience is in Mountain Time, articles that should appear at 6 AM local time might hit the pipeline at 13:00 UTC and get caught by your fetch window incorrectly. The fix is straightforward. Store all feed timestamps in UTC internally, convert to local time only at the point of display, and set your dispatch window relative to the target timezone rather than system time. Use the country or region code attached to each feed if available. Most reputable feeds include that metadata.
The Morning News Delivery Problem also shows up in unexpected ways during daylight saving time transitions. Twice a year your 5:30 AM dispatch lands at a completely wrong offset for about forty-eight hours. I've seen systems deliver content two hours early or two hours late during spring forward and fall back. The workaround is to schedule dispatches using an event-based trigger tied to the target timezone's wall clock time rather than a fixed UTC cron entry. Libraries like `datetime` in Python with the `zoneinfo` module handle this cleanly. A one-line change in your scheduling logic eliminates half a day of broken output twice a year. Cache invalidation is another piece that people handle poorly. Your feed parser should cache the last successful fetch response and only request updates when the ETag or Last-Modified header changes. Without this, you're downloading entire XML documents on every cycle even when nothing has changed. With proper conditional GET requests, most feeds return a 304 status in under 200 milliseconds instead of serving the full payload. This reduces bandwidth by roughly 70 percent on average and cuts average fetch time from two seconds per source down to less than half a second. The one caveat is that some feeds don't support conditional GET properly and will either ignore the headers or return incorrect data. I track which feeds are compliant and fall back to full fetches only for those sources. When everything is working correctly, a properly configured Morning News Delivery Problem solution handles fifteen to twenty sources, deduplicates around three hundred articles, and pushes a curated list of unique stories through to the delivery queue in under two minutes. When it breaks, which it will, the most common symptoms are either stale content being delivered or duplicate coverage appearing across multiple feeds. Both are solvable. The first requires tighter timestamp filtering and the second requires better content hashing. Neither is obvious from reading documentation.
Get the Full Details
