Working With The New Yorker Cartoon Of The Day
The New Yorker posts a cartoon every single day on their website. It goes live around 6 AM Eastern Time and stays up for as long as the editors want it there. Sometimes it gets taken down after a few weeks. Sometimes it hangs around for years. There is no official archive feed you can subscribe to that gives you high-resolution files. The website serves whatever resolution the browser demands, which means if you try to grab it with a script, you are usually getting a compressed web image rather than the original artwork. I spent about three months building a scraping pipeline for these because I wanted a local database I could search by theme, artist, or caption text. The problem hit me pretty quickly: the site uses dynamic loading. The cartoon itself lives inside an iframe or a JavaScript-rendered component, which means a simple requests call to the main page URL returns a shell with no actual image data. You have to either reverse-engineer the API calls they make in the background or use a headless browser to wait for the page to render fully before you pull anything. The workaround I ended up using was much simpler than the headless browser route. I found that the cartoon images are typically hosted on a CDN path that follows a predictable pattern based on the publication date. Rather than trying to scrape the rendered page, I started constructing URLs directly from the date stamp in the page metadata and checking whether the image existed. This cut the processing time from roughly 45 seconds per cartoon down to about two seconds, since I stopped waiting for JavaScript execution entirely.
Here is something most people who try this for the first time miss. The caption text is not always embedded in the same place across different cartoons. Older cartoons from before the site redesign have the caption in a completely different HTML structure than newer ones. If your parser only looks for one pattern, you will silently drop captions from a significant chunk of the archive and never realize it until you go back and check. I learned this the hard way after my initial database had maybe forty percent missing caption fields. I had to write a second pass that checked for multiple caption selectors and fell back through progressively broader patterns until it found something. Another thing worth knowing. The New Yorker occasionally pulls cartoons from the daily page after a while, usually when they rotate content but sometimes for legal reasons involving public figures. If you are building a personal archive and your scraper starts returning 404s on previously working URLs, the cartoon has been removed from the active site. There is no notification system for this. The only reliable backup is your own saved copy. There is no official download link provided by The New Yorker for individual cartoons. The site is designed for browsing, not for extraction. If you want a cartoon on your machine, you either save it manually through the browser or you build something yourself. The resolution you get from the browser is serviceable for personal use but nowhere near print quality. The original artwork is typically at least 2400 pixels on the longest side, and the web version usually sits somewhere around 600 to 800 pixels depending on the layout at the time of publication.
I also ran into an issue with the artist attribution data. Some cartoons list the artist name clearly in the page HTML. Others bury it in a JSON-LD structured data block that is easy to miss if you are not looking for it. A handful of older cartoons have no attribution at all visible on the page, and the artist name only exists in the original print publication. If you are doing this for research purposes, you will need to cross-reference with the printed archives or the New Yorker's own artist database, which is a separate system from the daily cartoon page. The whole process is not particularly difficult, but it is tedious. A well-written script can pull a day's cartoon in under ten seconds, including the metadata. The bottleneck is almost always the post-processing step where you clean up the captions and verify the artist names. That part does not scale well because you have to look at each one individually to make sure the automated parser did not attach a caption from a neighboring ad or related article to the wrong cartoon. If you are only interested in viewing the cartoons without building anything, the website works fine as is. There is a daily email subscription you can sign up for that sends you the cartoon each morning. It is lazy but it covers the basics. If you actually need the images or the text for any reason beyond casual reading, you are on your own for the technical work, and you should be prepared to handle cases where the data is incomplete or has disappeared entirely.