How Colorizing Historical Footage Actually Works
Most people think Speeches In Color is a single tool you run and get a finished video. It isn't. It's a post-production pipeline that chains several separate models together, and understanding that distinction is what separates someone who spends four hours on a thirty-second clip from someone who does it in forty minutes. The core concept is straightforward: you take a black and white source, extract the luminance channel, run it through a generative adversical network trained on color photography datasets, and then stabilize the output across frames so the colors don't flicker around. The trick is in the stabilization. A naive implementation will produce something that looks reasonable frame by frame but jumps between blue and green shirts every other second, which is worse than leaving it grayscale.
Speeches In Color: What You're Actually Getting
The project gained attention because it was applied to archival footage of public figures, and the results were visually striking. But the visual quality is only one part of the equation. The model used is typically based on DeOldify or a similar architecture, fine-tuned on historical photography. That fine-tuning matters because models trained on modern color photographs tend to oversaturate skin tones and push everything toward a warm amber cast that doesn't match mid-twentieth-century film characteristics. I ran into a specific issue last year when processing a 1930s newsreel. The background walls kept shifting between beige, pale gray, and off-white across consecutive frames. The individual frames looked fine in isolation, but the temporal instability was distracting. My workaround was to add a simple motion-compensated median filter across the color channels before feeding the frames into the colorizer. This isn't something most tutorials mention because it's not part of the model itself — it's a preprocessing step that stabilizes the input by reducing frame-to-frame noise in the luminance channel. The fix cut my post-processing time down significantly because I wasn't spending an hour on color grading just to mask the flicker. You also need to think about your source material before you start. Resolution matters more than most people expect. If you run a 480p original through a model that outputs 1080p, you're asking the AI to invent texture that doesn't exist. The result is soft, smeared color that looks artificial up close. I usually upscale the source first using a dedicated super-resolution model like Real-ESRGAN, then run the colorization pass. Doing it in that order gives you cleaner edges and more coherent color boundaries, especially around clothing fabric and skin.
Another thing beginners miss is the audio handling. The colored video looks polished until you realize the lip sync feels off because the colorization pass introduces a slight delay. This happens because many pipelines process frames in batches and the batching introduces a few milliseconds of latency per frame. The fix is trivial — just align the audio track to the original timing after colorization, but you have to check it because automated alignment tools often make it worse rather than better. I learned that the hard way on a project where the output had audio that drifted by about two seconds over a three-minute clip. There are also real limitations you should know about. The model cannot accurately reconstruct colors that have no reference data. If a speech takes place indoors with incandescent lighting, the model will guess at wall colors and furniture based on its training set, and those guesses will often be wrong. It has no way of knowing whether a suit is navy blue or charcoal gray when the source footage has no chroma information. For black and white newsreels from the 1940s and earlier, the uncertainty is highest because the training data skews toward outdoor daylight photography. If you need historical accuracy rather than aesthetic appeal, you should cross-reference with existing colorized photographs from the same era and use those as reference frames during processing. Some implementations of the pipeline allow you to inject reference images, which dramatically improves accuracy for consistent subjects like faces and uniforms. Without that step, you're getting artistic interpretation, not documentation.
Get the Full Details

The best freely available implementation I've worked with is the open-source DeOldify repository on GitHub, paired with a preprocessing script for temporal stabilization. There's also a Colab notebook version that lets you run it without setting up CUDA locally, though it's slower and occasionally crashes on clips longer than ninety seconds due to GPU memory constraints. For anything beyond short clips, I'd recommend running it locally on a machine with at least 8GB of VRAM. The raw footage isn't something you can download as a completed product from a single source — the project team released their methodology and model weights, but the final colorized videos are hosted on YouTube and the project's GitHub page. You can find the pipeline code and instructions at the official repository if you want to process your own material rather than using pre-rendered output.
Processing Workflow
Start by extracting the video to individual frames. Don't skip this step. Trying to run colorization directly on video files through most wrappers adds unnecessary overhead and makes debugging harder. I use ffmpeg to pull frames at the original frame rate, which for most historical footage is 24 or 25 fps. Next, run the super-resolution pass on all frames before colorization. This is non-negotiable if you care about edge coherence. The upscaler I use is Real-ESRGAN with the general purpose model, processed in batches of thirty-two frames at a time. Anything larger and you start seeing VRAM warnings that slow everything down. After upscaling, apply the temporal median filter I mentioned. A three-frame window is usually sufficient — too wide and you blur motion, too narrow and it doesn't stabilize anything. This step takes about two minutes per minute of footage on a decent CPU, so don't underestimate it.
Then run the colorization model. Batch size here depends on your GPU. Sixteen frames per batch is a safe starting point for an RTX 3060. Eight frames if you're on something smaller. The model processes each frame independently, which is why the stabilization step beforehand matters so much. After colorization, recombine the frames back into a video. Use the same frame rate as your source. Then align the audio and do a final check for any remaining flicker. Most of the time you won't see it on a television or monitor, but on a large screen it becomes obvious. If the result looks too saturated, reduce the color strength parameter in the model. The default setting is tuned for dramatic effect, not documentary accuracy. Dropping it to around 0.6 usually produces a more natural look for indoor historical footage.

I've found that the whole process for a five-minute clip takes roughly forty-five to sixty minutes on a machine with an RTX 3080, including preprocessing, colorization, and recombination. A weaker machine might take two to three hours for the same work. Planning around that timeline matters if you're processing multiple clips.