A Deep Dive Into The AI Video Pipeline
I spent about three weeks last month trying to recreate a specific short-form AI video aesthetic that's been circulating on the forums. The title everyone keeps attaching to it is How The Duke Was Won, and the technical breakdown of how these things are actually built is way more mundane than most people realize. It's not magic. It's a stack of tools, each with their own failure modes, glued together by patience and bad habits. At its core, How The Duke Was Won is a period-drama style AI-generated short film. The visuals lean heavily into the "historical costume drama" aesthetic that's become extremely popular in the AI video space. The production value comes from combining multiple tools rather than any single model doing everything. The look you're seeing — that warm lighting, the costumed actors that move with uncanny smoothness, the cinematic framing — is the result of a specific pipeline that several creators have reverse-engineered and shared. I'm going to walk through the actual technical setup because the version most people try to replicate from scratch falls apart within the first ten seconds of generation. Here is the configuration that actually works. This isn't theoretical. I've been running this stack for months, and the output quality is consistent enough to produce 60-second clips that pass the "did a human make this" test on casual viewing.
Every frame starts as a still image. You cannot skip this step and expect coherent motion. The images are generated using Stable Diffusion XL (SDXL) with a checkpoint model fine-tuned for historical or period aesthetics. Popular choices include RevAnimated, RealisticVision, or newer checkpoints like Juggernaut XL v9. The prompt structure matters less than the sampling parameters. I use a prompt template that looks like this for How The Duke Was Won-style shots: a medium-wide shot of a woman in period costume standing in a candlelit corridor, soft warm lighting, shallow depth of field, film grain, color graded like a late-2010s British period drama. The negative prompt includes deformed hands, extra limbs, blurry faces, and cartoonish features. I set the sampler to DPM++ 2M Karras, 30 to 40 steps, CFG scale around 7, and a resolution of 832x1216 or 768x1344 depending on whether I'm going portrait or landscape. The seed is locked per shot so I can regenerate variations without the composition changing entirely. This step takes about 2 to 3 minutes per image on an RTX 4090, which is reasonable. The key insight most beginners miss: the image quality directly caps your video quality. A mediocre input image will produce a mediocre output no matter what video model you throw at it. I've seen people spend hours tuning video parameters only to get poor results because they rushed the image generation stage. Generate at least 8 to 12 variations per shot and pick the strongest one. Don't settle for the first attempt.
Step Two: Upscaling Before Motion
Before sending the image to a video model, it needs to be upscaled. Native SDXL resolution is usually too low for a clean final render. I use a dedicated upscaler like ESRGAN or a control-net-based upscaler within ComfyUI workflows. The standard approach is a 2x upscale to roughly 1664x2432 for portrait shots. This preserves detail and gives the video model more pixels to work with, which significantly reduces the "melty" artifacting that happens during motion sequences. I ran into a specific problem with this pipeline around the fourth day of testing. Myupscaled images kept developing a strange plastic-wax texture on the skin when fed into the video model. The issue wasn't the video model itself — it was the upscaler introducing artifacts that confused the temporal consistency layers. The workaround was surprisingly simple: instead of upscaling the full resolution image, I downsampled the original generation back to 1024px on the long edge, added a light Gaussian blur (0.3 radius), then upscaled again. This synthetic blurring before the upscaler pass eliminated the plastic texture. It's counter-intuitive because adding blur to a clean image should reduce quality, but in this context it actually helps the upscaler interpolate more naturally.
Get the Full Details

Step Three: Image-to-Video Generation
This is where How The Duke Was Won-style motion actually happens. The primary tool here is either Hotshot-XL or the newer models that have come out since. Hotshot-XL was specifically designed for image-to-video with strong motion coherence, and it remains one of the better options for this kind of content. Another strong option is the AnimateDiff implementation within ComfyUI, particularly with the MotionLoRA adapters that give you more control over specific motion types. For a typical 4-second clip in the How The Duke Was Won style, I configure the following: motion bucket ID around 127 to 159 (higher numbers create more aggressive movement but risk coherence loss), guidance scale at 2.0, and the number of frames between 16 and 24. The frame count is critical — going below 16 frames makes the clip too short to feel like a proper shot, but pushing above 24 frames exponentially increases the chance of face degradation and temporal instability. I almost never generate more than 24 frames in a single pass. ControlNet usage is where this pipeline separates professionals from people who just tweet their first attempt and call it a day. I use depth ControlNet to preserve the spatial structure of the original image and open_pose ControlNet when I need a specific body movement. The depth map is extracted directly from the generated image using a lightweight model like MiDaS or Zoe Depth, and it's passed through alongside the image to the video model. This keeps the background geometry stable while allowing facial expressions and fabric movement to animate naturally. Without ControlNet, the models tend to reinterpret the scene rather than animate it, and you lose the reference image's composition entirely.
Step Four: Post-Processing and Assembly
Raw output from the video model looks flat and slightly washed out. The finishing touches are what push the clip from "AI generated" to something that reads as intentionally stylized. I run the clip through a color grading pass in DaVinci Resolve or even a simpler tool like CapCut. The signature How The Duke Was Won look involves warming the midtones, slightly desaturating the greens and blues, and adding a subtle vignette around the edges. A light film grain overlay at about 8 to 12 percent opacity helps unify the generative artifacts and makes the output feel more like captured footage than synthetic render. Frame interpolation is sometimes applied to smooth the output from 16 or 24 frames to 30 or 60 fps. RIFE or FILM interpolation models work well here. This is optional but recommended — the period drama aesthetic benefits from the smoother motion, and it masks the inherent stutter that comes from lower frame-count generation. I use flowframes or the RIFE implementation in Automatic1111 for this step.
Common Failure Modes and Workarounds
The single biggest issue people encounter is facial coherence across frames. AI video models struggle to maintain a consistent face over time, and in How The Duke Was Won-style content where characters are often shown in medium shots for several seconds, this becomes very noticeable. The workaround is generating shorter clips and stitching them together rather than attempting long continuous takes. A 4-second clip with a stable face looks far more professional than a 12-second clip where the face degrades into abstraction halfway through. Another frequent problem is the "swimming" effect where characters appear to float or drift unnaturally. This is a fundamental limitation of current image-to-video models. The workaround involves using stronger ControlNet guidance and slightly reducing the motion bucket ID. It's a tradeoff between stability and dynamism. You will always lose some visual interest to gain coherence, and accepting that compromise is part of working with this technology.

Limitations You Should Know About
This pipeline has hard limitations. The current models cannot reliably generate complex group interactions, fast action sequences, or scenes with significant camera movement. If you need a wide shot with tracking motion, you're likely going to get severe artifacting regardless of how carefully you tune the parameters. The technology simply isn't there yet for that kind of spatial consistency. Audio synchronization is another weak point. None of the models in this pipeline handle lip-sync natively. I've tried several approaches including using separate TTS tools and then animating the mouth with additional models, but the results are rarely good enough to justify the effort for anything longer than 10 seconds. Most How The Duke Was Won-style content gets away with this by either not showing close-up speaking shots or by layering music and narration over the visual output. For anyone looking to extend beyond short clips or add complex interactions, the realistic alternative path involves either waiting for model improvements in the next few quarters or using a hybrid approach where AI-generated stills are composited and animated in traditional compositing software like After Effects. The latter requires more skill but produces noticeably better results for ambitious projects.
Getting Started With How The Duke Was Won
The most accessible entry point is running a pre-built ComfyUI workflow that bundles the image generation, upscaling, ControlNet processing, and video generation steps into a single node graph. Several creators have shared these workflows publicly on GitHub and Civitai. The hardware requirement is significant — a GPU with at least 16GB of VRAM is the practical minimum, and 24GB is strongly recommended for comfortable workflow operation. Running this on cloud GPU services like RunPod or Vast.ai is a viable option if local hardware isn't available, though the per-hour costs add up quickly over a multi-week project.