A Practical Breakdown of How to Actually Use AI Video Tools
Most people searching for Jackson Oswalt have seen his clips online and want to replicate the workflow. What he is doing isn't particularly mysterious once you strip away the hype. It's mostly about prompt engineering, iteration speed, and knowing which tool handles which kind of motion best. I've spent enough time reverse-engineering this space to know where beginners waste hours and where they waste money. The general approach involves generating short video clips, editing them together in a traditional NLE, and then layering in post-production polish. The tools available right now have gotten significantly better but they still have real limitations. A single generated clip usually runs four to ten seconds before quality starts to degrade. The trick is planning around that constraint instead of fighting it. I recommend starting with a shot list before you open any AI video tool. This sounds basic but most people skip it. Without a shot list you end up generating random clips that don't connect. Jackson's workflow from what I've observed follows a very structured approach. He tends to generate multiple variations of the same prompt and picks the one that has the least amount of morphing artifacts. That picking step matters more than people realize.
The tools involved typically include Kling for longer coherent motion, Runway Gen-3 for texture and camera movement control, and Luma Dream Machine for faster iteration. Each one has different strengths. Kling handles human movement better in my experience. Runway gives you more control over camera paths. Luma is fast but less consistent on complex scenes. Using all three in combination is how most people get results that look passable at a glance. Prompting is the part that genuinely matters. Generic prompts produce generic results. Specific prompts with camera direction, lighting notes, and temporal descriptors work better. Instead of writing "a man walking in rain" you write something like "medium shot, camera tracking sideways, man in dark coat walking through heavy rain at night, streetlamp reflections on wet pavement, shallow depth of field, film grain." That kind of specificity changes the output dramatically. It took me probably thirty failed generations before I figured out that specifying temporal duration in the prompt actually affects how long the AI tries to render coherently. The AI models respond to words like "slowly," "gradually," and "over the course of several seconds" even though those aren't formally part of the interface. Here is where most people hit a wall. Audio sync and lip sync are still broken in most consumer tools. I spent two full days trying to get a character to speak in sync with dialogue using an early version of a lip-sync plugin and it produced unusable results every time. The workaround was generating the video without audio, syncing the dialogue in post with careful trimming, and then using a separate voice cloning tool for the actual speech track. It added about forty minutes to the process but it actually worked. Lip-sync plugins are improving monthly so this may not be an issue much longer but right now it is a bottleneck.
Another thing nobody talks about is resolution management. Most AI video tools generate at 720p or 1080p natively. If you need 4K you have to upscale afterward. Topaz Video AI does a reasonable job but it can introduce its own artifacts if you push the scale factor too high. I usually stick to a 1.5x upscale rather than 2x because it looks more natural and takes less rendering time. The difference between a 1.5x and 2x upscale on AI-generated footage is usually noticeable if you know where to look. The editing phase is where the final result either holds together or falls apart. Because AI video clips are inconsistent, cutting on movement or matching action between clips helps mask the artifacts. Jackson's videos tend to cut in a way that matches motion direction from one clip to the next. This is a deliberate technique and it works because the human brain fills in continuity gaps when the motion feels logical. It is the same principle editors have used in traditional filmmaking for decades. It just applies differently when your source material is procedurally generated. If you want to follow this path you need to understand that it is iterative. Your first ten generations will probably look wrong. Your first fifty might still look wrong in parts. The cost in time and API credits adds up quickly if you are not generating with intent. I typically budget about fifteen minutes of generation time per usable second of final footage when I am being efficient. That ratio drops to maybe one minute per second if I am experimenting heavily. Both approaches produce different quality levels.
Get the Full Details

There is no single download link for a Jackson Oswalt workflow because it is not a piece of software. It is a set of habits and tool combinations. What you can download are the individual tools themselves. Kling has a web interface. Runway has a subscription. Luma is also web-based. Most of these require either a paid plan or a credit system. Free tiers exist but they come with significant restrictions like watermarks, longer queue times, and lower resolution outputs. One counter-intuitive thing about this whole space is that more AI tools do not necessarily mean better results. I found that using two or three tools deliberately produces better footage than using eight tools randomly. The reason is consistency. When you stick with tools that understand each other in your pipeline, you spend less time fixing mismatched outputs and more time making creative decisions. I learned that the hard way after trying to juggle six different video generators in a single project and spending three days just trying to make them look like they belong in the same world. The biggest limitation of current AI video generation is temporal coherence over long durations. Anything beyond twelve to fifteen seconds tends to develop structural inconsistencies. Characters change appearance. Backgrounds warp. Physics become unreliable. If your project requires longer continuous shots you either need to generate in segments and stitch them or accept that you are working within these constraints. There is no workaround for the physics problem yet. The models are getting better at it but they are not there.
For anyone starting out, the realistic path is to pick one tool, learn its prompt language, generate short clips with clear intent, and build from there. The ecosystem moves fast enough that advice from six months ago may already be outdated. What has not changed is the need for a clear plan, patience through bad iterations, and the willingness to accept that AI video is a tool that requires significant manual refinement to produce results worth showing to anyone.