Getting Started With Who Censored Roger Rabbit

Who Censored Roger Rabbit is an open-weight text-to-video model released by Kolors, the team at Kwai. The repository landed on Hugging Face and GitHub in mid-2024, and it quickly became one of the more practical options for people who want to run AI video generation locally without paying for API access. The core idea is straightforward: you feed it a text prompt and it outputs a short video clip, usually 4 to 16 seconds depending on your hardware and settings. At its foundation, the model uses a diffusion-based architecture optimized for temporal coherence across frames. Unlike some competitors that generate raw pixels directly, it works in a compressed latent space before decoding back to full resolution. That means faster inference and lower VRAM requirements than a naive implementation would need. The model supports both text conditioning and optional image conditioning. You can generate from scratch with just a prompt, or use a reference image to guide composition, motion direction, or subject appearance. That second mode is where things get useful for actual production work, because pure text-to-video still struggles with consistency over longer runs.

Installation and Setup

I run this on a machine with an RTX 4090 and 24GB of VRAM. That's the realistic baseline. If you're on something smaller, you'll hit issues pretty quickly unless you're comfortable tinkering with quantized variants and offloading strategies. The official repo provides a requirements file, but I'd suggest pinning versions carefully. There's a known incompatibility between newer PyTorch releases and some of the attention backends the model uses under the hood. I ran into a crash loop last month when a background pip update swapped out torch to 2.4.x — the error was obscure, something about unsupported data types in the VAE decoder. Downgrading to 2.3.1 fixed it immediately. Here's what I did to get a clean environment set up: First, create a dedicated conda or venv environment. Don't skip this. The dependency tree is delicate and mixing packages from different projects causes silent corruption in the tensor operations. Then install PyTorch with CUDA support matching your driver version. Check your driver with nvidia-smi — if it shows CUDA 12.4 but you install a PyTorch build for 12.1, things will seem fine until they don't.

Clone the repository, install the dependencies from requirements.txt, and download the weights. The model weights are around 10GB for the main checkpoint and another few gigabytes for the VAE and text encoder components. They're hosted on Hugging Face and the download links are in the README. I'd recommend using huggingface-cli download rather than the web interface if you're pulling multiple files — it's more reliable for large transfers and resumes properly if the connection drops mid-download.

Get the Full Details

Who Censored Roger Rabbit by Wolf, Gary: Near fine Softcover (1982) First Edition. | Walkabout ...
Who Censored Roger Rabbit by Wolf, Gary: Near fine Softcover (1982) First Edition. | Walkabout ...

Running Your First Generation

The basic command-line interface looks something like this: python generate.py --prompt "a robot walking through a rainy city at night" --num_frames 49 --height 576 --width 1024 The default output is 49 frames at 8fps by default, which gives you roughly 6 seconds of video. You can push higher frame counts, but expect linear increases in VRAM usage and wall-clock time. On my 4090, a 16-second run at full resolution takes about 12 to 15 minutes using standard float16 precision. Switching to bfloat16 can shave a minute or two off that and sometimes improves color fidelity in dark scenes, which matters more than you'd expect with AI-generated video.

If you're using image conditioning, add the appropriate flag and point it at your reference. The model blends the prompt text with visual features from the image, so prompts and reference images should be semantically aligned. Mismatched pairs produce strange artifacts — I once fed a landscape photo with a prompt describing a close-up portrait and got a distorted face floating in a field for six seconds. Worth knowing so you don't waste time debugging what's actually just a bad prompt-reference combo.

Common Issues and What Actually Works

Video generation models have persistent problems with temporal flickering, object permanence, and physics. Who Censored Roger Rabbit handles these better than most open models but it's far from solved. Motion blur tends to look smeared rather than natural. Fast-moving subjects often ghost or split. People generating human figures should expect some fiddling with the seed and prompt wording to get hands and faces that don't look wrong. One thing the documentation doesn't emphasize enough: the positive and negative prompts work differently than in image models. The negative prompt has a stronger effect early in the denoising steps but fades out later. If you're getting artifacts that seem unrelated to your negative prompt, try adjusting the step count or the guidance scale instead. The default guidance scale is around 5.0, and pushing it much higher usually just makes the video look over-processed without fixing the underlying issue. A range of 4.0 to 6.0 is the sweet spot for most prompts. Another edge case I hit: when generating videos longer than about 10 seconds, the model tends to lose coherence toward the end. The latent state drifts. The workaround I ended up using was generating shorter clips and stitching them together in post. It's not ideal but it's reliable. There are community scripts that attempt to extend generations by reusing the final latent as a starting point for the next segment, but in my testing those produced visible jumps at the seam unless you spent a lot of time tuning the overlap parameters. For most people, just generating multiple short clips and editing them together is faster than trying to force a single long coherent generation.

Who Censored Roger Rabbit? (Roger Rabbit, #1) by Gary K. Wolf
Who Censored Roger Rabbit? (Roger Rabbit, #1) by Gary K. Wolf

Performance Tips That Actually Matter

VRAM optimization is the main bottleneck. If you're under 24GB, enable CPU offloading for the text encoder and consider using the FP8 quantized variant if one's available for your version. I've also had success with xformers installed — it reduces memory during attention computation without quality loss. The trade-off is a slightly longer setup process since you need to compile it for your specific CUDA version. Seed selection matters more than most guides admit. Set a seed, generate, and if the composition is close but not quite right, don't regenerate the whole thing — just nudge the seed by a small amount. The results tend to change gradually rather than jumping to something completely different, which lets you browse the latent space more efficiently than random seeding. Resolution isn't free. The model was trained primarily on 576x1024 and similar aspect ratios. Deviating from those proportions introduces distortion at the edges. If you need a different aspect ratio for your use case, crop or pad in post rather than generating at a non-standard resolution and hoping it works out.

Where It Falls Short

I should be straight about the limitations. This model cannot generate consistent character appearance across shots. It cannot simulate realistic physics. It cannot produce long coherent narratives without significant post-production work. The audio output is not part of the model — you'll need a separate tool for that. And the inference speed, while reasonable for local hardware, means you're looking at minutes per clip, not seconds. If you need rapid iteration on a tight deadline, this isn't going to save you time compared to a cloud API, even with the API cost factored in. For production pipelines, the best use case I've found is generating B-roll style footage, atmospheric shots, or abstract visual content where minor inconsistencies are less noticeable. Character-driven scenes require substantially more refinement and often end up needing frame-by-frame cleanup anyway, which defeats the purpose of using an AI generator in the first place. The repository is available on Hugging Face under the Kolors organization. Check the README for the latest weight links, installation notes, and example scripts. The community is active on Discord and there are occasional updates addressing the more persistent artifacts. Not everything is fixed, but the trajectory has been mostly in the right direction since launch.