Sketching in AI is harder than people admit
Most generators will give you something that looks like a kid's crayon drawing when you ask for a sketch. That's because "sketch" is one of those words that gets overused in prompt engineering. It means nothing specific to the model. What you actually want depends on whether you're looking for architectural line work, fashion croquis, rough charcoal studies, or ink wash gestural drawings. They all look different and require completely different prompt structures. Here's what's actually working right now across Midjourney v6.5, Stable Diffusion XL, and Flux. The short version: you need to specify medium, line quality, paper type, and confidence level. Not in that order necessarily, but all four elements matter. A prompt like "quick gesture sketch of a man on a horse, charcoal on toned paper, loose confident lines, visible erasure marks" produces something recognizably human-made. Add "photorealistic" to that and you get a photograph of a charcoal drawing, which is usually not what you want. The distinction matters more than you'd think.
I spent about three weeks last month trying to get consistent architectural elevation sketches from Flux. The default output kept looking like a blueprint crossed with a watercolor, which is useless if you need clean linework for presentation boards. The fix was adding "technical pen on tracing paper, 0.3mm mechanical pencil overlay, no shading, pure contour lines" and dropping the realism parameters to near zero. It took about eight tries before I stopped second-guessing myself on the results.
What actually moves the needle
Medium specificity is the single biggest factor. "Pencil sketch" is vague. "Hoarder's graphite on bristol board, 6B soft lead, heavy pressure variation" tells the model exactly what texture and contrast range to aim for. Same with ink. "India ink on cold-press watercolor paper" gives you completely different line behavior than "dip pen on smooth scratchboard." The absorption characteristics of the paper change how the lines render, and modern models do pick up on this if you're specific enough. Line confidence is the second factor people ignore. A real sketch has decisive lines mixed with hesitant ones. Prompts that include "confident single-weight lines with occasional build-up hatching" produce cleaner results than "messy sketchy lines," which just gives you visual noise. The word "sketchy" is basically a trap in prompt language. Models interpret it as adding random textures and artifacts rather than actual sketch aesthetics.
Get the Full Details

Common mistakes that waste time
The biggest pitfall is mixing too many mediums in one prompt. "Charcoal and ink and colored pencil sketch" doesn't give you a rich layered drawing. It gives you something that looks like the model couldn't decide what it was doing. Pick one primary medium and one secondary if you need layering. Keep it to two at most. Another one: asking for "rough" or "unfinished" without specifying what kind of rough. A fashion illustration rough has a completely different character than an anatomical study rough or a composition thumbnail. The word "rough" changes meaning entirely based on context, and models don't share your context unless you tell them. Resolution and aspect ratio also interact with sketch quality in ways that aren't obvious. Tall, narrow prompts (like 2:3 or 3:4) tend to produce more elongated, flowing lines that read as sketches. Square formats push the model toward more deliberate, composed drawing behavior. If you want that quick loose feeling, try a taller crop and tell the model the subject is captured in a single continuous gesture.
When this approach breaks down
These methods work well for figurative and architectural work. They break down for abstract or highly stylized sketching. If you're trying to do something like Hokusai-style rapid ink studies or Willem de Kooning–type gestural marks, the models don't have enough training data to reliably reproduce that aesthetic. You'll get something that vaguely resembles the style but lacks the actual decision-making behind it. In those cases, feeding a reference image through IP-Adapter or ControlNet with a sketch preprocessor gives you significantly better results than text alone. Stable Diffusion users should also know that the default checkpoints are terrible for sketch work out of the box. You need a checkpoint fine-tuned for line art or drawing, or you need to run your generation through a LoRA trained on sketch datasets. The difference between a base SDXL model and one with a sketch-specific LoRA applied is usually a 60 to 80 percent reduction in iteration count to get a usable result.
A practical workflow that actually saves time
Generate your base sketch at a low resolution first. Something like 768x1024. Don't go full resolution because you'll waste GPU time on compositions you're going to discard. Get the line arrangement right, then upscale with a dedicated line-art upscaler if you need print quality. For digital use, 150 DPI is usually sufficient for sketch work and the file sizes stay manageable. If you're working in Midjourney, use the --style raw parameter to reduce the model's tendency to beautify everything. Add --sref with a reference image of an actual sketch you like, and set the stylize value low, around 50 to 100. Higher stylize values push the output toward illustration territory, which defeats the purpose of a sketch prompt. For Stable Diffusion workflows, combine a lineart ControlNet preprocessor with your sketch prompt. This locks down the composition while the prompt handles the medium and line quality. Without ControlNet, you're essentially rolling dice on layout every generation, which turns a 5-minute process into a 45-minute one.

The prompts themselves are fairly reusable once you find a combination that works for your needs. Save them with their parameters, note what went wrong on each attempt, and build a small personal library. After two or three weeks of this, you should be able to generate acceptable sketch outputs in under ten minutes per concept instead of burning an hour on iterations.