Generating Clean Keyboard Imagery Without Losing Your Mind
Most people trying to get good images of mechanical keyboards from AI generators end up with either plastic-looking trash or completely impossible geometry. The problem isn't the prompt format - it's that everyone copies the same base prompts and then wonders why their results look identical. I spent about three months testing different approaches before landing on a system that actually produces usable renders. Let me walk you through how I set it up. The core issue with keyboard prompts is that AI struggles with small repetitive details like keycaps and switches. When you just type "mechanical keyboard aesthetic" into Midjourney or Stable Diffusion, you get either a blurry mess or a keyboard with 61 keys that somehow have overlapping letters. What actually works is building prompts around specific visual anchors rather than hoping the AI figures it out. Here's what I use as a starting point. "Product photography of a 75% mechanical keyboard on dark walnut surface, soft directional lighting from upper left, individual PBT keycaps visible with slight texture, brass weight on corner, shallow depth of field, shot on 85mm lens, moody ambient tones" - this gets you 80% of the way there. From there you layer in whatever style modifiers you want.
The breakdown matters more than the keywords. Lighting description should always come first in your prompt because it controls how the rest of the image reads. If you skip it, the AI defaults to flat diffuse lighting that makes everything look like a catalog photo from 2012. I usually specify either "soft window light with cool fill" or "single warm point source" depending on the mood I'm going for. The difference between those two choices alone shifts the entire aesthetic by about 40%. Camera angle and focal length are the next thing people get wrong. Every tutorial I've seen says "use a top-down shot" but that makes keyboards look flat and uninteresting. I shoot everything at a 30 to 45 degree angle, 50 to 85mm equivalent. This gives you enough perspective to see the keycap profiles and switch stem heights without the distortion that comes from wider lenses. If you need to show the full layout, stick to 50mm. If you're going for detail shots of just the upper right quadrant, 85mm lets you isolate specific keycaps and get that nice compression effect. I hit a real wall when I tried generating close-up shots of specific switches. The AI would consistently hallucinate the stems - either melting them together or making them the wrong shape entirely. My workaround was to generate the full board first, then use inpainting to fix just the switch area. I'd mask the switches, describe them specifically as "cherry MX compatible switch with rectangular stem and square housing, clear stem visible", and regenerate only that region. Takes about 90 seconds per correction. Not elegant but it works consistently.
The material descriptions are where most prompts fall apart. You need to be explicit about what things are made of. "Plastic keycaps" gets you glossy shiny garbage. "Textured PBT keycaps with matte finish" gives you something closer to reality. Same with the case - "brushed aluminum" versus "CNC milled aluminum with bead blast finish" produces very different results. I learned this the hard way when I spent two days trying to get a specific brushed aluminum look and realized I was just writing "metal case" in every prompt. Here are some prompt variations I've found useful depending on the aesthetic direction: For a clean minimal desk setup vibe: "Minimalist mechanical keyboard aesthetic, white and grey color palette, foam padding underneath, studio lighting, product shot, high detail, clean background"
Get the Full Details

For a darker moody build: "Dark themed mechanical keyboard, custom dye-sublimated keycaps, ambient RGB underglow, concrete surface, low key lighting, cinematic composition" For vintage or retro builds: "Retro mechanical keyboard aesthetic, beige and brown tones, vintage computer accessories nearby, warm incandescent lighting, film grain texture, 1990s office desk scene" For technical documentation style: "Mechanical keyboard blueprint style, technical illustration, exploded view of keyboard components, white background, precise line work, engineering diagram aesthetic"
There's a nuance that nobody talks about - the aspect ratio completely changes how the AI interprets the prompt. A 16:9 prompt will push the keyboard toward a wide product shot composition, while 4:3 feels more like a standard photograph, and 1:1 centers everything. If you're generating for social media specifically, lock in your aspect ratio before writing the prompt instead of after. The framing decisions affect which details the AI chooses to emphasize. My actual workflow takes about 12 to 15 minutes per batch. I generate eight variations at 1024x1024, pick the one that's structurally closest to what I want, then upscale it and run a second pass focusing only on the areas that need fixing. Using the V6 or Niji prompts in Midjourney saves me a step because they handle small repetitive geometry better than the default model. Stable Diffusion users should be running ControlNet with a depth map if they're generating keyboards regularly. It cuts the hallucination rate from something like 70% down to maybe 20%. The main bottleneck is consistency across multiple images for the same project. If you need five shots of the same keyboard from different angles, the AI will redesign it each time. I solved this by generating one hero image, extracting the keycap pattern and color scheme into a reference, then using that as an image prompt for the remaining shots. It keeps the visual identity coherent without requiring me to describe the exact same keyboard in different words every time.
If you need publication-quality renders, expect to spend the full 15 minutes per image going through the iterative process. These aren't one-prompt-and-done kind of results. But once you have the system dialed in, you can produce a solid batch faster than most people can assemble a keyboard from parts. That's honestly the real value here - not the individual images, but the ability to generate enough variations to pick from without being locked into whatever the AI throws at you first try.
