How People Actually Use Prompt Tools on a Steam Deck
Most people who ask about this aren't looking for anything fancy. They want to run AI models locally on the Deck, type prompts into a terminal, and get answers back without buying a laptop or leaving home. The Steam Deck runs Linux, so the usual command-line tools work fine. It's not seamless. But it gets the job done. The standard approach is installing Ollama, downloading a model, and chatting from the terminal. You switch to Desktop Mode, open a terminal, and run the install command. Ollama pulls in a small model—something like Phi-3 or Llama 3.2 3B—and then you just type your prompt. Responses come back at maybe 10 to 20 tokens per second on a healthy Steam Deck. Not blazing, but usable for casual tasks.
Steam Deck Prompts Simple workflow
I put together a simple wrapper script a while back because I got tired of typing the same commands every time. The idea was straightforward: a script that takes your prompt as an argument, runs it through Ollama, and saves the conversation to a text file so you can look back at it later. I wrote it in bash. It's roughly 40 lines, nothing elaborate. The script uses Ollama's API endpoint, sends the message, captures the response, and appends both to a log. I keep all my logs in a folder called ~/deck-prompts. Here's the basic structure: Install Ollama first:
curl -fsSL https://ollama.com/install.sh | sh Then pull a lightweight model: ollama pull phi3:mini
Get the Full Details

And run it: ollama run phi3:mini That last command drops you into an interactive chat. You type a prompt, hit Enter, and wait. If you want scripted use instead, you can call the API directly:
curl http://localhost:11434/api/generate -d '{"model":"phi3:mini","prompt":"explain recursion simply"}' This is where things get a bit messy. The Steam Deck's touch keyboard is awful for long prompts. I discovered this pretty quickly. After about ten minutes of fighting with the on-screen keyboard, I connected a cheap mechanical keyboard through USB-C. Suddenly the whole thing felt bearable. The keyboard issue is real and honestly the biggest barrier for most people. If you're not willing to carry an external keyboard, this workflow loses a lot of its appeal. Another thing nobody warns you about: thermal throttling. Running a local LLM for more than twenty minutes pushes the Deck's APU hard enough that it starts throttling. Token speeds drop from maybe 18 tokens per second down to 6 or 7. I learned this the hard way during a long session where I was feeding it code snippets and asking for refactors. The model didn't slow down in quality, just in speed. If you're doing quick one-off prompts, you won't notice. If you're running extended conversations, put the Deck on a flat surface, not in your lap, and maybe pop the back off for a few minutes to let it breathe.
For people who don't want to deal with Ollama at all, there's a lighter option called Jan. It has a GUI, which matters if you'd rather click through menus than type commands. It's slower to set up but more forgiving if you're not comfortable in a terminal. The trade-off is that it uses more RAM and CPU overhead, which makes theDeck warmer and the fans louder. Again, not ideal for handheld play, but fine if you're docked. There's also a project called Text generation webUI that some people port to the Deck. It's more feature-rich but significantly heavier. It supports larger models and has a browser-based interface. I tried it once. It ran, but the performance was rough. The Deck struggles with anything bigger than the 7B parameter range on this setup. For simple prompts and short responses, smaller models like Phi-3 or Qwen 2.5 7B work acceptably. Anything larger and you're just waiting around. One practical detail that matters more than it should: prompt formatting. Local models don't handle conversation history the way ChatGPT does. They usually expect the full context in each request unless you're using a framework that manages it for you. Ollama handles basic context internally, but it's limited. If you paste a huge block of text and then ask a question about it, the model will still process it. But if your prompt exceeds the context window—usually 8K tokens for the smaller models—the Deck either truncates it or throws an error. I keep a reference file of common prompt templates so I'm not rewriting them from scratch every time.

If you want something even simpler than all of this, there are pre-made prompt files you can just drop into a folder and feed to the model. A lot of community members share these on Reddit and GitHub. The quality varies wildly, but a few solid ones exist for code generation, creative writing, and general Q&A. I've found that the best results come from prompts that are specific and direct. Vague prompts waste tokens and time on a device that's already working this hard. One more thing worth mentioning: network prompts. If you don't want to run anything locally and just want to type prompts into an AI, you can access web-based services directly from the Steam Deck's browser. The interface is cramped but functional. Chrome and Firefox both work in Desktop Mode. Some people even use this as their primary method because it avoids the hardware limitations entirely. You just open the browser, go to whatever AI service you prefer, and type your prompt. No setup, no thermals, no token limits beyond what the service imposes. The whole setup takes about fifteen to twenty minutes from start to finish if you're familiar with the terminal. If you're not, give yourself an hour. Watch a couple of setup videos first. The community documentation is decent but scattered across different forums and GitHub repos.