Prompting isn't magic, it's just pattern matching at scale

The paper Language Models Are Few Shot Learners came out in 2020 and quietly changed how everyone approached GPT-like systems. Before that, the default assumption was that you needed to fine-tune a model for any new task. The paper showed that you could get surprisingly competent behavior just by describing the task in natural language and providing a handful of examples. Not dozens. Not hundreds. A few. Here is how it actually works under the hood. You format your prompt with input-output pairs that demonstrate the pattern you want, then append the query you actually want answered. The model completes the next token based on the statistical patterns it saw during pre-training, guided by the examples you provided in-context. That is what "few-shot" means: the model learns the task from the prompt itself rather than from gradient updates to its weights. I spent about three weeks trying to get a GPT-3 variant to consistently output structured JSON for a data extraction pipeline I was building. Zero-shot did not work at all. The model would output the right data but wrap it in conversational hedging like "Here is the JSON you requested:" followed by malformed objects. Five-shot worked reliably. I settled on providing exactly five clean input-output pairs with consistent formatting, including one edge-case example where the text contained missing or ambiguous entities. After that, the success rate jumped from maybe 40 percent to over 90 percent across my test set.

The key detail nobody emphasizes enough is that the examples need to match the difficulty of your actual query. If your few-shot examples are all simple, trivially clear cases and your real input is messy or ambiguous, the model will default to treating it like the easy examples and you will get garbage output. I learned that the hard way when a client asked me to extract names from legal contracts. The examples I wrote were from straightforward employment agreements. The actual contracts had nested clauses, redacted entities, and cross-references. I got consistently wrong results until I added two examples pulled directly from the messy contract subset. That one change fixed it.

How to set this up in practice

You do not need a GPU cluster or a custom training run. Pick your model through whatever API you have access to. OpenAI, Anthropic, or an open-source model served through vLLM or similar all support this approach. The main variable you control is the prompt structure itself. Build your prompt in this order: instructions if needed, then examples, then the target query. Keep the examples tightly aligned with the desired output format. If you want a classification, show the label format. If you want a generation task, show the exact structure. One example I have found useful is to include an example where the answer is "I don't know" or a similar refusal, especially for factual queries where the model might otherwise hallucinate confidently. That single example reduced my hallucination rate by roughly a third in a medical Q&A system I was testing. Temperature matters more than people expect. At 0.0 you get the most deterministic output, which is usually what you want for structured tasks. At higher temperatures the model explores more but consistency drops quickly. For the JSON extraction work I mentioned, I ran temperature at 0.0 and got stable results. When I bumped it to 0.7 just to see what would happen, the output quality degraded significantly even though the information content was roughly the same.

Get the Full Details

Forever a student: My 5 language learning tips
Forever a student: My 5 language learning tips

What people get wrong about few-shot learning

One common mistake is assuming more examples is always better. I tested six, ten, and fifteen-shot prompts on a summarization task and found that performance peaked around five to seven examples and then plateaued or slightly declined. The reason is that later examples compete for attention in the context window and the model starts to treat them as variations of a different task rather than reinforcement of the same one. Your mileage will vary depending on the model and task, but there is a practical ceiling. Another thing that catches people off guard is positional bias. Models tend to weight the first and last examples in your prompt more heavily than the middle ones. I verified this empirically by shuffling the order of five-shot examples across 200 test queries and measuring accuracy. Swapping the position of a difficult example from position three to position one improved overall accuracy by about eight percentage points. If you have a particularly tricky example, put it first or last, not buried in the middle. There is also the question of instruction tuning. Base models like GPT-3 175B respond differently to few-shot prompting than instruction-tuned variants like GPT-3.5-Turbo. The base model needs the task demonstrated through examples more explicitly. The instruction-tuned model can often follow a direct command without as many examples. This means you might use three-shot with an instruction-tuned model where you would need five-shot with a base model for the same quality. The tradeoff is that instruction-tuned models sometimes resist unusual formatting requests that a base model would happily follow if shown an example.

When few-shot breaks down

This approach does not work for everything. Tasks that require extensive domain-specific knowledge beyond what the model has seen during pre-training will still fail regardless of how many examples you provide. I tried using few-shot prompting on a legal compliance task where the model needed to apply a specific jurisdiction's recently updated regulations. No amount of examples could compensate for the model not knowing the actual regulation text. In that case, retrieval-augmented generation was the only viable path and it cut my error rate from about 35 percent down to under 8 percent. Long-context inconsistency is another real limitation. As your prompt grows beyond roughly 4000 tokens with many examples, the model starts dropping or misattending to earlier examples. I had a financial analysis prompt that included twelve shot examples plus long context windows and the output quality became erratic. Switching to a shorter six-shot prompt with the most representative examples restored reliability. The fix was not adding more context, it was pruning it. Cost and latency scale linearly with the number of tokens in your prompt. A fifteen-shot prompt on a 175B parameter model is noticeably more expensive and slower than a three-shot version. For high-throughput applications, this adds up fast. I moved a production pipeline from eight-shot to four-shot after benchmarking and the quality difference was within the noise margin while the per-request cost dropped by about 40 percent. Always benchmark your shot count against a held-out set before committing to a number.

Quick reference for getting started

Pick a model with sufficient capability for your task. Base models need more examples, instruction-tuned models need fewer. Format your prompt with clear input-output pairs. Include at least one example that matches the hardest case in your actual workload. Keep your total example count between three and seven for most tasks. Use temperature 0.0 for structured output. Position difficult examples at the start or end of your prompt. Measure performance on a held-out set before deploying. If the task requires knowledge the model likely does not have, consider retrieval or fine-tuning instead. The original paper is available through the OpenAI research page and the arXiv listing. Most modern model providers document their few-shot prompting approach in their developer docs now, though the details vary by provider. The core principles remain the same regardless of which API you use.

The Importance of Language – It Matters
The Importance of Language – It Matters