Building

How to prompt AI well: what the evidence says works (and what doesn't)

What prompting actually is, which techniques hold up under testing, which popular tricks don't, and a simple method you can use this week.

Manelink Research28 September 2026 · 17 sources · 10 min read

In brief

Prompt an AI the way you would brief a capable new colleague. Say what the task is, who it is for and what good looks like, put long material before the question, and show one to three varied examples. Tips, threats, expert personas and “think step by step” add little on current models. Test any prompt you reuse on real cases before you trust it.

Most people's first experience of AI is a text box and a blinking cursor. What you type into that box, the prompt, is the only steering wheel you have. Unsurprisingly, a whole industry of advice has grown up around it: magic phrases, "act as an expert" templates, offers to tip the model, even threats. Some of that advice is grounded in research. A lot of it was true for older models and quietly stopped being true. Some of it was never true.

This article sorts the evidence from the folklore. You'll get a plain-English explanation of why wording matters to a language model at all, what peer-reviewed and lab research says about the techniques people swear by, where prompting runs out of road, and a short, practical method for writing prompts that works across today's major models.

What a prompt actually does

A large language model (LLM) is a system trained to predict the next piece of text, one "token" (roughly a word or word-fragment) at a time. A prompt is everything the model reads before it starts writing: your question, any instructions, pasted documents, examples, and in most products a hidden "system prompt" written by the developer. The model doesn't look anything up in a manual of your intentions. It continues the text in the way that seems most likely given everything in front of it.

That's why wording matters. A prompt doesn't just ask a question. It sets up the situation the model thinks it's in. The 2020 paper that introduced GPT-3 made this concrete. Its authors showed that a large enough model could perform new tasks from a plain-text description plus a handful of worked examples placed in the prompt, with no retraining at all1. Researchers call this in-context learning: the model picks up the pattern of the task from the context you provide. Zero examples is "zero-shot" prompting. A few examples is "few-shot."

Two things follow. First, a prompt is a form of programming in plain language, and like any program it can be clear or ambiguous. Second, because the model is matching patterns rather than understanding your goals the way a colleague would, it can be thrown off by details that look irrelevant to a human. Both points show up repeatedly in the evidence.

The techniques with real evidence behind them

The research literature on prompting is large. A 2024 systematic review led by researchers at the University of Maryland catalogued 58 distinct text-prompting techniques and a vocabulary of 33 terms, partly because the field had so much conflicting jargon5. You don't need 58. A handful have strong support.

Show, don't just tell (few-shot examples). Giving the model examples of the input and the output you want remains one of the most dependable techniques, going back to the GPT-3 work1. But a 2022 study from the University of Washington and Meta found something surprising: across 12 models, randomly scrambling the answers in the examples barely hurt performance. What the examples mainly taught the model was the format, the range of possible answers, and the kind of input to expect4. In practice, that means examples are best at showing shape (length, structure, tone, labels) and you should vary them so the model doesn't copy an accidental pattern. Anthropic's current guidance recommends three to five relevant, diverse examples, clearly marked off from the instructions12.

Ask for step-by-step reasoning, at least on older models. In 2022, Google researchers showed that including worked examples with the reasoning written out, a technique called chain-of-thought prompting, let a 540-billion-parameter model reach state-of-the-art accuracy on a benchmark of grade-school maths word problems using just eight examples2. The same year, a University of Tokyo and Google team found that simply adding "Let's think step by step" raised one model's accuracy on a maths benchmark from 17.7% to 78.7%3. These are two of the most influential results in the field. As the next section shows, their value has shrunk a lot as models have changed.

Be clear, specific and give the reason. This sounds obvious, yet it is where most everyday prompts fail. Anthropic's documentation offers a useful test: show your prompt to a colleague with minimal context on the task, and if they'd be confused, the model will be too12. It also recommends explaining why a rule exists. "Never use ellipses" is weaker than "this will be read aloud by a text-to-speech engine, so never use ellipses, since it won't know how to pronounce them," because the model can generalize from the reason12. The same source advises stating what to do rather than what not to do: "write in flowing prose paragraphs" works better than "don't use markdown."

Put long material first and the question last. A 2024 Stanford-led study in Transactions of the Association for Computational Linguistics found that models use information best when it sits at the beginning or end of a long input, and noticeably worse when the key fact is buried in the middle11. Anthropic reports, from its own internal testing, that placing documents at the top and the question at the end "can improve response quality by up to 30 percent," especially with multiple documents12. That figure is a vendor's internal test, not an independent study, but the direction is consistent with the peer-reviewed work.

Structure the prompt. Separating instructions, background, examples and input with clear labels or headings (Anthropic suggests XML-style tags such as <instructions> and <document>) reduces the chance the model confuses one part for another12. OpenAI's current guide for its GPT-5.6 model recommends a similar skeleton: role, goal, success criteria, constraints, tools, output format and when to stop14.

The evidence on popular tricks

Since early 2025, Wharton's Generative AI Labs at the University of Pennsylvania has run a series of "Prompting Science Reports" that test common advice with a rigour most prompting tips never get: hard benchmarks, several current models, and each question asked dozens or hundreds of times so random variation can be measured. Their results are a useful corrective. The main benchmark used, GPQA Diamond, is a set of 198 PhD-level multiple-choice questions in biology, physics and chemistry, so the findings speak most directly to hard factual and reasoning questions, not to writing or creative work.

"Think step by step" now adds little. Testing chain-of-thought instructions with 25 runs per question, the team found average gains of 4.4 to 13.5 percentage points for non-reasoning models, at the cost of responses taking 35% to 600% longer. For reasoning models, which are trained to work through problems internally before answering, the gains were 2.9 and 3.1 points for two OpenAI models and a 3.3-point drop for one Google model, with responses 20% to 80% slower8. Many models already reason step by step on their own, so asking them to do it again mostly buys latency and cost.

Tipping and threatening don't work. Offering the model money or threatening it (advice once endorsed publicly by Google co-founder Sergey Brin) "generally has no significant effect on benchmark performance," across two hard benchmarks9. Individual questions moved up or down, but in ways nobody could predict in advance.

"You are a world-class expert" doesn't improve accuracy. Across six models and two graduate-level benchmarks, expert personas "generally did not improve accuracy relative to a no-persona baseline." Assigning low-knowledge personas, such as a layperson or a child, generally made results worse10. The authors note that personas may still be useful for setting tone or audience, which is a different job from getting facts right.

Politeness is a coin flip at the level of single questions. Comparing "please" with "I order you," the researchers found swings of up to 60 percentage points on individual questions in either direction, which washed out across the whole test set7. The same report found that whether a model looks "good" depends heavily on the bar you set: at a 51% majority-correct threshold the models looked strong. Demanding 100% consistency across 100 tries, they barely beat random guessing on these very hard questions7. Removing formatting instructions, by contrast, consistently hurt performance7.

Tiny formatting changes can matter enormously, sometimes. A 2024 ICLR paper from the University of Washington changed only meaning-preserving details, such as separators, spacing and capitalization, in few-shot prompts. On one open model (Llama-2-13B) accuracy varied by up to 76 percentage points depending on format, and the sensitivity didn't go away with larger models or more examples6. That was an older, smaller model tested on classification tasks. Frontier models in 2026 are generally more robust. The lesson that holds up is about measurement: judge a prompt on many runs, not one lucky output.

Why good prompting is changing

Two shifts stand out in 2025–2026.

Newer models follow instructions more literally, including bad ones. OpenAI's GPT-5 guide warned in August 2025 that "poorly-constructed prompts containing contradictory or vague instructions can be more damaging to GPT-5 than to other models," because the model spends effort trying to reconcile them13. Its July 2026 guidance for GPT-5.6 goes further: "conflicting rules can create more instability than missing detail," and it recommends describing "the destination rather than prescribing every step." OpenAI reports that in its internal testing, leaner system prompts improved evaluation scores by roughly 10–15% while cutting total tokens by 41–66%14. Anthropic similarly advises dialling back shouty emphasis ("CRITICAL: You MUST…") written for older models, because current models may now over-apply it12. Both are vendor claims about their own models, but they point the same way: the all-caps, belt-and-braces prompt is becoming a liability.

Prompting is becoming part of "context engineering." For AI agents that run over many steps with tools and documents, the prompt you type is a small part of what the model sees. Anthropic now describes the wider discipline as curating "the optimal set of tokens (information)" the model has at each step, and recommends system prompts pitched at the "right altitude": specific enough to guide behaviour, flexible enough to leave the model room to use judgment17. That topic has its own article in this series.

Meanwhile, machines are getting good at writing prompts too. A Google DeepMind method called OPRO, published at ICLR 2024, used one model to iteratively rewrite prompts for another and beat human-written prompts by up to 8% on maths problems and up to 50% on a set of hard reasoning tasks16. For repeated, measurable tasks, automatic prompt optimization is now a practical tool.

Limits and open questions

Your skill still matters, but not everywhere. The strongest human evidence comes from a large experiment by researchers at the University of Maryland, MIT, Microsoft Research and others, with 3,750 participants writing nearly 37,000 prompts15. In one task, people tried to recreate a target image using either an older or newer image model. Those given the better model wrote prompts 24% longer, and on this structured task their changing prompts accounted for roughly half of the performance gain from the upgrade. When the system silently rewrote their prompts, 58% of the new model's benefit was lost. On open-ended creative tasks, though, the model's capability dominated and prompting mattered much less. This study used image generators, not chatbots, and the latest version is a preprint that hasn't completed peer review.

Results are contingent. The Wharton team's own summary is that prompt engineering is "complicated and contingent"7. A trick that helps one model can hurt another, and a prompt that works for one question can fail on the next. That's why published lists of "the 10 best prompts" date so quickly.

Most research tests right-or-wrong questions. Benchmarks like GPQA are convenient because answers can be marked automatically. We have much weaker evidence on what makes prompts good for drafting, analysis, summarization or strategy work, which is most of what people actually use AI for.

No prompt removes the underlying limits. Prompting changes how a model uses what it knows. It can't give a model knowledge it lacks, and it doesn't eliminate confident errors. For facts that matter, you still need sources you can check.

What this means for you

Here is a simple method that reflects the evidence, and that works across today's major assistants:

  1. Brief it like a capable new colleague. Say what the task is, who it's for, what "good" looks like, and why. If a smart person with no context would need to ask questions, add the answers.
  2. Paste the material first, ask last. For anything longer than a page, put the documents at the top and your question or instructions at the end, with a clear label on each part.
  3. Show one to three examples of the shape you want, and vary them. Use them for format and tone, not as a way to smuggle in facts.
  4. Drop the rituals. Skip the tips, threats, all-caps warnings and "you are a world-class expert" preambles unless you've tested that they help on your task. Ask for step-by-step reasoning only if you're using an older, non-reasoning model.
  5. Test before you trust. If you'll reuse a prompt, run it on five to ten real cases, including a couple of hard ones, and read the results. One good answer tells you little.

Key takeaways

  • A prompt sets up the situation a language model responds to. Models learn a task's format from examples and instructions in context.14
  • Clarity, context about why, well-chosen examples, clear structure, and putting long material before the question have the best support.41112
  • Tipping, threats, expert personas and "think step by step" add little or nothing to accuracy on current models, and results vary unpredictably by question and model.78910
  • Newer models reward shorter, conflict-free prompts that describe the goal rather than every step.1314
  • Prompting skill matters most on structured, well-defined tasks. Always test a reusable prompt on several real cases.615

Sources

  1. Brown, T. B., Mann, B., Ryder, N., et al. (2020). Language models are few-shot learners. Advances in Neural Information Processing Systems 33. arxiv.org/abs/2005.14165
  2. Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., Chi, E., Le, Q., & Zhou, D. (2022). Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems 35. arxiv.org/abs/2201.11903
  3. Kojima, T., Gu, S. S., Reid, M., Matsuo, Y., & Iwasawa, Y. (2022). Large language models are zero-shot reasoners. Advances in Neural Information Processing Systems 35. arxiv.org/abs/2205.11916
  4. Min, S., Lyu, X., Holtzman, A., Artetxe, M., Lewis, M., Hajishirzi, H., & Zettlemoyer, L. (2022). Rethinking the role of demonstrations: What makes in-context learning work? Proceedings of EMNLP 2022. arxiv.org/abs/2202.12837
  5. Schulhoff, S., Ilie, M., Balepur, N., et al. (2024, rev. 2025). The Prompt Report: A systematic survey of prompt engineering techniques. arXiv:2406.06608. arxiv.org/abs/2406.06608
  6. Sclar, M., Choi, Y., Tsvetkov, Y., & Suhr, A. (2024). Quantifying language models' sensitivity to spurious features in prompt design, or: How I learned to start worrying about prompt formatting. ICLR 2024. arxiv.org/abs/2310.11324
  7. Meincke, L., Mollick, E., Mollick, L., & Shapiro, D. (2025, March 4). Prompting Science Report 1: Prompt engineering is complicated and contingent. Wharton Generative AI Labs. gail.wharton.upenn.edu/research-and-insights/tech-report-prompt-engineering-is-complicated-and-contingent
  8. Meincke, L., Mollick, E., Mollick, L., & Shapiro, D. (2025, June 8). Prompting Science Report 2: The decreasing value of chain of thought in prompting. arXiv:2506.07142. arxiv.org/abs/2506.07142
  9. Meincke, L., Mollick, E., Mollick, L., & Shapiro, D. (2025, August 1). Prompting Science Report 3: I'll pay you or I'll kill you – but will you care? arXiv:2508.00614. arxiv.org/abs/2508.00614
  10. Basil, S., Shapiro, I., Shapiro, D., Mollick, E., Mollick, L., & Meincke, L. (2025, December 5). Prompting Science Report 4: Playing pretend: Expert personas don't improve factual accuracy. arXiv:2512.05858. arxiv.org/abs/2512.05858
  11. Liu, N. F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., & Liang, P. (2024). Lost in the middle: How language models use long contexts. Transactions of the Association for Computational Linguistics, 12, 157–173. doi.org/10.1162/tacl_a_00638
  12. Anthropic. (2026). Prompting best practices. Claude Platform Docs (accessed September 28, 2026). platform.claude.com/docs/en/build-with-claude/prompt-engineering/claude-prompting-best-practices
  13. Kotha, A., Lee, J., Zakariasson, E., & Kavanaugh, E. (2025, August 7). GPT-5 prompting guide. OpenAI Cookbook. developers.openai.com/cookbook/examples/gpt-5/gpt-5_prompting_guide
  14. OpenAI. (2026). Prompt guidance for GPT-5.6. OpenAI API Docs (accessed September 28, 2026). developers.openai.com/api/docs/guides/prompt-guidance-gpt-5p6
  15. Jahani, E., Manning, B. S., Zhang, J., TuYe, H.-Y., Alsobay, M., Nicolaides, C., Suri, S., & Holtz, D. (2024, rev. January 7, 2026). Prompt adaptation as a dynamic complement in generative AI systems (earlier title: As generative models improve, people adapt their prompts). arXiv:2407.14333. arxiv.org/abs/2407.14333
  16. Yang, C., Wang, X., Lu, Y., Liu, H., Le, Q. V., Zhou, D., & Chen, X. (2024). Large language models as optimizers. ICLR 2024. arxiv.org/abs/2309.03409
  17. Rajasekaran, P., Dixon, E., Ryan, C., & Hadfield, J. (2025, September 29). Effective context engineering for AI agents. Anthropic Engineering. anthropic.com/engineering/effective-context-engineering-for-ai-agents