Prompt engineering is the practice of writing structured instructions that shape what a language model produces, covering zero-shot directives, few-shot examples, and chain-of-thought scaffolds used across GPT-4, Claude, and Gemini. OpenAI defines it as writing effective instructions so a model consistently generates content that meets your requirements, and OpenAI emphasizes that output is non-deterministic, which makes prompting a mix of art and science (OpenAI Platform, Prompt engineering guide). The techniques that move output quality the most are not clever phrasings. They are specificity, examples, role framing, output format constraints, reasoning scaffolds, and an evaluation loop that catches regressions when prompts or model versions change. Each major vendor exposes a slightly different control surface: OpenAI through the instructions parameter, Anthropic through the system prompt, and Google through clear-and-specific guidance plus the Gemini 3 thinking_level parameter. The patterns below cover the working stack practitioners apply across all three.
What Prompt Engineering Actually Controls

Prompt engineering controls the variables a model treats as guidance when generating a response: the role, the task description, the output format, and the examples that anchor what a good response looks like. The model is not consulting a knowledge base for each request. It is sampling from a probability distribution conditioned on the input tokens, which is why OpenAI frames model output as non-deterministic and why prompt design is an iterative discipline (OpenAI Platform, Prompt engineering guide). Google describes generative language models as advanced auto completion when given partial content, which makes clear, specific instructions the most reliable lever for steering behavior (Google AI for Developers, Gemini API prompting strategies). The practical implication is that vague instructions surface vague completions. A prompt that says "summarize this document" leaves the audience, length, tone, and structure unspecified; the model fills those gaps with its prior. A prompt that says "summarize this document for a product manager in five bullet points, each under twenty words, focused on customer-impact risks" narrows the plausible completions to a much tighter band. The taxonomy practitioners apply spans Zero-Shot prompting, Few-Shot prompting, chain of thought reasoning, ReAct loops, Self-Consistency sampling, and Tree of Thoughts search, with each named pattern targeting a different failure mode. For background on how the underlying model learns from data in the first place, see how machine learning models learn from data.
- Role: the identity or persona the model adopts for the response, set through role prompting or a System Prompt.
- Task instruction: the explicit description of what the model should do, ideally specific about scope, audience, and constraints.
- Output format: the shape the response must take, expressed as a schema, template, or worked example, often a JSON contract.
- Examples: input-output pairs that anchor the model to the expected pattern, central to Few-Shot prompting.
- Reasoning scaffold: an instruction to think step-by-step, the foundation of chain of thought, ReAct, and Tree of Thoughts techniques.
Zero-Shot and Few-Shot Techniques
Prompt engineering splits into two foundational modes based on whether examples are supplied: Zero-Shot sends only a task instruction, while Few-Shot prepends two to five worked input-output pairs that demonstrate the expected response shape. Zero-Shot relies on the model's pretraining alone, which works well for common tasks that the model has seen many variants of during training, such as summarization, translation, or sentiment classification. Few-Shot becomes valuable when the task is unusual, when the desired output structure is hard to describe in words, or when the model keeps drifting toward a default style that does not match the application. Anthropic explicitly recommends constraining responses with examples as one of the highest-impact moves for output consistency on Claude (Anthropic Docs, Increase output consistency). The mechanics matter. Each example narrows the range of plausible completions because the model treats the demonstrated pattern as the in-context distribution. Two well-chosen examples often outperform a paragraph of prose instructions because the model can pattern-match rather than parse natural language into behavior. Google's guidance for GPT-4 and Gemini users reinforces the same principle through its emphasis on clear and specific instructions as the foundation of effective prompt design (Google AI for Developers, Gemini API prompting strategies).
| Attribute | Zero-shot prompting | Few-shot prompting |
|---|---|---|
| Examples included | None | Two to five input-output pairs |
| Best for | Common tasks the model has seen in pretraining | Unusual formats or domain-specific patterns |
| Token cost | Lower; only the instruction occupies the context window | Higher; examples consume context window space |
| Output variance | Wider; the model fills gaps from its prior | Narrower; examples anchor the response shape |
| Failure mode | Drift toward default style or generic phrasing | Overfitting to example content rather than the pattern |
Chain-of-Thought and Reasoning Scaffolds
Prompt engineering unlocks multi-step reasoning by instructing the model to work through intermediate steps before giving a final answer, a pattern Google calls chain-of-thought (CoT) and that Gemini 3 models expose directly via a thinking_level parameter. In practice, a chain of thought prompt asks the model to lay out its intermediate reasoning before stating an answer, which improves accuracy on multi-step arithmetic, logic, and planning tasks. Google's Gemini 3 documentation states that Gemini 3 series models use dynamic thinking by default to reason through prompts, and that the thinking_level parameter controls the maximum depth of the model's internal reasoning process before it produces a response (Google AI for Developers, Gemini 3 model documentation). Lower thinking levels minimize latency and cost; higher levels trade those for deeper reasoning. The pattern extends through ReAct (Reasoning and Acting), where the model interleaves reasoning traces with tool calls, observing the result of each action before deciding the next step. ReAct turns a single completion into a controlled loop, which is the basis for most RAG and agentic workflows. Two adjacent techniques harden the same idea: Self-Consistency samples multiple chain of thought traces and majority-votes the answer, while Tree of Thoughts branches the reasoning into parallel paths and prunes weak ones. Other documented patterns include Least-to-Most prompting, Step-Back prompting, Reflexion, Generated Knowledge prompting, Skeleton-of-Thought, and Automatic Prompt Engineer, each targeting a distinct reasoning or reliability failure. Frameworks such as DSPy and LangChain expose these patterns as composable modules over the GPT-4, Claude, and Gemini APIs. CoT and ReAct work best on larger models with the capacity to follow long reasoning chains; smaller instruction-tuned models often skip the trace and jump to a conclusion, which is why Anthropic recommends pairing chain of thought with system prompt role-framing and format constraints to reinforce the behavior (Anthropic Docs, Increase output consistency). For deeper background on tool-use loops, see the retrieval-augmented generation explainer.
- Instruct the model to reason step-by-step: a single line such as "think through this problem step by step before answering" reliably surfaces intermediate reasoning on capable models.
- Provide a worked reasoning example: Few-Shot CoT, where one demonstration shows the full reasoning trace plus the final answer, anchors the model to the expected pattern.
- Use the platform's thinking control where available: Gemini 3 exposes a thinking_level parameter; raise it for hard problems and lower it for latency-sensitive paths.
- Move to ReAct for tool-using workflows: instruct the model to alternate Thought, Action, and Observation steps so each tool call is grounded in an explicit reasoning trace.
- Escalate to Self-Consistency or Tree of Thoughts for hard problems: sample multiple traces and aggregate, or branch the reasoning, when a single pass is unreliable.
- Chain prompts for complex tasks: Anthropic recommends breaking complex tasks into smaller subtasks across multiple prompts rather than asking for everything in one shot.
System Prompts and Role Framing
Prompt engineering at the session level happens inside the System Prompt, where role framing, behavioral constraints, and persistent tone instructions are set once and govern every turn that follows. Anthropic states that system prompts define Claude's role and personality, and recommends using them for role-based applications where a stable identity and persona matter (Anthropic Docs, Increase output consistency). OpenAI exposes the same control surface through the instructions parameter, which gives the model high-level instructions on tone, goals, and examples of correct responses; OpenAI further states that any instructions provided this way take priority over a prompt in the input parameter (OpenAI Platform, Prompt engineering guide). The practical pattern is layered. The System Prompt carries the role, the operating rules, the output format defaults, and any safety constraints that must hold across the session. The user prompt carries the specific task. Role prompting inside the System Prompt is more than flavor; it shifts which patterns from pretraining the model surfaces, which is why "you are a senior security engineer reviewing this code" produces different output than the same request without the role. For a comparison of how different model families handle instruction following, see comparing language models and their instruction-following capabilities.
- Set the role in the System Prompt: name the persona, the expertise level, and the audience the model is writing for.
- State the operating rules: what the model should always do, never do, and how it should handle ambiguity or missing information.
- Anchor the default output format: declare the shape of normal responses so each user turn does not have to re-specify it.
- Reserve user prompts for task details: per-request specifics belong in the user turn, not in repeated re-statements of the role.
- Keep the System Prompt stable: changing it mid-session resets the operating context and can introduce inconsistency across turns.
Output Format and Consistency Constraints
Prompt engineering gains consistency through format constraints that tell the model exactly what shape the output must take, whether a JSON object, a numbered list, or a custom template. Anthropic recommends precisely defining the desired output format using JSON, XML, or custom templates so that Claude understands every output formatting element you require (Anthropic Docs, Increase output consistency). The same source draws a sharp line for stricter cases: when guaranteed JSON schema conformance is required, Anthropic recommends using Structured Outputs instead of prompt engineering techniques, because schema-enforced decoding removes the residual variance that prose instructions cannot. The same principle applies on the OpenAI side, where the instructions parameter can carry examples of correct responses that lock the response shape across sessions. Format constraints reduce hallucination surface area in a measurable way. A model asked for "a list of risks" can return a paragraph, three bullets, or a table; a model asked for "a JSON array of objects with fields risk, severity, and mitigation, where severity is one of low, medium, high" returns exactly that or fails in a structured, catchable way. For why constraint-tightening reduces output risk overall, see AI output risks including hallucination and why constraints reduce them.
- Specify the container: JSON object, XML document, Markdown table, or a labeled template; pick one and state it.
- Declare every field: name each key, its type, and the allowed value range so the schema is unambiguous.
- Show one worked example: a single concrete output anchors the model more reliably than a paragraph of schema prose.
- Use Structured Outputs when conformance is non-negotiable: Anthropic recommends Structured Outputs over prompt-only techniques for guaranteed JSON schema validity.
- Constrain enumerations explicitly: list the exact allowed strings for any categorical field rather than describing them in prose.
- Define error behavior: tell the model what to emit when input is malformed or a field cannot be determined.
Evaluating and Iterating Prompts
Prompt engineering is not complete at first draft: OpenAI recommends building evals that measure prompt behavior across a representative sample before any iteration or model version change. OpenAI describes the practice directly as building evals that measure the behavior of your prompts so you can monitor prompt performance as you iterate, or when you change and upgrade model versions (OpenAI Platform, Prompt engineering guide). The same documentation recommends pinning production applications to specific model snapshots to preserve consistent behavior, a practical safeguard against silent drift when a vendor rolls a new model version into the default endpoint. The evaluation loop is the difference between prompt design as folklore and as a discipline. A representative test set of inputs, paired with reference outputs or graded rubrics, lets a team measure whether a prompt change improved behavior on the cases that matter or only on the one example the change was tuned against. The pattern is the same across GPT-4, Claude, and Gemini: write the prompt, run it against the eval set, score the results, and iterate. Frameworks such as DSPy push this further by treating prompts as compiled artifacts whose parameters are optimized against a metric. Anthropic's guidance reinforces the loop through its emphasis on chaining prompts for complex tasks and validating each subtask in isolation. Model output remains non-deterministic even with the strongest prompt, so the eval set also functions as a regression guard when the model snapshot is upgraded.
- Build a representative input set: twenty to fifty examples covering the easy, hard, and edge cases the prompt will see in production.
- Define a scoring rubric: exact-match for structured outputs, graded scoring or LLM-as-judge for open-ended responses.
- Pin the model snapshot: per OpenAI guidance, lock the production prompt to a specific model version so behavior changes are intentional.
- Run the eval before and after each prompt change: compare aggregate scores rather than spot-checking one or two outputs.
- Re-run the eval on every model upgrade: treat a model version bump as a code change that requires the same regression coverage.
Prompt Engineering Across Models
Prompt engineering translates differently across GPT-4, Claude, and Gemini because each model exposes a distinct set of control surfaces for instruction priority, reasoning depth, and output structure. OpenAI's instructions parameter takes priority over input-prompt text and carries high-level behavior including tone, goals, and examples of correct responses (OpenAI Platform, Prompt engineering guide). Anthropic's Claude uses a System Prompt to set role and personality, with format constraints expressed through JSON, XML, or custom templates (Anthropic Docs, Increase output consistency). Google's Gemini 3 series uses dynamic thinking by default and exposes a thinking_level parameter that controls reasoning depth, with lower levels minimizing latency and cost (Google AI for Developers, Gemini 3 model documentation). The differences matter for portability. A prompt that relies on the instructions parameter to override later input text behaves differently when ported to Claude, where the same separation is handled through the System Prompt. A reasoning-heavy prompt tuned for Gemini 3 may need an explicit chain of thought instruction on a model without a thinking-level control. The reusable core across all three is the same: specific instructions, role framing, examples, and an explicit output format. The platform-specific surfaces sit on top of that core. For broader selection criteria when choosing a platform, see how to evaluate AI and machine learning platforms for your use case.
| Control surface | OpenAI GPT-4 | Anthropic Claude | Google Gemini 3 |
|---|---|---|---|
| Persistent instructions | instructions parameter; takes priority over input | System prompt; sets role and personality | System instruction field; clear and specific guidance |
| Reasoning depth control | Prompt-driven chain-of-thought | Prompt-driven chain-of-thought; pair with role framing | thinking_level parameter; dynamic thinking by default |
| Structured output | Structured Outputs with JSON schema | Structured Outputs for guaranteed schema conformance | JSON mode and response schema controls |
| Version pinning | Model snapshot identifiers | Versioned Claude model identifiers | Versioned Gemini model identifiers |
References
- OpenAI Platform, Prompt engineering guide
- Anthropic Docs, Increase output consistency
- Google AI for Developers, Gemini API prompting strategies
- Google AI for Developers, Gemini 3 model documentation
Further reading
Frequently Asked Questions
What is prompt engineering?
Prompt engineering is the practice of writing effective instructions so a model consistently generates the output you need. It covers technique selection (zero-shot, few-shot, chain-of-thought), format constraints, system prompt design, and iterative evaluation to close the gap between an initial response and a production-ready result.
What is the difference between zero-shot and few-shot prompting?
Zero-shot prompting asks the model to complete a task with no worked examples; few-shot prompting supplies two to five input-output pairs before the request. Per OpenAI guidance, few-shot examples narrow the range of plausible completions, which reduces variance and lifts consistency on tasks where the expected format is hard to describe in words alone.
Does chain-of-thought prompting work on all models?
Chain-of-thought prompting is most effective on large models that have enough capacity to follow multi-step reasoning traces. Smaller or instruction-limited models often ignore the reasoning scaffold and skip straight to a conclusion; Anthropic recommends pairing chain-of-thought with system prompt role-framing and output format constraints to reinforce the behavior.
When should I use a system prompt instead of a user prompt?
Use a system prompt to set persistent role, tone, and behavioral constraints that apply across every turn in a session. Anthropic states that instructions in a system prompt take priority over the input prompt and provide a stable identity and persona for role-based applications, whereas user-turn instructions handle per-request task details.









