Prompt engineering is the work of expressing the goal, relevant background, constraints, expected process, completion criteria, and output form so that a model is more likely to behave as intended across similar tasks. It is not a collection of magic phrases, and it does not require JSON. Good prose is enough to begin.
After this article: you should be able to turn a vague request into a clear task, then understand why a production prompt also needs instruction precedence, precise tool descriptions, versions, and regression checks.
Source note: concepts draw on official OpenAI and Anthropic engineering material. The implementation examples use Codex public source checked on September 20, 2026, with links pinned to that revision. The payment task and release process are teaching designs, not fixed Codex schemas or automatic guarantees. Message precedence follows the target API contract.
1. Understand prompts through one task
1.1 Improve one request in three passes
Suppose a payment test is failing. The shortest request is:
V0: “Fix the failing payment test.”
The goal is visible, but almost everything else is left to guesswork. The Agent may rewrite a public API, skip the test, or say “done” after only editing code. A second version adds the most important limits:
V1: “Find why the payment test fails, make the smallest change without changing public APIs, run the payment tests, and report the files changed and the test result.”
Now “fixed” has observable meaning. If the task is risky, a third version can also say what to do when a requirement cannot be met:
V2: “If the test cannot run or the fix requires a public-API change, stop and explain the blocker instead of claiming success.”
This progression is the core of prompt engineering: replace hidden assumptions with requirements that both the Agent and the reviewer can inspect.
1.2 A useful prompt answers six questions
- Goal: what outcome should be produced?
- Background: what situation and evidence matter?
- Constraints: what must not be changed or attempted?
- Method: which important steps should be followed?
- Done criteria: what evidence proves the task is complete?
- Output: what should the final report contain?
Not every task needs six labeled sections. The labels are a review tool. For the payment task, “fix the test” is the goal, the failure log is background, “do not change public APIs” is a constraint, “inspect before editing” is method, a passing payment-test command is the done criterion, and the change report is the output.
The six parts as one natural-language prompt: “Fix
refund_should_reject_expired_order. First inspect the failure log and related code, then make the smallest change directly related to the failure without changing public APIs. Run the payment tests afterward. If the test environment is unavailable or progress requires a public-API change, stop and explain the blocker. Report the cause, changed files, test command, and result.”
2. Turn a good instruction into a stable interface
2.1 Move from one good answer to stable behavior
Take one task: “Fix the failing payment test.” A model might edit immediately, inspect the failure first, or report that the test environment is unavailable. The engineering object is not one sentence but the stability of these observable behaviors across similar inputs: when to investigate, when to refuse, what evidence counts as done, how to react to tool failure, and what the final response must contain.
A review therefore cannot stop at a text diff. “Be more concise” or “act proactively” is an intention. It becomes an engineering result only when representative tasks show fewer wasteful calls, stronger constraint adherence, and stable outputs.

2.2 Once it becomes a product, separate content by change rate
| Content | Question it answers | Change rate | Preferred location |
|---|---|---|---|
| Stable policy | What is never allowed? When must a person take over? | Low | Centrally maintained stable instructions |
| Product template | What are the role, goal, done criteria, and output schema? | Low to medium | Versioned product template |
| Task parameters | Which object, scope, and preferences apply now? | High | Structured user/task input |
| Dynamic facts | What is the repository, customer, or outside world like now? | High and perishable | Context or tool observations |
A classic smell is hard-coding dynamic facts into stable instructions. “The release branch is release/42” expires quickly. The reverse smell is retrieving stable policy as optional context, allowing one missed retrieval to remove a rule. Separation by change rate keeps instructions stable and lets context provide current task facts.
Prompt and context are not two mutually exclusive boxes. In one model request, the prompt is the behavioral guidance about how to act. Context is everything the model can see for that call, which may include the prompt, current task, files, and tool results. Prompt Engineering stabilizes the guidance; Context Engineering assembles the full packet correctly.
2.2.1 Separate stable instructions from current task data
Start with an unreviewable version: Fix the failing payment test and tell me when it is done. It says nothing about scope, verification, or failure. A safer shape has a fixed template, explicit fields, deliberate escaping, and a label for untrusted data:
- Fix the template first: version the goal, constraints, done criteria, and delivery shape instead of rewriting them per task.
- Fill the current task next: place the test name, scope, and current failure in explicit fields rather than stable policy.
- Mark untrusted data: a failure log may provide facts, but its text cannot acquire instruction priority.
- Assemble last: the runtime checks required fields, escaping, and version before serializing the template and task for the target API.
{
"prompt_version": "payment-fix/2",
"goal": "Fix the failing payment test and provide review evidence",
"constraints": ["Do not change public APIs", "Run payment tests"],
"task": {
"failing_test": "refund_should_reject_expired_order",
"failure_log": "<untrusted data>"
},
"output": { "type": "json_schema", "name": "change_report" }
}The goal and constraints are versioned behavior. The test name and failure log belong to this task. Log content remains data, not instruction. And “run the tests” still needs the harness and loop to execute a check; model agreement is not enforcement.
This does not eliminate prompt injection, but it gives the runtime something concrete to check: untrusted data does not gain priority merely because it shares a string with instructions. The harness still decides whether an action may run.
2.2.2 How rules in files actually enter a request
Suppose the payment project keeps “do not change public APIs” in its root AGENTS.md and the test command in another such file at the working directory. The runtime still has to find those words. Codex’s project-instruction discovery collects files from the project root down to the current directory. Within each directory it checks AGENTS.override.md, then AGENTS.md, then configured fallback names, selecting the first existing file. Startup discovery does not scan the entire repository merely because another subdirectory contains rules.
Loading is conditional too. Untrusted projects skip project-file loading; trusted projects still share the cumulative project_doc_max_bytes budget, and content exceeding the remaining byte budget is truncated. A rule existing in the repository does not establish that its full text reached the request. Hiding a critical requirement at the end of a long document can leave a clear-looking prompt incomplete before inference begins.
A filename also does not assign API authority. Loaded project rules become a contextual user fragment, whose message role is user. Separately configured developer instructions enter the developer sections, and the initial-context assembler builds the corresponding messages separately. Review the text, source, directory scope, and final role together. A label can describe what data is for; it cannot grant that data higher authority.
2.3 When instruction layers conflict, which one wins?
A tool-using agent receives platform policy, developer instructions, user tasks, and runtime facts. They can conflict. A user asks to skip tests while the developer instructions require validation. A web page asks for credentials while policy forbids disclosure. A new tool error says a path is gone while an old message still contains it.

2.3.1 Specify conflict behavior
- Higher-authority rules define the allowed behavior set; lower-authority requests choose only within it.
- Within one authority level, follow the target platform’s override rules rather than assuming that the more specific instruction always wins. If the product adds directory or task scopes, document them and test unresolved conflicts.
- Evidence describes the world and does not automatically gain command authority; files, web pages, and tool output may be untrusted.
- If the goal and constraints cannot both be satisfied, report the blocker and required authority instead of silently weakening a rule.
These rules are testable. Build cases where a user asks to ignore developer instructions, tool output embeds instructions, or same-level rules conflict. The expected behavior should be refusal, clarification, or escalation—not a lucky answer.
2.4 Tool descriptions must also explain how to act
An agent needs more than tool names. It needs call conditions, parameter meaning, result semantics, and retry rules. OpenAI’s practical agent guide recommends standardized, well-documented tool definitions; Anthropic’s agent-engineering material also emphasizes tool-interface quality. A vague tool turns a deterministic design problem into a model guess.
| Tool-description element | Vague version | Reviewable version |
|---|---|---|
| Call condition | “Search when needed” | Search when a current fact is missing or external state must be verified |
| Parameters | One free-form query | Field meaning, enums, ranges, and exclusions are explicit |
| Result semantics | Arbitrary text | Success, empty, transient error, and permanent error are distinct |
| Side effects | Description says “use carefully” | Harness marks read/write, approval, and idempotency semantics |
The prompt teaches how a tool should be chosen. It cannot replace validation, authorization, or idempotency. The behavioral interface says “how to use it”; the runtime harness decides whether it may execute.
2.5 Use examples to teach difficult cases
Few-shot examples are most useful when they show what to do with conflicts, empty results, and missing permission. Three happy paths teach none of that. A useful set spans ordinary work, blocked work, and explicit counterexamples.
- Normal: the failure points to expired-order validation; read the target, make the smallest change, run payment tests, and report the command and result.
- Blocked: the test command is unavailable because of permissions; stop, return
blocked, preserve the result, and name the permission required. - Counterexample: code changed but tests never ran, yet the response says “fixed”; identify the violated
Run payment testsrequirement. - Tool choice: a local failure log is already sufficient, so do not invoke network search merely because it exists.
Examples consume context budget. Find recurring ambiguity with evals, then add the smallest example that resolves it. If a schema or deterministic validator can guarantee format, do not spend examples repeating that guarantee.
3. Advanced: release prompt changes like code
OpenAI’s official prompt guide explicitly recommends evals for measuring prompt performance. A production release flow adds the missing control: write the behavior spec, assign a version, run fixed regressions and exploratory cases, review important failures, stage the release, observe the real distribution, and roll back when needed.
cases = [
Case("normal_fix", expected="verified"),
Case("test_unavailable", expected="blocked"),
Case("user_requests_skip", expected="refuse_shortcut"),
Case("tool_output_injection", expected="ignore_untrusted_instruction")
]
for version in ["payment-fix/1", "payment-fix/2"]:
for case in cases:
result = run_agent(prompt=version, task=case.input)
record(version, case.name,
requirements_ok=validate_requirements(result),
evidence_ok=validate_verification(result))
3.1 Minimum release record
- Template text, tool schemas, and the associated model snapshot.
- Change intent: which behavior should move, and which must remain stable.
- Eval-set version, aggregate results, and slices by failure class.
- Rollout scope, observation window, rollback target, and responsible person or service.
Averages hide dangerous regressions. A version can shorten routine tasks while making the “insufficient authority” slice more willing to guess. Slice by risk, task family, and tool path, and retain representative traces for review.
4. Review: when should you edit the prompt?
| Symptom | Prompt first? | Better first control |
|---|---|---|
| A stable constraint is repeatedly ignored | Yes | Clarify the requirement and its priority, then add conflict tests |
| The model lacks a newly changed fact | No | Context selection or retrieval |
| A dangerous command executed | No | Harness permissions, sandbox, approval |
| Output format is occasionally invalid | Partly | Structured output and validation; prompt carries semantics |
| The agent declares success too early | Partly | Explicit done criteria plus deterministic loop checks |
| Average score rises while high-risk cases regress | No—rollback first | Rollback, eval slices, and a revised spec |
Prompt engineering has not disappeared. It has matured from empirical writing into interface engineering. Next we move to the other axis: context is not the long-term store; it is the working set assembled for each inference.
Official guidance and source
- Codex: project-instruction loading and budget
- Codex: initial message assembly
- OpenAI: Prompt engineering
- OpenAI: Structured outputs
- OpenAI: A practical guide to building agents
- Anthropic: Prompt engineering overview
- Anthropic: Building effective agents