myibrahim.cloud

Vibe Coding · Workflow & Prompts

Prompt engineering for engineers

Stop calling it 'engineering' until you treat it like one. Versioning, eval-driven iteration, and the prompt patterns that survive contact with production.

The phrase "prompt engineering" gets sneered at, mostly fairly. Most "prompt engineering" is incantation: people typing magic words and hoping the model behaves. That's not engineering. Engineering is hypothesis-test-iterate against a metric.

This is how to do the latter.

Treat prompts like code#

The single shift that fixes 80% of prompt problems: store prompts in version control, just like everything else.

# prompts/customer_support_v3.md  (in the repo, code-reviewed)
SYSTEM_PROMPT = open("prompts/customer_support_v3.md").read()

response = client.messages.create(
    model="claude-sonnet-4-6",
    system=SYSTEM_PROMPT,
    messages=[{"role": "user", "content": user_query}],
)

Why this matters:

  • You can git-blame prompt changes when behavior shifts.
  • Code review applies. Bad prompts get caught before shipping.
  • A/B testing becomes "deploy two versions, compare metrics."
  • Your prompts get reviewed by people, not invented in a chat panel.

If your team's prompts live as inline strings in 12 different functions, the rest of this article doesn't apply to you yet. Fix that first.

The structure that holds up#

A prompt that survives production has roughly this shape:

1. Role definition. Concise, specific.
   "You are a senior support engineer at AcmeCo helping with API issues."

2. Scope and constraints.
   "Only answer questions about the AcmeCo API. Refuse off-topic queries.
   If you don't know, say so. Never invent endpoint names."

3. Reference material (cached).
   <docs>{full API reference}</docs>

4. Format spec.
   "Respond in markdown. Cite each claim with the section number from the docs.
   If the user's question can't be answered from the docs, respond with
   exactly: 'I don't have that in the documentation.'"

5. Few-shot examples (optional but recommended).
   <example>...</example>
   <example>...</example>

6. The user's input goes here.

The order matters. Put stable content first — the role, scope, format, references — so prompt caching covers it. The user's query, the dynamic part, goes last.

Few-shot beats explanation#

The fastest way to teach a model a format: show, don't tell.

Bad:
"Answer in JSON with keys 'category' and 'urgency'."

Better:
"Examples of the format:
<example>
{ "category": "billing", "urgency": "low" }
</example>
<example>
{ "category": "outage", "urgency": "critical" }
</example>"

Three examples is usually enough. More than five rarely helps. Pick examples that span the corner cases — easy ones, edge cases, the format you want for ambiguous input.

XML tags > markdown headers#

For prompt structure, XML tags work better than markdown. The model's been trained on a lot of XML-tagged data and it parses these reliably:

<task>...</task>
<context>...</context>
<input>{user_input}</input>
<output_format>...</output_format>

Markdown headers (## Task, ## Context) work too but the model occasionally spills out of them. XML tags give clearer boundaries.

"Think step by step" still works, just smarter#

Chain-of-thought has been the prompt-engineering hammer for two years. Modern models often do it implicitly when needed. But for hard reasoning tasks, encourage explicit structure:

<task>Determine whether the user's request matches our refund policy.</task>
<input>{user_message}</input>
<policy>{refund_policy}</policy>

<instructions>
Before answering, work through these steps inside <thinking></thinking> tags:
1. What is the user actually asking for?
2. Which policy clauses apply?
3. Does the request meet the criteria?

Then provide your final answer in <answer></answer> tags.
</instructions>

The <thinking> block:

  • Lets you inspect the model's reasoning.
  • Gives the model space to work through the problem before committing.
  • Can be stripped from the user-facing output.

You pay for the thinking tokens, so use this when the task is hard. Don't reach for it on simple classification.

The "what to do when X" pattern#

Models tend to over-helpful. They'll try to answer questions that should be refused. They'll guess when they should ask. Spell it out:

If the user asks about anything outside billing, respond with:
"I can only help with billing questions. For other topics, contact support@acmeco.com."

If the user's question is ambiguous, ask one clarifying question
before answering. Don't guess.

If you don't have enough information from the documentation, respond:
"I don't have that information. Try the API reference at docs.acmeco.com."

These rules need to be tested in the eval set. Otherwise they drift.

Iteration, with measurement#

The bad workflow:

  1. Write prompt
  2. Test on a few demo queries
  3. Ship
  4. Customers complain
  5. Tweak prompt
  6. Hope

The good workflow:

  1. Write 50-question eval set, with criteria for "correct"
  2. Write prompt
  3. Run eval, get baseline %
  4. Iterate prompt — every change increases or decreases the metric
  5. Ship the version that beat baseline
  6. Sample production traffic, add new examples to eval set
  7. When the metric stagnates, try a different model

Without an eval set, prompt iteration is vibes-based. With one, it's debuggable engineering.

Eval accuracy across prompt iterations
  • Initial v158
  • Added few-shot72
  • XML structured78
  • Added refusal rules84

Cost-aware prompts#

Tokens cost money. Tokens cost latency. A few discipline points:

  • Cap output tokens. Set max_tokens to the smallest value that produces good answers. Most "give me a 3-sentence summary" responses are fine at max_tokens=200.
  • Avoid asking the model to repeat input. "Restate the user's question, then answer it" is wasted tokens. The model has the question; you don't need it back.
  • Push reference material into cached prefix. Anthropic's prompt caching is 90% off for stable content. Long instruction blocks belong in the cached prefix.
  • Drop polite filler. "Please respond as best you can" wastes tokens and doesn't change behavior.

What the model can't tell you#

A few things that prompts genuinely can't fix:

  • The model doesn't know your data. Tell it (RAG).
  • The model can't do math reliably. Have it call a tool.
  • The model can't access the internet at inference. Use a tool.
  • The model doesn't know today's date. Pass it in.
  • The model has a knowledge cutoff. Don't ask about events past it.

When you find yourself prompting harder and harder for something the model fundamentally can't do, the prompt isn't the problem. The architecture is.

Further reading#

  • ai
  • prompts
  • llms
  • workflow
Need this built? I build ai product or mvp projects for clients worldwide. Tell me about yours.