Prompt Engineering Is Too Small a Mental Model, Think Context Architecture
I was knee‑deep in a fintech app that tried to answer user questions about credit‑score changes. My first version was a classic “prompt engineering” hack:
“You are a friendly finance bot, explain the score drop in plain English.”
It worked for the first few queries, then the answers got weird – sometimes the bot repeated the same disclaimer, other times it started spilling the policy docs verbatim. I was losing token budget fast and the model kept “remembering” the old disclaimer even after I removed it. I realized I was treating the whole prompt as a monolith, when in reality I was juggling five different pieces of context that should have their own lanes.
Enter context architecture
Think of an LLM conversation like a pizza kitchen.
- System instructions – the chef’s recipe book; they stay on the board all night.
- User state – the current order ticket (what the customer just asked, plus any previous choices).
- Retrieved knowledge – the pantry shelf you pull ingredients from (a database lookup of the user’s credit history).
- Tool results – the side‑dish the sous‑chef prepares (e.g., a risk‑score API call).
- Policy – the health‑inspection sticker you slap on the box to make sure you never serve raw chicken.
The magic happens in precedence. The model first looks at system instructions, then layers user state, then fetched knowledge, then tool results, and finally policy overrides anything that would break compliance. If you drop a policy clause at the bottom of a huge prompt, it might never be seen because the model runs out of tokens before it gets there. That’s why you always put the “stable‑prefix” (system + policy) at the front – it’s cached, cheap, and guarantees the model never forgets the rules.
Token budget = pizza dough budget
You only have so many tokens per request; if you waste them on stale context (like a disclaimer you already sent), you run out of room for the juicy answer.
Contamination risk is when old user state leaks into a new conversation – imagine the kitchen forgetting to clear the ticket and adding yesterday’s pepperoni to today’s vegan pizza. To avoid that, we clear the user state each session or use a fresh context window.
Stable‑prefix caching
(OpenAI Prompt Caching) lets you store the system + policy block once and reuse it across millions of calls. In my edutech project, we cached a 120‑token prefix that described the “tutor bot” persona and compliance rules. The cache cut latency by 30 % and saved us $2 K a month on token spend.
Case study: Fintech fraud‑alert bot
| Piece | Content |
|---|---|
| System | “You are a compliance‑aware fraud analyst.” |
| Policy | “Never reveal raw transaction IDs.” |
| User | “My card was declined twice.” |
| Retrieved | “Last 5 transactions: $23, $0, $0.” |
| Tool | “Risk engine returns confidence 0.92.” |
The final prompt stayed under 350 tokens, the answer was spot‑on, and no policy leakage occurred.
So the next time you’re tempted to dump a giant prompt into the model, pause and ask yourself:
- What belongs in system instructions?
- What’s the user state?
- What do I need to fetch?
- What tool results am I adding?
- What policy must sit at the very front?
Build that context architecture, and you’ll spend less time chasing bugs and more time building cool features.
What weird context mishap have you run into? Drop a comment and let’s swap stories!
If you are someone who loves to know the technical work and architecture design I have shared more details based on my experience on this here:
https://github.com/SalmonJoy/My_guide_for_building_AI_systems/blob/main/Context_engineering_vs_prompt_engineering.md
Top comments (0)