Ensiklopedia VibeKoding: An Introduction to Context Engineering.Ensiklopedia VibeKoding: An Introduction to Context Engineering.
> ๐ก Study Guide: Prompt engineering solves "how to say things clearly"; context engineering solves "how to let the model see the right information at the right time." This chapter revolves around one question: Within a limited context window, how do you make the model understand you without burning through your budget?> ๐ก Study Guide: Prompt engineering solves "how to say things clearly"; context engineering solves "how to let the model see the right information at the right time." This chapter revolves around one question: Within a limited context window, how do you make the model understand you without burning through your budget?
Before diving in, here are two "foundation bricks" to review:Before diving in, here are two "foundation bricks" to review:
------
Many people encounter similar situations when using large language models:Many people encounter similar situations when using large language models:
Intuitively, we tend to think: "This model has a bad memory."Intuitively, we tend to think: "This model has a bad memory."
But most of the time, the problem isn't that the model "can't remember" โ it's that we haven't designed the context it can see properly.But most of the time, the problem isn't that the model "can't remember" โ it's that we haven't designed the context it can see properly.
Faced with these challenges, relying solely on "writing better prompts" falls short. We need a more systematic engineering approach to ensure the model always receives the most critical information โ within a limited window and budget. This is exactly what context engineering aims to solve.Faced with these challenges, relying solely on "writing better prompts" falls short. We need a more systematic engineering approach to ensure the model always receives the most critical information โ within a limited window and budget. This is exactly what context engineering aims to solve.
------
Let's start with a concise working definition, then look at a few typical scenarios.Let's start with a concise working definition, then look at a few typical scenarios.
> Context engineering is an engineering discipline for building and managing the "information environment" for LLMs. It determines what the model "sees, ignores, and when it sees it," enabling stable task completion within a finite context window.> Context engineering is an engineering discipline for building and managing the "information environment" for LLMs. It determines what the model "sees, ignores, and when it sees it," enabling stable task completion within a finite context window.
Simply put, it boils down to three things: organizing information, controlling the window, and managing costs.Simply put, it boils down to three things: organizing information, controlling the window, and managing costs.
Common use cases include:Common use cases include:
Next, we'll start with the "hard-won lessons" of a real team to see how they gradually evolved from "only knowing how to write prompts" to "mastering context engineering."Next, we'll start with the "hard-won lessons" of a real team to see how they gradually evolved from "only knowing how to write prompts" to "mastering context engineering."
------
The case studies in this chapter come from Manus (a general-purpose AI Agent).The case studies in this chapter come from Manus (a general-purpose AI Agent).
Unlike ordinary conversations, Manus needs to autonomously plan and invoke tools to complete long tasks (involving dozens or even hundreds of interaction rounds).Unlike ordinary conversations, Manus needs to autonomously plan and invoke tools to complete long tasks (involving dozens or even hundreds of interaction rounds).
This creates a core contradiction:This creates a core contradiction:
After multiple architectural refactors, the Manus team realized one truth: Context can't just be "written" โ it must be "designed."After multiple architectural refactors, the Manus team realized one truth: Context can't just be "written" โ it must be "designed."
Manus co-founder Ji Yichao shared their "pitfall history":Manus co-founder Ji Yichao shared their "pitfall history":
| Iteration | Problem Encountered | Thinking at the Time | Result |
|---|---|---|---|
| 1st | AI kept forgetting mid-conversation | "Just write more prompts" | Longer prompts, higher costs |
| 2nd | Important info kept getting squeezed out | "Copy the important stuff multiple times" | Even longer text, even higher costs |
| 3rd | The bill was shockingly high | "Can we reuse previous computations?" | Found ways to reduce repeated computation costs |
| 4th | Long documents couldn't be processed | "Can we look things up only when needed?" | Built a "library + on-demand retrieval" solution |
Core insight: It's not about remembering more, but remembering more cleverly.Core insight: It's not about remembering more, but remembering more cleverly.
Traditional computer memory = Hard drive:Traditional computer memory = Hard drive:
AI context = Small whiteboard:AI context = Small whiteboard:
Manus's lesson: Use the whiteboard sparingly and cleverly โ don't use it to store an encyclopedia.Manus's lesson: Use the whiteboard sparingly and cleverly โ don't use it to store an encyclopedia.
------
Let's look at how your money is spent during a typical AI conversation:Let's look at how your money is spent during a typical AI conversation:
CODE ๐ฐ Cost Breakdown (per conversation): โโ 70% re-reading old content ("What did we just talk about?") โโ 20% processing new content ("What are you saying now?") โโ 10% generating a reply ("How should I respond?")
Surprising finding: 70% of the money is spent having the AI re-read what you've already said!Surprising finding: 70% of the money is spent having the AI re-read what you've already said!
Before discussing pricing, we need to understand a core technical concept: KV Cache (Key-Value Cache).Before discussing pricing, we need to understand a core technical concept: KV Cache (Key-Value Cache).
Don't be intimidated by the jargon โ it's essentially the AI's "short-term memory cheat sheet."Don't be intimidated by the jargon โ it's essentially the AI's "short-term memory cheat sheet."
It's like:It's like:
> You walk into an exam room.> You walk into an exam room.
> Scenario A: Every time, you have to read the entire textbook from cover to cover before starting to answer questions. (Slow, exhausting, expensive)> Scenario A: Every time, you have to read the entire textbook from cover to cover before starting to answer questions. (Slow, exhausting, expensive)
> Scenario B: You've already memorized the textbook content (Cache), so you sit down and start answering immediately. (Fast, effortless, cheap)> Scenario B: You've already memorized the textbook content (Cache), so you sit down and start answering immediately. (Fast, effortless, cheap)
In cloud providers' billing tables, "memorized material" (Cache Hit) is typically over 90% cheaper than "new material" (Cache Miss).In cloud providers' billing tables, "memorized material" (Cache Hit) is typically over 90% cheaper than "new material" (Cache Miss).
Using Claude as an example:Using Claude as an example:
Manus's practice: By having the AI "memorize the textbook," they reduced costs from $0.15 to $0.02 โ saving 87%!Manus's practice: By having the AI "memorize the textbook," they reduced costs from $0.15 to $0.02 โ saving 87%!
Many developers habitually put the "current time" as the first line of the System Prompt, thinking it's rigorous.Many developers habitually put the "current time" as the first line of the System Prompt, thinking it's rigorous.
But this is one of the biggest anti-patterns in context engineering.But this is one of the biggest anti-patterns in context engineering.
Imagine this: you've memorized an entire history book (System Prompt), only to have the first line read "the current second."Imagine this: you've memorized an entire history book (System Prompt), only to have the first line read "the current second."
If that line changes every second, then everything you memorized a second ago becomes invalid โ you have to re-memorize it from scratch.If that line changes every second, then everything you memorized a second ago becomes invalid โ you have to re-memorize it from scratch.
This is the Achilles' heel of prefix reuse (KV Cache): If the beginning changes, everything after it must be recomputed.This is the Achilles' heel of prefix reuse (KV Cache): If the beginning changes, everything after it must be recomputed.
text System: The current time is 2024-01-01 12:00:01. You are an assistant... (One minute later) System: The current time is 2024-01-01 12:01:01. You are an assistant...
Consequence: Although only a few characters changed, because they're at the very beginning, 99% of the subsequent fixed content can't benefit from cache reuse. Every request is as slow and expensive as the first one.Consequence: Although only a few characters changed, because they're at the very beginning, 99% of the subsequent fixed content can't benefit from cache reuse. Every request is as slow and expensive as the first one.
text System: You are an assistant... (thousands of characters of fixed rules and knowledge base go here) User: (Pass in the current time via a tool call or user message)
Benefit: The thousands of characters of rules never change; the AI only needs to "memorize" them once. Subsequent requests directly use cached memory, making them blazing fast.Benefit: The thousands of characters of rules never change; the AI only needs to "memorize" them once. Subsequent requests directly use cached memory, making them blazing fast.
๐ Try it yourself:๐ Try it yourself:
Toggle the switch below to enable "memorization acceleration," then click "Send New Request" multiple times.Toggle the switch below to enable "memorization acceleration," then click "Send New Request" multiple times.
Observe: when the first block becomes "memorized," what happens to the Time to First Token (TTFT)?Observe: when the first block becomes "memorized," what happens to the Time to First Token (TTFT)?
------
As conversations grow longer, the first problem you hit is: What happens when the window is full?As conversations grow longer, the first problem you hit is: What happens when the window is full?
The simplest memory management is the Sliding Window: new stuff comes in, old stuff goes out.The simplest memory management is the Sliding Window: new stuff comes in, old stuff goes out.
It sounds fair, but in real tasks, it's a disaster.It sounds fair, but in real tasks, it's a disaster.
Scenario replay:Scenario replay:
text Conversation log: [1] User: I'm Zhang San, responsible for the payments system [2] User: The project uses Go [3] User: The database is PostgreSQL ... [20] User: Write an API endpoint for me
Result: By the time you reach message 20, message 1 ("I'm Zhang San") has already been pushed out of the window. The AI has completely forgotten who you are or what system you're responsible for.Result: By the time you reach message 20, message 1 ("I'm Zhang San") has already been pushed out of the window. The AI has completely forgotten who you are or what system you're responsible for.
The core problem: This strategy treats critical information (identity, tech stack) and filler ("ok", "got it") as equals โ kicking them all out together.The core problem: This strategy treats critical information (identity, tech stack) and filler ("ok", "got it") as equals โ kicking them all out together.
Besides "forgetting fast," AI has another quirk: it also "overlooks" things.Besides "forgetting fast," AI has another quirk: it also "overlooks" things.
Research shows: AI is most sensitive to the beginning and end, and most likely to ignore the middle. This is the famous Lost in the Middle phenomenon.Research shows: AI is most sensitive to the beginning and end, and most likely to ignore the middle. This is the famous Lost in the Middle phenomenon.
U-shaped memory curve:U-shaped memory curve:
text Position: Start โ Middle โ End Memory: High โ Low โ High
๐ Try it yourself:๐ Try it yourself:
Solution: Place critical information at the beginning (system prompt) or the end (user query).Solution: Place critical information at the beginning (system prompt) or the end (user query).
------
If "first in, first out" doesn't work, what should we do?If "first in, first out" doesn't work, what should we do?
Manus's answer: Establish an "information hierarchy."Manus's answer: Establish an "information hierarchy."
Stop treating every piece of information equally โ decide what stays and what goes based on importance:Stop treating every piece of information equally โ decide what stays and what goes based on importance:
| Tier | Information Type | Treatment | Cost Impact |
|---|---|---|---|
| VIP | System setup, user identity | Keep forever | +15% cost |
| Important | Current task objective | Keep for task duration | +10% cost |
| Normal | Regular conversation history | Keep last 5 turns | Baseline cost |
| Disposable | Retrievable knowledge | Look up when needed | -60% cost |
Core idea: Spend 25% more to retain 90% of critical information.Core idea: Spend 25% more to retain 90% of critical information.
Think of the context window as a whiteboard:Think of the context window as a whiteboard:
๐ Try it yourself:๐ Try it yourself:
In the demo below, try "pinning" an important conversation message.In the demo below, try "pinning" an important conversation message.
Observe: as you keep chatting, does the pinned information stay in place while unpinned messages get pushed out?Observe: as you keep chatting, does the pinned information stay in place while unpinned messages get pushed out?
------
Sometimes there's simply too much information to process (like hundreds of pages of technical documentation), and the whiteboard can't hold it all. That's when you need an external brain โ RAG (Retrieval-Augmented Generation).Sometimes there's simply too much information to process (like hundreds of pages of technical documentation), and the whiteboard can't hold it all. That's when you need an external brain โ RAG (Retrieval-Augmented Generation).
When Manus faced millions of words of technical documentation, they compared two approaches:When Manus faced millions of words of technical documentation, they compared two approaches:
99% cost savings and 87% time savings!99% cost savings and 87% time savings!
Manus's experience distilled:Manus's experience distilled:
๐ Try it yourself:๐ Try it yourself:
Enter a question in the search box (e.g., "how to reset password") and see how the system pulls only the most relevant snippets from a mountain of documents.Enter a question in the search box (e.g., "how to reset password") and see how the system pulls only the most relevant snippets from a mountain of documents.
------
What if all the information is important, nothing can be deleted, and you don't want to look things up?What if all the information is important, nothing can be deleted, and you don't want to look things up?
Then you have no choice but to write smaller โ this is context compression.Then you have no choice but to write smaller โ this is context compression.
| Compression Method | Compression Ratio | What's Preserved | Use Case | Cost Savings |
|---|---|---|---|---|
| Summarization | 70% | Main ideas | Quick overview | Save 30% |
| Key Points | 50% | Critical points | Structured output | Save 50% |
| Tabular | 30% | Core data | Programmatic processing | Save 70% |
๐ Try it yourself:๐ Try it yourself:
Choose different compression strategies and see how walls of text become shorter and more refined.Choose different compression strategies and see how walls of text become shorter and more refined.
------
So far, we've learned various independent strategies like building blocks:So far, we've learned various independent strategies like building blocks:
Now it's time to assemble these blocks into a complete castle โ what we call Manus's "Memory Palace."Now it's time to assemble these blocks into a complete castle โ what we call Manus's "Memory Palace."
Don't think of context as a messy pile of text; think of it as a layered building. Each floor has its unique function and "rules of residence."Don't think of context as a messy pile of text; think of it as a layered building. Each floor has its unique function and "rules of residence."
๐ Try it yourself:๐ Try it yourself:
Click "Start Building" and see how we construct this palace floor by floor.Click "Start Building" and see how we construct this palace floor by floor.
The design philosophy of this palace aims to resolve three contradictions:The design philosophy of this palace aims to resolve three contradictions:
After the Manus team deployed this architecture, the results were immediate:After the Manus team deployed this architecture, the results were immediate:
------
To help you intuitively understand how this mechanism works, we've prepared a full-chain simulation.To help you intuitively understand how this mechanism works, we've prepared a full-chain simulation.
Choose a scenario and click "Next" to see how the Memory Palace dynamically retrieves, assembles, and cleans up context in the few seconds between a user query and the AI's response.Choose a scenario and click "Next" to see how the Memory Palace dynamically retrieves, assembles, and cleans up context in the few seconds between a user query and the AI's response.
If you're designing a system like Manus, don't just focus on how to write prompts โ pay more attention to how the system architecture orchestrates context.If you're designing a system like Manus, don't just focus on how to write prompts โ pay more attention to how the system architecture orchestrates context.
Below are system design blueprints for two classic scenarios, including both prompt design and code logic (pseudocode).Below are system design blueprints for two classic scenarios, including both prompt design and code logic (pseudocode).
> Core challenge: Long task cycles make it easy to forget the original requirements and project context.> Core challenge: Long task cycles make it easy to forget the original requirements and project context.
> Strategy: System layer (identity) + Task layer (pinned objectives) + Chat layer (sliding window).> Strategy: System layer (identity) + Task layer (pinned objectives) + Chat layer (sliding window).
1. System Prompt (Layer 1 & 2)1. System Prompt (Layer 1 & 2)
markdown # Layer 1: Identity Setup (System Prompt) - Never changes, leverages KV Cache You are a senior full-stack engineer, proficient in Python and Vue3. Coding style: - Variable naming strictly follows PEP8 - Critical logic must include comments - Prioritize existing project utility functions # Layer 2: Task Pinning (Task Context) - Must not be deleted during the task Current task: Refactor payment module (payment_module) Core constraints: 1. Must maintain compatibility with legacy API v1.0 2. Database migration scripts must be idempotent 3. Deadline: This Friday
2. Context Assembly Logic (Pseudocode)2. Context Assembly Logic (Pseudocode)
python def build_engineer_context(user_input, chat_history, task_info): context = [] # 1. Foundation layer: Identity setup (leverages KV Cache) # This part doesn't change for hundreds of rounds; computation cost is near $0 context.append(SYSTEM_PROMPT) # 2. Pillar layer: Task pinning (Pinned) # No matter how long the conversation, this is always inserted after System context.append(f"Current task: {task_info}") # 3. Retrieval layer: Code snippets (RAG) # Search the codebase for relevant code based on the user's question relevant_code = search_codebase(user_input) if relevant_code: context.append(f"Reference code:\n{relevant_code}") # 4. Interaction layer: Conversation history (Sliding Window) # Only take the last 10 turns to avoid blowing up the context recent_chat = chat_history[-10:] context.extend(recent_chat) # 5. Latest input context.append(user_input) return context
> Core challenge: Cost-sensitive, and absolutely cannot fabricate answers.> Core challenge: Cost-sensitive, and absolutely cannot fabricate answers.
> Strategy: System layer (strong constraints) + RAG layer (dynamic injection).> Strategy: System layer (strong constraints) + RAG layer (dynamic injection).
1. System Prompt (Layer 1)1. System Prompt (Layer 1)
markdown # Layer 1: Identity Setup (System Prompt) You are a professional e-commerce customer service representative. Response principles: 1. A warm, professional, and concise tone 2. **Absolutely prohibited** from fabricating facts; only answer based on [Reference Materials] 3. If the materials don't contain an answer, respond directly with "I'm very sorry, I'll need to transfer you to a human agent for this question"
2. Context Assembly Logic (Pseudocode)2. Context Assembly Logic (Pseudocode)
python def build_support_context(user_input): context = [] # 1. Foundation layer: Identity setup context.append(SYSTEM_PROMPT) # 2. Library layer: Dynamic retrieval (RAG) # Only in customer service scenarios does RAG take center stage, placed in the middle docs = vector_db.search(user_input, top_k=3) context.append("ใReference Materials Startใ") for doc in docs: context.append(doc.content) context.append("ใReference Materials Endใ") # 3. Interaction layer: Very short history # Customer service typically doesn't need long memory; keep the last 3 turns context.extend(get_recent_chat(limit=3)) context.append(user_input) return context
------
| English Term | Chinese Translation | Explanation |
|---|---|---|
| Context Window | ไธไธๆ็ชๅฃ | The maximum text length a model can process in a single pass (including input and output). Content exceeding this limit is truncated or forgotten. |
| Token | ่ฏๅ | The smallest unit of text processed by an LLM. Roughly 1 token โ 0.75 English words or 0.5 Chinese characters. Both billing and window limits use tokens as the unit. |
| KV Cache | KV ็ผๅญ | An inference acceleration technique that caches already-computed attention key-value pairs, avoiding redundant computation on repeated prefixes and significantly reducing latency and cost. |
| RAG | ๆฃ็ดขๅขๅผบ็ๆ | Before answering a question, relevant information is retrieved from an external knowledge base and provided as context to the model, reducing hallucination and expanding knowledge boundaries. |
| Sliding Window | ๆปๅจ็ชๅฃ | The most basic context management strategy. Keeps the token count within the window constant; when new content enters, the earliest old content is automatically removed. |
| Lost in Middle | ไธญ้ด่ฟทๅคฑ | A limitation of large models. Research shows models remember information at the beginning and end of long contexts best, while easily overlooking information in the middle. |
| System Prompt | ็ณป็ปๆ็คบ | Instructions placed at the very beginning of a conversation, used to set the model's identity, behavioral norms, response style, and core tasks. |
| Few-shot | ๅฐๆ ทๆฌๅญฆไน | Providing several "question-answer" examples in the prompt to help the model quickly understand the task pattern and output format. |
| Chain of Thought | ๆ็ปด้พ | Guiding the model to output reasoning steps before giving a final answer. This approach significantly improves the model's ability to solve complex logical and mathematical problems. |
| Hallucination | ๅนป่ง | The phenomenon where a model confidently generates information that appears plausible but is actually incorrect or non-existent. |
| Embedding | ๅ้ๅ | The technique of converting text into high-dimensional numerical vectors. Semantically similar texts are closer in vector space, forming the foundation of semantic search. |
| Vector DB | ๅ้ๆฐๆฎๅบ | A database specifically designed for storing and retrieving vector data. Supports similarity search to quickly find the document fragments that best match a query. |
| Temperature | ๆธฉๅบฆ | A hyperparameter controlling the randomness of model output. Higher values (e.g., 0.8) produce more diverse, creative outputs; lower values (e.g., 0.2) produce more deterministic, precise outputs. |
| TTFT | ้ฆๅญๅปถ่ฟ | Time to First Token โ the time elapsed from when a user sends a request to when the model outputs its first token. A key metric for measuring interactive experience. |
------
Manus's four refactors taught us:Manus's four refactors taught us:
From a practical standpoint: It's not about remembering more, but about remembering with more structure and selectivity.From a practical standpoint: It's not about remembering more, but about remembering with more structure and selectivity.
From a cost perspective:From a cost perspective:
The goal: within a given model and context limit, ensure every token invested serves a clear purpose.The goal: within a given model and context limit, ensure every token invested serves a clear purpose.