Ensiklopedia VibeKoding: Principles of Large Language Model Operation.Ensiklopedia VibeKoding: Principles of Large Language Model Operation.
> π‘ Learning Guide: This chapter requires no programming background. Through interactive demonstrations, it takes you deep into the inner workings of Large Language Models (LLMs). We'll start from the most basic concept of tokenization and go all the way to how GPT is trained and performs inference.> π‘ Learning Guide: This chapter requires no programming background. Through interactive demonstrations, it takes you deep into the inner workings of Large Language Models (LLMs). We'll start from the most basic concept of tokenization and go all the way to how GPT is trained and performs inference.
Humans communicate through language; computers compute through numbers.Humans communicate through language; computers compute through numbers.
The essence of a Large Language Model (LLM) is a bridge connecting these two worlds.The essence of a Large Language Model (LLM) is a bridge connecting these two worlds.
Its core task is singular: transform the problem of "understanding language" into a problem of "mathematical computation."Its core task is singular: transform the problem of "understanding language" into a problem of "mathematical computation."
To achieve this goal, we need to solve three core challenges:To achieve this goal, we need to solve three core challenges:
This tutorial will take you from zero, step by step, through the construction of this bridge.This tutorial will take you from zero, step by step, through the construction of this bridge.
------
A computer cannot read the word "hamburger" β it only understands numbers.A computer cannot read the word "hamburger" β it only understands numbers.
So our first task is: break text into the smallest units a computer can process.So our first task is: break text into the smallest units a computer can process.
Tokenization is the process of splitting a sentence into individual "word units" (Tokens).Tokenization is the process of splitting a sentence into individual "word units" (Tokens).
I love AI).English: Words are naturally separated by spaces, making tokenization easy (e.g., I love AI).ζη±δΊΊε·₯ζΊθ½).Chinese: There are no spaces, so algorithms are needed to segment the text (e.g., ζη±δΊΊε·₯ζΊθ½).The program that performs tokenization is called a Tokenizer.The program that performs tokenization is called a Tokenizer.
It acts like an interpreter, responsible for translating human text into sequences of numbers that machines can read.It acts like an interpreter, responsible for translating human text into sequences of numbers that machines can read.
Modern LLMs (such as GPT-4) typically use Subword Tokenization techniques (like the BPE algorithm).Modern LLMs (such as GPT-4) typically use Subword Tokenization techniques (like the BPE algorithm).
The clever part: common words stay intact, rare words get split.The clever part: common words stay intact, rare words get split.
Here is a real BPE tokenization example (based on the GPT-4 Tokenizer):Here is a real BPE tokenization example (based on the GPT-4 Tokenizer):
Input: "The quick brown fox jumps over the lazy dog. \nδ»ε€©ε€©ζ°ηδΈιοΌ"Input: "The quick brown fox jumps over the lazy dog. \nδ»ε€©ε€©ζ°ηδΈιοΌ"
Token List:Token List:
text index=791, string='The' index=4062, string=' quick' index=14198, string=' brown' index=39935, string=' fox' index=83368, string=' jumps' <-- if split, might be ' jump' + 's' index=927, string=' over' index=279, string=' the' index=16053, string=' lazy' index=3290, string=' dog' index=13, string='.' index=198, string='\n' <-- newline character index=33838, string='δ»ε€©' <-- common word merged directly index=54580, string='倩ζ°' index=20265, string='η' index=57672, string='δΈι' index=171, string='οΌ'
> On handling rare characters:> On handling rare characters:
> If a character not present in the vocabulary is encountered (suppose "δ»" were very rare), the model falls back to Byte-level encoding.> If a character not present in the vocabulary is encountered (suppose "δ»" were very rare), the model falls back to Byte-level encoding.
> 1. Raw Input: δ»> 1. Raw Input: δ»
> 2. Bytes: \xE4 \xBB \x8A> 2. Bytes: \xE4 \xBB \x8A
> 3. BPE Lookup: first look for \xE4\xBB\x8A -> not found -> split into \xE4\xBB (ID=1001) + \x8A (ID=2002).> 3. BPE Lookup: first look for \xE4\xBB\x8A -> not found -> split into \xE4\xBB (ID=1001) + \x8A (ID=2002).
> 4. Final Tokens: [1001, 2002].> 4. Final Tokens: [1001, 2002].
>>
> This mechanism guarantees that no matter what characters are input, the model can process them β there is never an OOV (Out Of Vocabulary) problem.> This mechanism guarantees that no matter what characters are input, the model can process them β there is never an OOV (Out Of Vocabulary) problem.
Key point: LLMs process not words, but Token IDs (a sequence of numeric indices).Key point: LLMs process not words, but Token IDs (a sequence of numeric indices).
------
Our task is to process language. But computers only understand numbers.Our task is to process language. But computers only understand numbers.
The most direct idea: assign each word an ID number.The most direct idea: assign each word an ID number.
If we only use IDs, the computer would think "10" and "20" are just two unrelated numbers.If we only use IDs, the computer would think "10" and "20" are just two unrelated numbers.
Moreover, if the vocabulary has 100,000 words, we might need an array of length 100,000 to represent a single word (One-Hot encoding), where 99,999 positions are 0 and only one position is 1.Moreover, if the vocabulary has 100,000 words, we might need an array of length 100,000 to represent a single word (One-Hot encoding), where 99,999 positions are 0 and only one position is 1.
To represent a word efficiently and meaningfully, we invented Embedding.To represent a word efficiently and meaningfully, we invented Embedding.
Instead of using a long 0/1 array, it uses a shorter array filled with decimals (e.g., 512 numbers) to describe a word.Instead of using a long 0/1 array, it uses a shorter array filled with decimals (e.g., 512 numbers) to describe a word.
[0.8 (is fruit), 0.1 (red), 0.9 (sweet)...]For example: [0.8 (is fruit), 0.1 (red), 0.9 (sweet)...]This way, we not only compress the data but also turn word meanings into computable "coordinates."This way, we not only compress the data but also turn word meanings into computable "coordinates."
------
Having solved the problem of representing "a single word," we now need to solve the problem of representing "a sentence."Having solved the problem of representing "a single word," we now need to solve the problem of representing "a sentence."
Because a sentence contains many words.Because a sentence contains many words.
This is a matrix.This is a matrix.
The reason we assemble them into matrices is that the core hardware of modern computers β GPUs (Graphics Cards) β are inherently designed for matrix operations.The reason we assemble them into matrices is that the core hardware of modern computers β GPUs (Graphics Cards) β are inherently designed for matrix operations.
Only by turning language into matrices can we leverage the parallel processing power of GPUs to achieve efficient inference and training.Only by turning language into matrices can we leverage the parallel processing power of GPUs to achieve efficient inference and training.
Let's review how data flows:Let's review how data flows:
------
Before diving into specific architectures, let's understand the term "model" in plain terms.Before diving into specific architectures, let's understand the term "model" in plain terms.
In the AI field, a Model is essentially a super complex function or black box.In the AI field, a Model is essentially a super complex function or black box.
An analogy:An analogy:
You can think of a model as an experienced master chef:You can think of a model as an experienced master chef:
What we call Training is the process of having this chef start as an apprentice, making mistakes billions of times. If it's too salty, tweak the "salt knob"; if too bland, tweak the "heat knob" β until they can consistently produce delicious dishes.What we call Training is the process of having this chef start as an apprentice, making mistakes billions of times. If it's too salty, tweak the "salt knob"; if too bland, tweak the "heat knob" β until they can consistently produce delicious dishes.
Today's LLMs are "super chefs" who have "read every book in human history" β except what they cook isn't food, but text.Today's LLMs are "super chefs" who have "read every book in human history" β except what they cook isn't food, but text.
With data (Tokens) and a chef (Model) in hand, let's look at how this chef thinks.With data (Tokens) and a chef (Model) in hand, let's look at how this chef thinks.
In AI's evolutionary history, there have been two main "ways of thinking" (architectures): RNN and Transformer.In AI's evolutionary history, there have been two main "ways of thinking" (architectures): RNN and Transformer.
Early models (RNNs, Recurrent Neural Networks) processed a sentence like playing the Telephone Game.Early models (RNNs, Recurrent Neural Networks) processed a sentence like playing the Telephone Game.
How it works:How it works:
This brings two fatal drawbacks:This brings two fatal drawbacks:
In 2017, Google proposed a completely new architecture β the Transformer. It fundamentally changed the rules, turning the "Telephone Game" into a Roundtable Meeting.In 2017, Google proposed a completely new architecture β the Transformer. It fundamentally changed the rules, turning the "Telephone Game" into a Roundtable Meeting.
How it works:How it works:
The Transformer no longer passes messages one by one; instead, it has all words sit down at the table at once.The Transformer no longer passes messages one by one; instead, it has all words sit down at the table at once.
This perfectly solves RNN's pain points:This perfectly solves RNN's pain points:
> In summary:> In summary:
>>
> - RNN: Like navigating a maze, step by step, easy to get lost.> - RNN: Like navigating a maze, step by step, easy to get lost.
> - Transformer: Like viewing a map from a bird's-eye view β both the start and finish are in sight.> - Transformer: Like viewing a map from a bird's-eye view β both the start and finish are in sight.
Because the Transformer processes everything at once, without special handling it can't tell the difference between "I love you" and "You love I" (same words, just in a different order).Because the Transformer processes everything at once, without special handling it can't tell the difference between "I love you" and "You love I" (same words, just in a different order).
So we give each word a number tag (Positional Encoding) to tell the model who is in position 1, who is in position 2, and so on.So we give each word a number tag (Positional Encoding) to tell the model who is in position 1, who is in position 2, and so on.
> A small reminder: many LLMs are autoregressive (predicting the next word), so during generation they still output one token at a time; but in the internal computation of each generation step, the Transformer still excels at leveraging matrix parallelism and caching optimizations.> A small reminder: many LLMs are autoregressive (predicting the next word), so during generation they still output one token at a time; but in the internal computation of each generation step, the Transformer still excels at leveraging matrix parallelism and caching optimizations.
You may have heard that when generating long texts, it gets slower toward the end, or VRAM usage grows. This is usually because the model needs to "remember" everything it has generated so far.You may have heard that when generating long texts, it gets slower toward the end, or VRAM usage grows. This is usually because the model needs to "remember" everything it has generated so far.
How does the Transformer "take notes"?How does the Transformer "take notes"?
In the Transformer's attention mechanism, every word generates two vectors, Key (K) and Value (V), for later words to "query."In the Transformer's attention mechanism, every word generates two vectors, Key (K) and Value (V), for later words to "query."
The role of KV Cache:The role of KV Cache:
KV Cache is like an "incremental notebook."KV Cache is like an "incremental notebook."
This is why long-context conversations consume so much VRAM β it's not that the model is getting bigger, but that the notes (KV Cache) are getting too thick.This is why long-context conversations consume so much VRAM β it's not that the model is getting bigger, but that the notes (KV Cache) are getting too thick.
------
Many people mistakenly think ChatGPT truly understands what we're saying, but its fundamental instinct is only one thing: predicting the next token (Next Token Prediction).Many people mistakenly think ChatGPT truly understands what we're saying, but its fundamental instinct is only one thing: predicting the next token (Next Token Prediction).
If you give a base model the input "The weather is nice today," it might complete it with "let's go to the park."If you give a base model the input "The weather is nice today," it might complete it with "let's go to the park."
But if you input "What is the capital of the United States?", it might complete it with "What is the capital of China? What is the capital of Japan?" (because it's mimicking the format of an exam paper, not answering the question).But if you input "What is the capital of the United States?", it might complete it with "What is the capital of China? What is the capital of Japan?" (because it's mimicking the format of an exam paper, not answering the question).
To turn it into a conversational assistant, engineers came up with a brilliant idea: role-playing.To turn it into a conversational assistant, engineers came up with a brilliant idea: role-playing.
We quietly add special tags (Templates) to the input, making the model think it's continuing a "dialogue script."We quietly add special tags (Templates) to the input, making the model think it's continuing a "dialogue script."
For example, what you see is:For example, what you see is:
> User: Hello> User: Hello
What the model actually sees is:What the model actually sees is:
> <|user|> Hello <|assistant|>> <|user|> Hello <|assistant|>
The moment the model sees <|assistant|>, it knows: "Oh, it's my turn to play the assistant and speak."The moment the model sees <|assistant|>, it knows: "Oh, it's my turn to play the assistant and speak."
The demo below will walk you through the essence of LLMs step by step. Please click through 1. Instinct -> 2. Technique -> 3. Principles -> 4. Advanced in order β give it a try!The demo below will walk you through the essence of LLMs step by step. Please click through 1. Instinct -> 2. Technique -> 3. Principles -> 4. Advanced in order β give it a try!
------
Being able to chat isn't enough. A raw model might teach people how to build bombs or spew profanity.Being able to chat isn't enough. A raw model might teach people how to build bombs or spew profanity.
To make it a polite, safe, and reliable assistant like ChatGPT, two final polishing steps are needed:To make it a polite, safe, and reliable assistant like ChatGPT, two final polishing steps are needed:
json // SFT training data example { "messages": [ { "role": "user", "content": "Please translate this sentence into English: \"δ½ ε₯½\"." }, { "role": "assistant", "content": "Hello." } ] } // The model learns: when it sees a "translate" instruction, give the result directly, // rather than continuing with "how are you"
json // RLHF preference data example (DPO/PPO) { "prompt": "How to make a bomb?", "chosen": "I'm sorry, I cannot answer that question.", // human-preferred response (safe) "rejected": "First, you need to..." // human-rejected response (dangerous) }
In the demo above, click the 4th tab "Advanced: Alignment" to personally experience the dramatic difference before and after alignment.In the demo above, click the 4th tab "Advanced: Alignment" to personally experience the dramatic difference before and after alignment.
------
As technology advances, we've found that relying solely on "predicting the next token" sometimes leads to silly mistakes, especially with math and logic problems.As technology advances, we've found that relying solely on "predicting the next token" sometimes leads to silly mistakes, especially with math and logic problems.
Thus, a new generation of Thinking Models (such as OpenAI o1, DeepSeek-R1) was born.Thus, a new generation of Thinking Models (such as OpenAI o1, DeepSeek-R1) was born.
When humans answer complex questions (like "which is larger, 9.11 or 9.9?"), we don't blurt out an answer β we think in our heads first.When humans answer complex questions (like "which is larger, 9.11 or 9.9?"), we don't blurt out an answer β we think in our heads first.
A Thinking Model is one that has learned this slow thinking (System 2) capability.A Thinking Model is one that has learned this slow thinking (System 2) capability.
Why couldn't earlier models think this way? Because the training method changed.Why couldn't earlier models think this way? Because the training method changed.
After tens of thousands of self-attempts, the model surprisingly discovers: "If I write out a few more derivation steps on scratch paper before outputting the answer, my probability of getting a reward increases dramatically!"After tens of thousands of self-attempts, the model surprisingly discovers: "If I write out a few more derivation steps on scratch paper before outputting the answer, my probability of getting a reward increases dramatically!"
Thus, this "think first, then answer" behavioral pattern gets reinforced and solidified. It's like AlphaGo playing against itself and ultimately surpassing human game records.Thus, this "think first, then answer" behavioral pattern gets reinforced and solidified. It's like AlphaGo playing against itself and ultimately surpassing human game records.
When using Thinking Models (like DeepSeek-R1, OpenAI o1), your prompting strategy needs a complete overhaul.When using Thinking Models (like DeepSeek-R1, OpenAI o1), your prompting strategy needs a complete overhaul.
| Feature | Traditional Models (GPT-4o, Claude 3.5) | Thinking Models (R1, o1) |
|---|---|---|
| Core Logic | System 1 (Intuition) | System 2 (Logic) |
| Prompting Tips | Need to guide Chain of Thought (CoT) e.g., "Please think step by step..." | Don't overdo it The model has built-in CoT; manual guidance actually interferes |
| Instruction Clarity | Need to break complex tasks into subtasks | Give the final goal directly and let the model decompose it |
| Suitable Scenarios | Creative writing, simple translation, chitchat | Complex math, code refactoring, logical reasoning |
> β οΈ Note: With Thinking Models, the less intervention, the better. You only need to clearly define "what constitutes a perfect task result," not "how to do it."> β οΈ Note: With Thinking Models, the less intervention, the better. You only need to clearly define "what constitutes a perfect task result," not "how to do it."
In the future, we may no longer need to distinguish between "thinking models" and "regular models."In the future, we may no longer need to distinguish between "thinking models" and "regular models."
The ideal AI should, like humans, possess Adaptive Compute capability:The ideal AI should, like humans, possess Adaptive Compute capability:
As models grow larger (like GPT-4, DeepSeek-V3), if every single token generation required computing every neuron, the speed would become unbearably slow.As models grow larger (like GPT-4, DeepSeek-V3), if every single token generation required computing every neuron, the speed would become unbearably slow.
Thus, the MoE (Mixture of Experts) architecture emerged.Thus, the MoE (Mixture of Experts) architecture emerged.
The essence of MoE lies in native token-level routing. It is absolutely not about dividing work by "task type" (e.g., sending all math problems to the math expert); rather, it divides work in real time by "the current token being generated."The essence of MoE lies in native token-level routing. It is absolutely not about dividing work by "task type" (e.g., sending all math problems to the math expert); rather, it divides work in real time by "the current token being generated."
def", route to the code expert.When the model generates "def", route to the code expert.love", route to the literature expert.When the model generates "love", route to the literature expert.3.14", route to the math expert.When the model generates "3.14", route to the math expert.This means that even within the same sentence, different tokens are often handled by different experts.This means that even within the same sentence, different tokens are often handled by different experts.
Beyond MoE, there's another core pain point: context length.Beyond MoE, there's another core pain point: context length.
Traditional Transformers (like GPT-4) use standard attention mechanisms, whose computational cost explodes quadratically as token count increases.Traditional Transformers (like GPT-4) use standard attention mechanisms, whose computational cost explodes quadratically as token count increases.
To solve this, models like MiniMax (abab series) and RWKV adopted Linear Attention.To solve this, models like MiniMax (abab series) and RWKV adopted Linear Attention.
The fundamental difference is: do you choose to "keep all original words," or do you choose to "summarize as you go"?The fundamental difference is: do you choose to "keep all original words," or do you choose to "summarize as you go"?
| Architecture | Core Mechanism | Complexity (Length N) | Parallel Training | Inference Speed | Forgetting Problem | Representative Models |
|---|---|---|---|---|---|---|
| RNN | Sequential recurrence | $O(N)$ (Low) | β No | Slow (sequential) | Severe (long-distance forgetting) | LSTM, GRU |
| Transformer | Global attention | $O(N^2)$ (Very High) | β Yes | Medium (KV Cache) | None (but window-limited) | GPT-4, Llama |
| RWKV / Linear | Linear attention | $O(N)$ (Low) | β Yes | Fast (constant VRAM) | Mild (compression loss) | RWKV, MiniMax |
> RWKV / Linear Attention attempts to combine the strengths of both: parallel training like Transformer, efficient inference like RNN.> RWKV / Linear Attention attempts to combine the strengths of both: parallel training like Transformer, efficient inference like RNN.
------
You've now connected the dots from "tokenization" all the way to "ChatGPT":You've now connected the dots from "tokenization" all the way to "ChatGPT":
Next steps:Next steps:
transformers library.If you want hands-on practice, try loading a small model (like GPT-2) using Python's transformers library.------
| Term | Full Name | Explanation |
|---|---|---|
| LLM | Large Language Model | A large language model. An AI model trained on massive amounts of text, capable of understanding and generating human language. |
| Token | - | Tokenization. The smallest unit into which text is split (such as a word, character, or subword fragment). Models read and write Token IDs. |
| Embedding | - | Word vector. A numerical vector that maps a Token into a high-dimensional space (e.g., 4096 dimensions), capturing semantic relationships between words. |
| Transformer | - | The core architecture of modern LLMs. Based on the attention mechanism, capable of processing long texts in parallel. |
| Attention | Attention Mechanism | Attention mechanism. Allows the model to dynamically focus on other relevant words in context when processing a given word. |
| Context Window | - | Context window. The maximum number of Tokens a model can "remember" in a single inference pass (e.g., 128k). |
| Pre-training | - | Pre-training. Training a model on massive amounts of unlabeled text so it learns the fundamental patterns of language and world knowledge. |
| SFT | Supervised Fine-Tuning | Supervised fine-tuning. Using high-quality Q&A pair data to teach the model to follow human instructions. |
| RLHF | Reinforcement Learning from Human Feedback | Reinforcement learning from human feedback. Using human scoring to further adjust model behavior, aligning it with human values (Alignment). |
| CoT | Chain of Thought | Chain of thought. A technique that guides the model to generate reasoning steps before producing the final answer. |
| MoE | Mixture of Experts | Mixture of experts. A model composed of multiple "expert" sub-models, where the appropriate experts are selectively activated based on the problem β more efficient. |
| Temperature | - | Temperature. A parameter controlling the randomness of model generation. Higher temperature β more creative but less controllable; lower temperature β more deterministic. |