As we mentioned before, starting from this large chapter, we will divide the content into two directions for explanation. One is LLM, and the other is CV.
This chapter serves as an introductory chapter for LLM, and will lay the groundwork before we explain KV-cache, alignment, MoE, etc.
Since modern large models like ChatGPT 6 Astra, Claude Mythos 5.1, Grok 4.6 have abandoned the Encoders and Encoder-Decoder architecture, we will skip these two architectures. The focus of this chapter will be on decoders.
Alright, let's begin.
After the explanation of the Attention mechanism in the previous chapter, I believe everyone has a certain grasp of the Attention mechanism. Therefore, we will dive right in.
First, we won't directly discuss the theory; let's talk about the background first. What exactly are Encoders and Decoders used for? They actually serve to link our inputs and outputs together, acting as a bridge.
Now let's talk theory:
First, let's discuss Encoders. This part will be relatively simple, and the content won't be much.
Last time we talked about self-Attention, which is essentially about finding out what a certain word means in the context.
The work of Encoders is like an assembly line, processing data in four steps.
1. First step: Embedding + Positioning
We will still use an example sentence: 'Apple Is Delicious, but Apple's stock price went down.'
The Encoder will break this sentence into individual tokens, like this:
'Apple'. 'Is' 'Delicious' 'but' and so on.
For these Tokens, you might think that the machine cannot distinguish which Apple refers to the fruit and which refers to the company, so the Positional Encoding we discussed last time comes into play. Encoders assign a Position number to each Token, thus clarifying the distinction. The first step is complete.
2. Second step: Self Attention and Multi-Head Attention.
We discussed these two in detail last time, so I won't repeat myself.
The machine gains a thorough and broad understanding of each word through Self Attention and Multi-head Attention.
At this point, the machine knows that the first Apple refers to the edible fruit (because the weight of Delicious is very high), and the second Apple refers to the company (the weight of Stock is high, giving this Apple a technological company connotation).
At this time, there are multiple experts (experts studying subject-verb and subject-object relationships), which is one Head, researching.
3. Third step: Residual Connection
In the Encoder, if all the multiple experts of Multi-head attention output their own understanding one by one, it would become chaotic, and the original meaning of the word would be lost.
We need to connect the final result with the original meaning to achieve the best understanding.
We prevent each expert from deleting or modifying previous results.
\($$\text{输出} = X + F(X)$$\)
Thus, each expert that takes over only looks for deficiencies in the previous results (just like the previous XG boost) to make supplements, rather than overturning them.
Step four: throw it into the FNN (Feed Forward Network)
The first step is dimensionality increase,
The FNN will increase something that was originally 768 dimensions by multiple times, for example, 4 times, so this vector becomes 3072 dimensions, a vector with richer information.
The second step is to introduce an activation function, such as RELU.
This activation function will make a nonlinear judgment, keeping certain information if it exceeds a certain standard, and vice versa.
Step three: dimensionality reduction.
After the activation function works, we obtain a brand new vector that returns to 768 dimensions, which is more comprehensive and meaningful.
Finally, we introduce normalization, and our Encoder's work is complete, coming to a conclusion.
Due to space reasons, we will focus on explaining how the Decoder outputs in the next chapter.
Thank you for watching, goodbye~