Skip to content
RWRui Wang / Ideas
← All notes

Deep Learning

Deep learning - PyTorch

Deep learning - PyTorch cover
This chapter was originally supposed to be published before Transformers, but due to some reasons, it was too difficult to write and got backlogged for a long time. By the time this article was finally completed, Transformers had already updated to version 3/4.

Learning PyTorch should take priority over Transformers because T is essentially a high-level encapsulation of P. Understanding P will help with T, and obviously, P is much simpler.


Without further ado, let's get straight to the point:


First, let's explain the uses of PyTorch:


PyTorch uses tensors, along with forward and backward propagation, greatly accelerating Transformers' computation speed and significantly reducing its space complexity, making it the foundation for many high-level encapsulations in the LLM field.


Personally, the aspect I use PyTorch for the most is computing gradients because its performance is truly outstanding.


We'll explain PyTorch in 5 stages:


First, let's cover some theory:


Tensors:It's just a niche name—essentially, they're Arrays. However, for PyTorch, high-dimensional Arrays are so predominant that they adopted the specialized name for high-dimensional Arrays: tensors.


I believe everyone has some understanding of arrays from learning NumPy Arrays and Pandas DataFrames. The difference with tensors is that tensors can be computed on GPUs, making them much faster.


Since this PyTorch theory chapter is simple and its application isn't difficult either, unlike Transformers where we only covered theory, let's actually apply it here:


import torch

# 1.1 直接从列表创建
x = torch.tensor([1.0, 2.0, 3.0])
print(x)  # tensor([1., 2., 3.])

# 1.2 创建特定形状的张量
zeros = torch.zeros(2, 3)          # 2行3列的全 0 张量
ones = torch.ones(2, 3)            # 2行3列的全 1 张量
rand_tensor = torch.randn(2, 3)    # 2行3列的标准正态分布随机数

print(rand_tensor)


As you can see, creating a tensor is actually quite simple.


Shape manipulation

Next, let's talk about tensorshape manipulation, which is also a step in Embedding:


x = torch.arange(12)  # 生成 0 到 11 的一维张量 [0, 1, ..., 11]

# 将 1D 张量重塑为 3行4列 的 2D 张量
x_2d = x.view(3, 4)    

#或者

x_2d = x.reshape(3, 4)

# 使用 -1 让 PyTorch 自动推导该维度的大小
x_auto = x.view(2, -1) # 自动算出另一维是 6(因为总数 12 / 2 = 6)

print("x_2d 形状:", x_2d.shape)
print("x_auto 形状:", x_auto.shape)


You can see that tensors are highly mutable and very flexible, which is why they're so popular.


We also mentioned that the higher the dimensionality of a token's content vector, the more information it can carry. That's why we prefer to increase the dimensionality of vectors and then scale them.


Tensor operations and broadcasting mechanism

Tensor addition, subtraction, multiplication, and division are performed element-wise by default. When two tensors have mismatched shapes, PyTorch will attempt to automatically expand dimensions (broadcasting).


a = torch.tensor([[1.0, 2.0], [3.0, 4.0]])
b = torch.tensor([[10.0, 20.0], [30.0, 40.0]])

# 按元素相加/相乘
print("按元素相加:\n", a + b)
print("按元素相乘:\n", a * b)

# 矩阵乘法(线性代数中的矩阵乘法,使用 @ 或 torch.matmul)
print("矩阵乘法:\n", a @ b)

Automatic differentiation (Autograd)

Earlier, we mentioned that gradients describe the direction of descent. To update parameters, PyTorch's Autograd helps us automatically compute gradients for every parameter in the loss function.


# 假设 x 是我们需要优化的模型参数(比如权重)
x = torch.tensor(2.0, requires_grad=True)

# 构建计算过程:y = x^2 + 3x + 1
y = x**2 + 3*x + 1

print("y 的值:", y.item())  # y = 2^2 + 3(2) + 1 = 11.0

Backward propagation (loss.backward())

When PyTorch computes a scalar value, such as loss, we initiate backward propagation to let torch automatically calculate gradients (requires_grad = True) and store them in .grad.


# 执行反向传播,计算 dy/dx
y.backward()

# 查看 x 的梯度
print("x 的导数 (dy/dx):", x.grad)  # 输出: tensor(7.)

Gradient accumulation and clearing (extremely important)

PyTorch's characteristic is that each time gradients are computed, they are automatically accumulated into .grad, so we need to manually clear them:


x = torch.tensor(2.0, requires_grad=True)

# 第一轮前向 + 反向
y1 = x**2
y1.backward()
print("第1轮计算后的梯度:", x.grad)  # 2 * 2 = 4

# 第二轮前向 + 反向(如果不手动清零)
y2 = x**2
y2.backward()
print("第2轮计算后的梯度(累加了):", x.grad)  # 4 + 4 = 8

# 正确做法:清除梯度
x.grad.zero_()
print("清零后的梯度:", x.grad)  # tensor(0.)


Disabling gradient tracking:

During the model evaluation phase, we no longer need to modify gradients, so gradient tracking is disabled.


Let's practice this: For


\(y = w^2\)


Finding the minimum of a function:


Below is the solution:


# 1. 随机初始化参数 w
w = torch.tensor(10.0, requires_grad=True)
learning_rate = 0.1

print("初始 w:", w.item())

# 2. 迭代 20 次更新 w
for step in range(20):
    # 前向传播:计算 Loss
    loss = w ** 2
    
    # 反向传播:计算 dLoss/dw
    loss.backward()
    
    # 更新参数:w = w - lr * grad
    # 注意:更新参数本身的操作不需要记录梯度,所以放在 no_grad 中
    with torch.no_grad():
        w -= learning_rate * w.grad
        
    # 必须把梯度清零,否则下一轮会累加
    w.grad.zero_()

print("优化 20 步后的 w:", w.item())  # 极其接近 0


Phase 2: Dataset/Dataloader data loading pipeline:


First, let's briefly introduce Dataset and Dataloader.


Dataset: The core is retrieving individual samples, solving the problem of 'where is the data and how to read it';


Dataloader: The core is how to efficiently feed data to the model; batching the data, shuffling, and accelerating multi-threaded loading onto the GPU;


PyTorch supports both map-style and iterable-style datasets, but map-style is used in most daily development scenarios.


Moving to application:


Custom map-style dataset: (three essential elements) The following three methods must be implemented;

import torch
from torch.utils.data import Dataset

class CustomDataset(Dataset):
    def __init__(self, data_list, labels):
        """
        1. 初始化:传入数据源、路径或配置参数,完成初始化设置。
        """
        self.data = data_list
        self.labels = labels

    def __len__(self):
        """
        2. 返回数据集总样本量:len(dataset) 时会被自动调用。
        """
        return len(self.data)

    def __getitem__(self, idx):
        """
        3. 按索引获取单个样本:dataset[idx] 时调用。
           在这里进行数据读取(如从磁盘读图片)和预处理变换(Transform)。
        """
        x = self.data[idx]
        y = self.labels[idx]
        
        # 返回格式通常为 (feature, label) 元组或字典
        return torch.tensor(x, dtype=torch.float32), torch.tensor(y, dtype=torch.long)


Reading images & data augmentation:

from PIL import Image
import torchvision.transforms as T

class ImageDataset(Dataset):
    def __init__(self, image_paths, labels, transform=None):
        self.image_paths = image_paths
        self.labels = labels
        # 预处理变换管线
        self.transform = transform or T.Compose([
            T.Resize((128, 128)),            # 调整图像大小
            T.ToTensor(),                    # 转为 Tensor 并且像素值归一化到 [0, 1]
            T.Normalize(mean=[0.485, 0.456, 0.406], 
                        std=[0.229, 0.224, 0.225]) # 标准化
        ])

    def __len__(self):
        return len(self.image_paths)

    def __getitem__(self, idx):
        # 延迟读取:只在请求该索引时才从磁盘加载图片(节省内存)
        img_path = self.image_paths[idx]
        image = Image.open(img_path).convert("RGB")
        
        if self.transform:
            image = self.transform(image)
            
        label = torch.tensor(self.labels[idx], dtype=torch.long)
        return image, label


Using dataloader to encapsulate the dataset:


from torch.utils.data import DataLoader

dataloader = DataLoader(
    dataset=dataset,          # 实例化的 Dataset 对象
    batch_size=32,            # 批次大小:每个 Batch 包含的样本数
    shuffle=True,             # 是否每个 Epoch 打乱数据(训练集设为 True,测试集设为 False)
    num_workers=4,            # 多进程加载数据的线程数(通常设为 CPU 核心数,Windows 上设 0 或 2)
    pin_memory=True,          # 如果在 GPU 上训练,设为 True 可加速数据转移到 GPU
    drop_last=False           # 当总样本数不能被 batch_size 整除时,是否丢弃最后一个不够批次的数据
)


Iteratively using DataLoader in the training loop:


for epoch in range(num_epochs):
    for batch_idx, (inputs, targets) in enumerate(dataloader):
        # inputs 形状:  [batch_size, channels, height, width]
        # targets 形状: [batch_size]
        
        # 将数据搬运到训练设备 (GPU / CPU)
        inputs = inputs.to(device)
        targets = targets.to(device)
        
        # 此处接模型的训练逻辑(前向传播、反向传播)
        ...


Aligning sample data:


from torch.nn.utils.rnn import pad_sequence

def custom_collate_fn(batch):
    """
    batch 结构为列表: [(data1, label1), (data2, label2), ...]
    """
    data = [item[0] for item in batch]
    labels = [item[1] for item in batch]
    
    # 对变长文本/序列填充 padding,使其长度一致
    data_padded = pad_sequence(data, batch_first=True, padding_value=0)
    labels = torch.tensor(labels, dtype=torch.long)
    
    return data_padded, labels

# 在 DataLoader 中指定 custom_collate_fn
dataloader = DataLoader(dataset, batch_size=16, collate_fn=custom_collate_fn)

Phase 3: Building neural network architectures:

In deep learning, we build neural networks primarily to overcome the limitations of traditional machine learning and rule-based programming with high-dimensional, nonlinear, and complex structured data.


Thus, we introduce neural networks to help solve this pain point.


The main feature of neural networks is their rapid processing of nonlinearity, utilizing two components: __init__ and forward


First, let's directly build a neural network.


import torch
import torch.nn as nn

class MultiLayerPerceptron(nn.Module):
    def __init__(self, input_dim, hidden_dim, output_dim):
        # 1. 必须首先调用父类的 __init__
        super(MultiLayerPerceptron, self).__init__()
        
        # 2. 定义网络层组件
        self.fc1 = nn.Linear(input_dim, hidden_dim)  # 第一层线性变换 (全连接层)
        self.relu = nn.ReLU()                        # 激活函数
        self.fc2 = nn.Linear(hidden_dim, output_dim) # 第二层线性变换

    def forward(self, x):
        # 3. 定义数据流向
        # x 形状: [batch_size, input_dim]
        out = self.fc1(x)       # -> [batch_size, hidden_dim]
        out = self.relu(out)    # 引入非线性激活
        out = self.fc2(out)     # -> [batch_size, output_dim]
        return out

# 实例化网络
model = MultiLayerPerceptron(input_dim=10, hidden_dim=64, output_dim=2)
print(model)



We have just set up the network, but looking at this class alone might still be a bit abstract. Let's actually feed it some data and see what it outputs.


# 假设这一批有 8 个样本,每一个样本有 10 个 features
inputs = torch.randn(8, 10)

logits = model(inputs)

print("输入形状:", inputs.shape)   # torch.Size([8, 10])
print("输出形状:", logits.shape)   # torch.Size([8, 2])
print(logits[0])


The output here is not the final answer, but Logits. We setoutput_dim=2so each sample will get two logits. Later, the Loss Function will compare these Logits with the true Label to tell the model how far off its guess was.


There's another important point here: we didn't manually save the weights infc1andfc2one by one. As long as the Layer is written intonn.ModulePyTorch will automatically register them.


for name, parameter in model.named_parameters():
    print(name, parameter.shape)

# 输出大概会是:
# fc1.weight torch.Size([64, 10])
# fc1.bias   torch.Size([64])
# fc2.weight torch.Size([2, 64])
# fc2.bias   torch.Size([2])


This is whymodel.parameters()can be directly handed over to the Optimizer. The Optimizer doesn't need us to tell it where each parameter is; the Model has already prepared the list.


Activation Function and Loss Function

Earlier, we used ReLU. If there were no Activation Function, stacking many Linear Layers would essentially still be a larger Linear Function, and no matter how many layers are added, it wouldn't learn truly complex nonlinear relationships.


The logic of ReLU is simple:


\(\mathrm{ReLU}(x)=\max(0,x)\)


negative numbers are zeroed out, positive numbers are retained. It looks very brutal, but it really works well. Common Activations also include Sigmoid, Tanh, GELU, etc. Especially in the Transformers we discussed earlier, you often see activation functions like GELU.


After getting the output, we also need a Loss Function. Loss is the quantified result of the model's mistakes. The larger the Loss, the more off the prediction was; the smaller the Loss, the more likely the direction is correct.


loss_fn = nn.CrossEntropyLoss()

inputs = torch.randn(4, 10)
targets = torch.tensor([0, 1, 1, 0])

logits = model(inputs)
loss = loss_fn(logits, targets)

print("logits shape:", logits.shape)
print("loss:", loss.item())


Here's a common mistake: when usingCrossEntropyLossdon't apply Softmax to the output yourself. This Loss internally performs a more stable Log-Softmax process, so we can directly feed in the raw Logits.


Stage 4: Optimizer and the Complete Training Loop

We have finally arrived at the actual Training.


A training loop essentially repeats the following five steps endlessly:


1. Clear the gradients from the previous round;

2. Forward pass to obtain predictions;

3. Calculate how much is wrong using the loss function;

4. Backward pass to compute gradients;

5. The optimizer updates parameters based on gradients.


Let's directly create a small binary classification dataset. The data here is randomly generated with a simple goal: if the sum of two features is greater than 0, it's Class 1; otherwise, it's Class 0.


import torch
import torch.nn as nn
from torch.utils.data import TensorDataset, DataLoader

torch.manual_seed(42)

# 制造 1000 个二维样本
features = torch.randn(1000, 2)
labels = (features[:, 0] + features[:, 1] > 0).long()

dataset = TensorDataset(features, labels)
train_loader = DataLoader(dataset, batch_size=32, shuffle=True)

class TinyClassifier(nn.Module):
    def __init__(self):
        super().__init__()
        self.network = nn.Sequential(
            nn.Linear(2, 16),
            nn.ReLU(),
            nn.Linear(16, 2)
        )

    def forward(self, x):
        return self.network(x)

model = TinyClassifier()
loss_fn = nn.CrossEntropyLoss()
optimizer = torch.optim.Adam(model.parameters(), lr=0.01)


Adamis a very common optimizer. The learning ratelrcontrols how big each update step is. Too large may overshoot the correct answer, while too small may take forever to train.


Next comes the most important training loop:


num_epochs = 20

for epoch in range(num_epochs):
    model.train()
    total_loss = 0.0

    for inputs, targets in train_loader:
        # 1. 清除上一轮累加的梯度
        optimizer.zero_grad()

        # 2. Forward
        logits = model(inputs)

        # 3. 计算 Loss
        loss = loss_fn(logits, targets)

        # 4. Backward
        loss.backward()

        # 5. 更新 Parameters
        optimizer.step()

        total_loss += loss.item()

    average_loss = total_loss / len(train_loader)

    if (epoch + 1) % 5 == 0:
        print(f"Epoch {epoch + 1:02d} | Loss: {average_loss:.4f}")


It's best to memorize the order of these five steps because whether you're training a simple MLP, CNN, or a large Transformer, the outer framework remains quite similar.


optimizer.zero_grad()must be placed at the beginning for the reason we've mentioned earlier: PyTorch accumulates gradients by default. Forgetting to clear them will mix this round's gradients with the previous one. Unless you're intentionally doing gradient accumulation, this is usually a bug.


After training, we simply check the accuracy:


model.eval()
correct = 0
total = 0

with torch.inference_mode():
    for inputs, targets in train_loader:
        logits = model(inputs)
        predictions = logits.argmax(dim=1)

        correct += (predictions == targets).sum().item()
        total += targets.size(0)

accuracy = correct / total
print(f"Accuracy: {accuracy:.2%}")


model.eval()switches layers like Dropout and BatchNorm, which behave differently during training and evaluation, to evaluation mode.torch.inference_mode()tells PyTorch that we're only looking at answers and don't need to build a gradient graph. This saves memory and is faster.


However, note that we're still using training data here, so this accuracy only proves that the model roughly learned the rules, not that it performs well on unseen data. In real projects, you must also split validation and test sets, or you might get seemingly good but actually overfitted results.


Phase 5: Device, Saving the Model, and Inference

As mentioned earlier, one major advantage of Tensors over Numpy Arrays is that they can be moved to accelerators. The most common is NVIDIA CUDA; on Apple Silicon, you can try MPS; if neither is available, fall back to CPU.


if torch.cuda.is_available():
    device = torch.device("cuda")
elif torch.backends.mps.is_available():
    device = torch.device("mps")
else:
    device = torch.device("cpu")

print("Using device:", device)

model = model.to(device)


A very classic error is 'Expected all tensors to be on the same device.' This usually happens when the model is on GPU but the data is on CPU, or vice versa. Therefore, in the training loop, the data must also be moved accordingly:


for inputs, targets in train_loader:
    inputs = inputs.to(device)
    targets = targets.to(device)

    optimizer.zero_grad()
    logits = model(inputs)
    loss = loss_fn(logits, targets)
    loss.backward()
    optimizer.step()


If you don't save the trained model, everything will be lost after closing Colab or Python next time. The most common method is to save the state_dict, which are the model's parameters.


# 保存参数
torch.save(model.state_dict(), "tiny_classifier.pth")

# 重新建立同样的结构
loaded_model = TinyClassifier().to(device)

# 加载参数
state_dict = torch.load(
    "tiny_classifier.pth",
    map_location=device,
    weights_only=True
)
loaded_model.load_state_dict(state_dict)
loaded_model.eval()


Why rebuild the structure first? Because state_dict only saves weights and biases, not the entire Python class. The model architecture during loading must match the one at saving time.


Finally, we give it two new unseen samples to perform real inference:


new_samples = torch.tensor([
    [2.0, 1.0],
    [-1.5, -0.8]
], dtype=torch.float32).to(device)

with torch.inference_mode():
    logits = loaded_model(new_samples)
    probabilities = torch.softmax(logits, dim=1)
    predictions = probabilities.argmax(dim=1)

print("probabilities:", probabilities.cpu())
print("predictions:", predictions.cpu())


Note that Softmax finally appears here. During training, we feed raw logits directly to CrossEntropyLoss; during inference, we manually apply Softmax to show more intuitive probabilities to humans.


Some pitfalls I'm sure you'll encounter:

1. Shape mismatch. Often the model isn't failing to train - your [batch, features] just doesn't match the layer's required shape. When encountering issues, first print(tensor.shape), this really solves many problems.

2. Forgetting to clear gradients. The loss behaves increasingly strangely, only to discover later that zero_grad() was missing.

3. Device inconsistency. Model, input and target should be on the same device.

4. Forgetting model.eval() during evaluation. Especially when the model contains Dropout or BatchNorm, results will change.

5. Only watching training loss. Good training loss doesn't mean the model is truly good - validation is where you check generalization ability.


Summary

Alright, by now we've actually completed a minimal but complete PyTorch workflow:


Tensor
  -> Dataset / DataLoader
  -> nn.Module
  -> Forward
  -> Loss
  -> Backward
  -> Optimizer.step()
  -> Evaluation
  -> Save / Load
  -> Inference


The real challenge in PyTorch isn't how to write any particular Function, but rather connecting the entire chain of Data, Shape, Model, Loss, and Gradient. Once this pipeline is established, subsequent CNNs, Transformers, or even larger models essentially just involve adding new Layers and techniques to this framework.


Of course, this article only covers the Fundamentals. PyTorch still has a plethora of advanced topics like Learning Rate Scheduler, Mixed Precision, Gradient Clipping, and Distributed Training. But if you can rewrite this Training Loop yourself, you've already crossed the most critical threshold.


After such a long delay, this piece is finally complete. Next time you encounter nn.Linear, model.parameters(), or loss.backward() in Transformers, they shouldn't feel like magical incantations appearing out of thin air anymore.


Thanks for watching, and see you in the next episode~