Descent on the dense biases — the MLP rungs' bias columns #
SgdDescent.Linear and SgdDescent.Mlp state one-step descent for each dense WEIGHT matrix of the
Chapter-2 MLP (dense → relu → dense → relu → dense). This file states it for each dense BIAS:
- output bias
b₂—linear_bias_sgd_descends, constant2/(1 − 2D); - hidden bias
b₁—mlp_hidden_bias_sgd_descends, one frozen mask, constant2·d₃·w₂²/(1 − 2·w₂·D); - input bias
b₀—mlp_input_bias_sgd_descends, two frozen masks, constant2·d₃·d₂²·w₁²·w₂²/(1 − 2·w₂·d₂·w₁·D).
Each is its weight rung with the layer input replaced by the constant 1: the bias moves its
pre-activation by exactly the step (dense_bias_drift, no input bound a), so the weight rung's
a becomes 1 in the margins and the constants. The segment-Lipschitz step is
MlpSlot.loss_grad_lipschitz at a bias map (σ = ρ = 1 for the hidden layer, σ = w₁,
ρ = d₂·w₁ for the input layer), the gradient's row the channel indicator (pdiv_dense_b).
The rungs are generic in the layer's input, so the Chapter-3 CNN's dense-head biases are literal
instances at the pooled activation, as its head weights are of the weight rungs
(SgdDescent.Cnn). The oracle accuracy, the margins, the small-step and the two dominance
conditions remain hypotheses; no binary32 twin is stated for the biases.
The linear classifier's loss as a function of its bias.
Equations
- Proofs.linearBiasLoss W x label b = Proofs.crossEntropy n (Proofs.dense W b x) label
Instances For
Segment-Lipschitz gradient for the output-bias loss: under 2D < 1 the gradient entries
drift by at most (2/(1−2D))·(t·D) along [v, v+d].
One inexact SGD step on the output bias decreases one example's cross-entropy loss —
linear_sgd_descends with the input replaced by 1: example (x, label), the bias moving,
the weights fixed, constant C = 2/(1−2D) at step radius D = lr·(‖∇L‖₁ + n·η). The oracle
accuracy, the small-step and the two dominance conditions remain hypotheses.
The MLP's loss as a function of the hidden bias b₁.
Equations
- Proofs.mlpHiddenBiasLoss W₁ W₂ b₂ a₀ label b = Proofs.crossEntropy d₃ (Proofs.dense W₂ b₂ (Proofs.relu d₂ (Proofs.dense W₁ b a₀))) label
Instances For
The MLP's loss as a function of the input-layer bias b₀.
Equations
- Proofs.mlpInputBiasLoss W₀ W₁ b₁ W₂ b₂ x label b = Proofs.crossEntropy d₃ (Proofs.dense W₂ b₂ (Proofs.relu d₂ (Proofs.dense W₁ b₁ (Proofs.relu d₁ (Proofs.dense W₀ b x))))) label
Instances For
The input-bias loss is differentiable wherever both pre-activations are off the kinks.
Closed form of the input-bias loss gradient at a two-margin point:
∂L/∂b₀ⱼ = relu'(z₀ⱼ)·∑ₗ W₁ⱼₗ·relu'(z₁ₗ)·∑ₖ W₂ₗₖ·(softmax − onehot)ₖ.
The middle pre-activation moves by at most w₁·‖e‖₁ per entry under a step e of the input
bias — one dense crossing after a 1-Lipschitz ReLU.
Segment-Lipschitz gradient for the input-bias loss: MlpSlot.loss_grad_lipschitz at the
middle pre-activation, σ = w₁, ρ = d₂·w₁, the row relu₀'s frozen mask times W₁'s row.
Constant 2·d₃·d₂²·w₁²·w₂²/(1−2·w₂·d₂·w₁·D).
One inexact SGD step on the MLP's input bias decreases one example's cross-entropy loss —
mlp_input_sgd_descends with the layer input replaced by 1: the margins D < |z₀ⱼ| and
w₁·D < |z₁ₗ| at the step radius D = lr·(‖∇L‖₁ + d₁·η) freeze both masks, constant
C = 2·d₃·d₂²·w₁²·w₂²/(1−2·w₂·d₂·w₁·D). With this and the output and hidden bias rungs, each
dense bias of the MLP has a single-layer, single-example descent statement.