SmallParamGrad — the per-example kit for the chapter nets' loss gradients #
The seven ImageNet nets state their parameter gradients through ParamGradNodes, at the batched
*GradB nodes. The chapter nets (linear, MLP, MNIST CNN, the CIFAR CNNs) run one example at a time
and emit the per-example GradNode ops (SgdNodes). This file is their kit:
- Per node kind.
convW_hasGradAt,convB_hasGradAt,denseW_hasGradAt,denseB_hasGradAt: a node fed the gradient ofGat its op's output is∂G/∂θwith that one parameter varied. The fused*Sgdops the SGD renders emit areθ − lr·these nodes (convWeightSgd_eq_grad,convBiasSgd_eq_grad,weightSgd_eq_grad,biasSgd_eq_grad). - Per stage.
hasGradAt_dense,hasGradAt_relu,hasGradAt_conv: the loss read at a stage's output, pulled back through the stage's certified VJP, lands on the emitted chain's own cotangent. - The 2×2 pool, through a fixed selection.
σnames one cell of each window (poolSelIdx); the pool with its routing frozen atσis a gather, linear and differentiable everywhere (hasGradAt_gatherRelu), and its backward routes each window's cotangent to that ONE cell (selScatter) — what the renderedselect_and_scatter(select =GE) does.maxPool_relu_eventuallyEq_selis why that is the loss's gradient in the parameters at a tied window: if every window is dead, or tied only between cells that are the same function of the moving parameter (MaxPool2SmoothUpTo), the pooled ReLU IS the gather along the parameter, near the point. - The loss.
hasGradAt_crossEntropy: softmax cross-entropy at a hard label has gradientsoftmax − onehotin the logits.
Which cells are twins is per net: each net file names them (cells equal at every value of the
weights upstream of the pool) and discharges maxPool_relu_eventuallyEq_sel along each
parameter.
Conv weight node = ∇_W G.
Conv bias node = ∇_b G.
Dense weight node = ∇_W G.
Dense bias node = ∇_b G.
The fused *Sgd op is θ − lr· its un-fused *Grad peer, at any cotangent: the SGD renders
(the linear, MLP, MNIST-CNN and CIFAR arms) step by exactly the node the lemmas above identify.
Through a dense layer: the backward is W · dy (emitDenseBack).
Through a ReLU off its kink: the backward is the mask relu'(z) ⊙ dy (emitReluBack).
Through a stride-1 conv with odd kernels: the backward is the rendered reversed-kernel conv
(Back3.conv, conv_flatten_bridge).
The selection names a maximum of every window of u. The rendered select_and_scatter
(select = GE) makes such a choice; the canonical argmax is one (poolSelDom_argmax).
Equations
- One or more equations did not get rendered due to their size.
Instances For
Smooth, dead, or tied only between twins at the 2×2 windows (WindowSmoothUpTo), on the
pool's PRE-activation: a window is dead when its cells are all ≤ 0. The pre-activation margin
MaxPool2MarginQUpTo δ T implies it at any δ ≥ 0 (windowSmoothUpTo_of_margin).
Equations
Instances For
The 2×2 pool is continuous (a max of coordinates), so a pre-activation computed through earlier pools moves continuously with the parameters.
Through ReLU then a gather (y ↦ relu y ∘ σ), off the ReLU kinks: the backward is the scatter
along σ, then the ReLU mask. No pool hypothesis: the gather is linear.
Along a parameter, ReLU then the 2×2 pool is the gather at a fixed selection. Z θ is the
pool's pre-activation as the parameter moves; at θ₀ it has no zero entry, every window is
dead or tied only between T-twins, σ names a maximum of every window of its ReLU, and
T-twins are equal at EVERY θ. Then near θ₀ the pooled ReLU reads each window at σ: a
dead window stays negative, a strict maximum stays strict, and a twin stays tied with it.
The step ties' pool backward is the scatter at the first argmax. The Back3 maxpool node
(maxPoolBackDenote, the den of the rendered maxPoolBack) routes each window's cotangent
to maxPool2Argmax's cell, the window's first maximum, so through the flatten it is
selScatter along that selection, at every point, ties included. With poolSelDom_argmax,
the chain the step ties read is the loss-gradient chain at σ = maxPool2Argmax.