The MNIST MLP — every parameter gradient node IS the loss's derivative #
mlp_train_step_tied_certified ties each of the six SGD updates to the certified per-layer
Jacobian contracted with the cotangent the emitted chain threads to it (mlpCotOut1,
mlpCotOut0), and leaves open whether that cotangent is the loss gradient at the hidden layers.
mlp_net_lossGrad closes it: at the same cotangents, the un-fused weightGrad / biasGrad node
of each layer is the gradient of the loss in that parameter, for any loss L of the logits with
gradient g there; mlp_net_lossGrad_CE instantiates it at the softmax cross-entropy the render
emits. The fused weightSgd / biasSgd ops are θ − lr· these nodes
(SmallParamGrad.weightSgd_eq_grad, SmallParamGrad.biasSgd_eq_grad).
How. The loss read at the logits is pulled back one certified stage at a time
(SmallParamGrad.hasGradAt_dense, SmallParamGrad.hasGradAt_relu), each landing on the emitted
chain's cotangent; at each layer's output the node lemma (SmallParamGrad.denseW_hasGradAt)
turns it into the parameter gradient.
Hypotheses. Both hidden pre-activations off the ReLU kink (the pair mlpHasVJPAt takes).
Scope. One example (the emitted module batch-contracts; den is per-example).
Every MLP parameter node is the gradient of L in that parameter: the six nodes, each at
the cotangent the emitted chain threads to its layer (g at the logits, mlpCotOut1,
mlpCotOut0), stated against L of mlpForward with that one parameter varied.
Equations
- One or more equations did not get rendered due to their size.
Instances For
Every MLP parameter node is the gradient of L in that parameter, whenever g is L's
gradient at the logits and both hidden pre-activations are off the ReLU kink.
The artifact's loss: every node is the gradient of the softmax cross-entropy at label,
g the emitted loss cotangent.