The batched BCE-with-logits loss, and its gradient is the emitted cotangent #
BceLossCot identifies the bce := true renders' three-op cotangent ROW BY ROW: at example n
it is ∂bceLogits/∂z at that example's logits over the baked N·K. This file states the loss as
one function of the flat batch of logits — bceBatchLoss, every example's class-summed
BCE-with-logits over N·K, i.e. the mean over B×K — and proves its full gradient is the emitted
cotangent, read at the head's N·K index (bceBatchLoss_grad). It is SmoothedBatchLoss's twin
for ResNet-50's BCE artifacts, and like bceLossCotGraph_row it needs no hypothesis on the target.
The batched BCE-with-logits loss: Σ_n bceLogits(tₙ, zₙ) / (N·K), as a function of the
flat logits — the mean over B×K that timm's BinaryCrossEntropy computes.
Equations
- Proofs.bceBatchLoss N K t z x✝ = ∑ n : Fin N, Proofs.bceLogits K (Proofs.targetRow N K t n) (Proofs.logitRow N K z n) / (↑N * ↑K)
Instances For
The emitted BCE cotangent is the batched loss's gradient. Read at the head's N·K index
(unrowB), the three-op chain at the logits rowB z, with the committed divisor N·K, is
∇ bceBatchLoss at z — for every target.