Memory footprint, gradient checkpointing buffer, and backward pass activation overhead calculation for neural network architecture #250.
Activations and KV-cache scale linearly with batch size, requiring proportional GPU memory buffer allocations.
Gradient checkpointing recomputes activations during backward passes, reducing memory by up to 60% at the cost of ~25% compute overhead.
Yes, FlashAttention-2 and FlashAttention-3 avoid materializing the N×N attention matrix in HBM, reducing memory complexity from quadratic O(N²) to linear O(N).