This animation compares two loss functions used when training a small student network to mimic a large teacher network. It shows cross-entropy loss measuring the gap between a model's predictions and hard true labels, then contrasts this with distillation loss, which measures the gap between student and teacher soft probability outputs using temperature-scaled softmax. Useful for university students learning how model compression and teacher-student training combine both signals into one weighted objective.
Narrated · 16:9 · Preview before teaching · automatic layout checks do not establish subject accuracy
Explain cross-entropy loss and distillation loss in knowledge distillation Audience: freshman in the university