Here's something that surprises a lot of people. These two predictions are both correct.
Prediction A
Cat: 51%
Dog: 49%
Prediction B
Cat: 99.9%
Dog: 0.1%
Accuracy treats them exactly the same. Cross Entropy doesn't. It rewards confidence only when the model is correct.
If the true class is "Cat":
Prediction A gets a relatively high loss. Prediction B gets a very small loss.
Now flip the prediction.
Cat: 0.1%
Dog: 99.9%
The loss explodes. That's because Cross Entropy isn't asking:
Did you get it right?
It's asking:
How confident were you in the correct answer?
That's why neural networks optimize Cross Entropy instead of accuracy.
Accuracy is too coarse to guide learning.


