A generated benchmark runs SGD and Adam, both at lr=0.01, and reports "Adam is worse on our problem." What's wrong with the experiment?
A generated benchmark runs SGD and Adam, both at lr=0.01, and reports "Adam is worse on our problem." What's wrong with the experiment?