promptdojo_

Optimizers and learning-rate schedulers — step 2 of 7

A generated benchmark runs SGD and Adam, both at lr=0.01, and reports "Adam is worse on our problem." What's wrong with the experiment?