promptdojo_

Experiment tracker lite — step 2 of 7

Run 12 (new features, new lr, new seed, more data, new threshold) beats run 7 by two points, and the notebook declares the new features responsible. What does the one-change-per-comparison rule say?