Your prompt tweak moved eval accuracy from 80% to 84% — measured on 25 examples. What's the defensible conclusion?