promptdojo_

Precision, recall, and threshold tradeoffs — step 2 of 7

Two features ship this quarter: (a) auto-refund — the model's positive verdict refunds money with no human in the loop; (b) support-ticket triage — flagged tickets get a human look. Which metric does each one optimize at the threshold?