If it is not measured, it is a vibe
Prompt changes feel like progress until support volume ticks up. We keep a golden set in git and run it on every prompt PR.
Clients get a simple dashboard: accuracy vs. the set, refusal rate, and latency. That is enough to decide whether a model swap is worth it.
