Build a representative question set
Include common questions, difficult edge cases, permission-sensitive requests and questions that should not be answered. Test against realistic source documents rather than a small showcase set.
Score retrieval and generation separately
Measure whether useful evidence appears in the retrieved context, then assess whether the answer is supported, complete and appropriately uncertain. This prevents a fluent response from hiding poor retrieval.
Set release thresholds
Define acceptable quality, citation, latency, safety and cost thresholds before launch. Failed cases need documented fallbacks, escalation and improvement ownership.
References and further reading
Official and independent guidance used to support this practical overview. Product capabilities and pricing can change; verify current provider documentation before making a final decision.