Generative AI

Evaluating an enterprise chatbot

A claim of "92% accuracy" means almost nothing unless you state which question set it was measured on and what counted as correct.

BigAI data engineering team
· 1 min read · Updated

1. Why RAG quality is hard to measure {#hard}

Unlike a classification problem, a RAG answer has no single correct string. Two differently worded answers can both be right, and a fluent answer can be completely wrong.

2. Four metrics BigAI uses {#metrics}

  • Citation accuracy. Does the cited document actually contain the information in the answer? This is the single most important metric.
  • Correct refusal rate. When the documents do not cover the question, does the system say “not found” instead of guessing?
  • Document coverage. What share of real user questions have a corresponding document at all? A low score here means you need to write documentation, not tune the model.
  • Real satisfaction. Measured through in-product feedback at the moment of use, not a survey sent afterwards.

3. Acceptance process {#process}

We build the test question set together with the customer at the start of the project — typically 150–300 questions taken from the real history of requests sent to HR and IT.

That set is then frozen and used to measure before, during and after rollout. Freezing it matters: a test set that drifts alongside the system cannot tell you whether anything improved.

Related content

How RAG works and when to fine-tune instead

The two approaches are often framed as alternatives. RAG solves a knowledge problem; fine-tuning solves a behaviour problem. Confusing them is a common way to overspend.

Cutting Spark cluster cost by 60%

Most Spark spend is wasted in places that are easy to fix: the wrong storage format, partitions that are too small, avoidable shuffles, and clusters idling overnight.

Start with a free 60-minute data assessment

A BigAI solution engineer will review your current data estate with you, identify the highest-value problem to solve and sketch a realistic roadmap. No commitment.

  • Data maturity assessment
  • 2–3 use cases with clear ROI
  • Budget and timeline estimate

By submitting this form you agree to the BigAI Privacy Policy.

Hotline Free consultation