Skip to content

Customer service20:04

Measuring AI assistant quality across 500 test pairs and a 100-response panel

How to combine automated evaluation with a human panel: 8–14 intents, 500–1000 test pairs, 100 responses every 2 weeks, and an audit every 3 months.

AI-generated image.

Listen to the article summary

Synthetic voice.
0:00
0:00

Here we document how these things work in our implementations: this is an open guide, not legal advice tailored to your company, so leave the final word to your lawyer. In internal meetings about assistant quality, people increasingly claim that automated evaluation is enough. You enable RAGAS or TruLens, read faithfulness 0.92, answer relevancy 0.87, and announce that the bot responds well. That sounds convincing: you get numbers, they are repeatable, they do not depend on reviewer fatigue, and they can scan the entire conversation history overnight. There is a lot of sense in that. The problem starts when we treat the LLM judge score as the only proof that the assistant does no harm.

The strongest counterargument

Advocates of hard metrics are right about one thing: at 10,000 conversations a day, a human will not read even 1% of the logs, whereas an automated system will compare every response with the context retrieved from the database. Tools like RAGAS and TruLens measure whether a response sticks to the sources and addresses the question. In our tests, such monitoring catches regression after updating a knowledge base in 30–40 minutes, before reaching even a single user. If an assistant handles 1,000 queries and 920 end with a correct answer, accuracy is 92%, and that is a metric you can show to the board.

AI-generated image.

What this argument overlooks

Automated metrics measure consistency with retrieved context, not the truthfulness of the context itself. If the knowledge base contains an outdated fee table or obsolete delivery times, a faithfulness score of 0.94 means the system faithfully repeated the error. The second blind spot is overconfidence: a response can be plausible while promising a refund absent from the policy. No LLM-as-a-judge metric will catch this. In e-commerce, a single overconfident delivery promise can generate dozens of returns; in banking, a wrong answer about fees results in a complaint an analyst reads for two hours. On top of that comes the cost of eroded trust: after one bad response, a user rarely returns to the bot, and post-interaction CSAT drops by double digits before any automated tool detects the issue.

Our position

Sound evaluation is a hybrid: automated testing after every change, a human review panel every two weeks, and a quarterly audit on 500–1,000 question–acceptance criterion pairs. In our process, automated tools cover 100% of the logs, but we make the decision to keep the system in production only after the panel review. We build the test set from the 8–14 most frequent production intents, collecting 40 to 120 questions for each intent: 30–35 examples per class is the threshold for a margin of error around 10 percentage points with a 95% confidence interval. Below this sample size, the difference between 85% and 88% accuracy is statistical noise.

Beyond accuracy, we also measure the human escalation rate, time to resolution, raw CSAT, and cost of appeals. An escalation rate above 25% or a CSAT drop of more than 10 points at the same accuracy level signals that the assistant hands off to humans too often, or that users are not getting what they need. Regulatory requirements also apply: the AI Act mandates data quality monitoring and transparency, and we pseudonymize conversation logs before human review, because analyzing user inputs constitutes personal data processing under GDPR.

AI-generated image.

For this quarter's decision, this means one thing: a one-off RAGAS report is not enough. Build a dataset of 500–1000 pairs from production, run a bi-weekly panel on 100 responses, and schedule the first quarterly audit. Only then can you see whether accuracy is improving or drifting along with seasonal promotions. We would revise this position if the assistant operated in a closed domain, such as order status from a single system, where every answer can be verified deterministically and automation fully replaces human review. That is not the case for open questions about products, fees, and timelines.

Source materials

Sources consulted during research. The text above is our own; these third parties are not responsible for its content and have not authorised it.

No-obligation call

Schedule a workshop with Paul

Thirty minutes or a full workshop, choose what suits you. Book a slot directly in the calendar and get instant confirmation.

Paul Lazniak

HEXART Founder