Home/ Articles/ [ a-written-test-set-catches-wrong-answers-before-your-custome ]

A written test set catches wrong answers before your customers do

Before any artificial intelligence (AI) assistant goes live, it should face a list of real questions with the right answers written down. This article explains why that written test set matters more than the AI model you choose.

A neat stack of index cards fanned out on a wooden desk under warm side lighting.

Before you let an artificial intelligence (AI) assistant talk to a customer, you need proof it answers correctly. Not a demonstration. Not a supplier's confidence. Proof, written down, that you checked.

That proof is a test set. It is a list of real questions your customers actually ask, with the correct answer written next to each one. Before launch, someone runs every question through the AI model and marks each answer right or wrong. If the model changes later, you run the list again. It costs little and it tells you, in plain figures, whether the assistant is ready.

Why this matters more than the model

Owners tend to assume a wrong answer means a weak model. Usually it does not. A survey of enterprises found that 64% had traced a confidently wrong AI agent answer back to missing or inconsistent company data in the past six months, and 32% had this happen more than once (Venturebeat). A separate survey found 56% of organisations had problems from AI systems running on incomplete or poorly governed data, with 17% calling the harm significant and 39% needing remediation after moderate disruption (Data Technologies). Another survey looking specifically at customer service found the real obstacles were compliance, security, disconnected systems, legacy infrastructure, skills gaps and unclean data, not the AI models themselves (SiliconANGLE). A test set is how you find these problems on your own terms, before a customer does.

Boundaries matter as much as facts. One test found that simply telling an AI model, in plain instructions, that anything not explicitly listed was off-limits cut its full failure rate from over 50% of attempts to under 10% (Yahoo Tech). Your test set should include questions that try to push the assistant outside its remit, not only questions it should answer well.

A test set is proof, written down, that you checked.

Despite how common these problems are, discipline around them is thin. Only 12% of organisations consistently measure whether AI's value justifies its cost (Technology For You), and in Malaysia only 18% of businesses have a documented process for escalating AI errors when they happen (The Vibes). A written test set, kept and rerun, is one of the simplest ways to close that gap.

What to ask a supplier

Ask whether a written test set is part of the build, not an afterthought. Ask who writes the correct answers, since it should be someone from your business who knows the right answer, not only the people building the assistant. Ask how often the test set is rerun after launch, and what happens when an answer comes back wrong. If a supplier cannot describe this process clearly, treat that as your answer.

Sources

  1. 64% of enterprises find AI errors in data, Venturebeat, 6 October 2026
  2. The hidden data costs missing from AI budgets | TechTarget, Data Technologies, 2 October 2026
  3. AI in customer experience has an orchestration problem, not an adoption problem, SiliconANGLE, 3 October 2026
  4. Are AI Agents Going Rogue? Here’s What Business Leaders Should Know, Yahoo Tech, 1 October 2026
  5. New KPMG AI Pulse Survey: As AI maturity converges, leading organizations show what AI at scale requires, Technology For You, 5 October 2026
  6. Malaysian workers race ahead of employers in AI adoption, The Vibes, 20 September 2026