Library

Research · PDF, 14 pages

BackPro hallucination benchmark

120 compliance questions, half of them unanswerable by design, put to three systems on identical infrastructure so that only the pipeline differed. Method and results, March 2026.

Written for compliance, risk and technology teams evaluating AI.

Download the PDF Free to download. No form.

What you will take away

  1. 01

    Overall made-up answers across all 120 questions: BackPro 0.8% (one borderline case), a plain model 15.8%, standard retrieval (RAG) 28.3%.

  2. 02

    On the 60 answerable questions, accuracy was 96.7% for BackPro, 30.0% for standard RAG and 28.3% for the plain model.

  3. 03

    On the 60 questions with no answer in the documents, BackPro correctly refused 98.3%, the plain model 70.0% and standard RAG 55.0%.

  4. 04

    Adding retrieval made fabrication worse, not better, because a search always returns something. The difference came from checking whether an answer is supported and allowing the system to refuse.

Inside the research

  • What was tested: three systems, one model, one infrastructure
  • Why half the questions cannot be answered
  • Scoring by an independent model judge
  • Results on answerable and unanswerable questions
  • Why the pipeline made the difference

Questions it answers

Why were half the questions unanswerable?
Because in regulated work the ability to refuse is a feature under test, not an edge case. The 60 unanswerable questions were plausible but impossible: other companies, numbers in no document, overseas rules, detail a DDQ never contains, and real terms mixed with invented specifics.
Why does 28.3% appear twice?
They are two different measures that happen to match: standard RAG’s overall rate of made-up answers, and the plain model’s accuracy on the answerable questions.