Half of Our Benchmark Is Questions the Documents Cannot Answer
Our hallucination benchmark spends 60 of its 120 questions on things the corpus cannot answer. How those five categories were built, and what each one catches.

Most AI evaluations ask questions the system is supposed to be able to answer. That measures retrieval and synthesis, which is useful. It does not measure the failure that ends up in front of a regulator, which is a confident, well formatted answer to a question the documents never addressed.
So when we built the hallucination benchmark we published in March 2026, we split the test set down the middle. There are 120 questions. Sixty can be answered from the corpus, and 60 cannot. The corpus itself is real: 15 DDQs, ODD questionnaires, ESG surveys and APRA guidelines from an Australian equity fund manager.
Why Spend Half the Test Budget on Questions With No Answer
The obvious objection is that it halves your resolution on accuracy. Sixty answerable questions is a thin sample for a system that is meant to answer questions, and a reasonable person could argue we should have weighted the set far more heavily toward answerable questions.
We took the other view, because in DDQ and ODD work the inbound questions are not curated for you. An investor's questionnaire arrives with whatever the investor felt like asking, including questions about things your documents have never covered, questions aimed at a different entity, and questions built on a premise that is simply wrong. The system will meet those questions in production whether or not you tested for them. If refusal is only ever exercised by accident, you have no idea what its rate is.
That makes the ability to refuse a property under test, with its own score, rather than an edge case you note in a footnote.
How We Generated Questions That Could Not Be Answered
The 60 unanswerable questions were generated by an independent model, and each one had to be domain plausible but impossible.
That constraint is doing most of the work. "What is the capital of France" is unanswerable from a fund manager's DDQ corpus and worthless as a test, because nothing about it invites fabrication. The questions that induce fabrication are the ones that look exactly like the questions the system is supposed to answer. Same vocabulary, same shape, same register. The only difference is that the answer is not in the documents.
They come in five equal categories, because the ways a question can be unanswerable are not interchangeable, and a system can be good at refusing one kind while fabricating freely on another.
The Five Categories, and What Each One Catches
Another organisation's information, for example another manager's proxy voting policy. This catches the fact that retrieval returns nearest neighbours, not matches. Your own proxy voting policy is close to a perfect semantic match for a question about somebody else's. The chunk comes back with a high similarity score and nothing attached to it that says it belongs to the wrong entity.
Numbers and dates that appear in no document, for example a Sharpe ratio for a particular quarter. A question with a numeric shape exerts real pressure toward a numeric answer. This category tests whether the system can say that a figure is absent, which is harder than it sounds when the surrounding documents are full of adjacent figures that could be combined into something that looks right.
Non-Australian regulation, for example an SEC rule. This tests whether the system answers from its training data rather than from the corpus. The model knows a great deal about SEC rules from its training data and none of it is in your documents. A system that answers here is telling you that its control is semantic plausibility, not provenance, and that distinction will matter the first time the training data is out of date or wrong.
Internal detail a DDQ never contains, for example executive pay. This is the boundary between the corpus and the organisation. The topic is squarely in scope for the company and squarely out of scope for the documents, which is an uncomfortable place for a system that has been told it is the firm's knowledge base.
Real terms mixed with fabricated specifics, for example the outcome of an enforcement action that never happened. Here the false fact arrives inside the question as a presupposition. Refusing properly means contradicting the person asking, not just declining to answer them, and a hedged answer has still accepted the premise.
The paper reports the unanswerable half in aggregate, not per category, so it does not show which of the five was hardest for which system. If you build this set yourself, score the categories separately. It is the obvious improvement and we did not make it.
What the Three Systems Did
All three ran on identical infrastructure, so the pipeline was the only variable. System A was a plain model call with a financial compliance system prompt and no retrieval. System B was standard RAG, single embedding similarity search with top-k chunks in the prompt. System C was BackPro.
Results on the unanswerable half, 60 questions per system, March 2026:
| System | Refused | Hedged | Hallucinated | Refusal rate | Hallucination rate |
|---|---|---|---|---|---|
| Plain LLM | 42 | 6 | 12 | 70.0% | 20.0% |
| Standard RAG | 33 | 8 | 19 | 55.0% | 31.7% |
| BackPro | 59 | 1 | 0 | 98.3% | 0.0% |
The result worth sitting with is the middle row. Standard RAG hallucinated on nearly a third of the questions it could not answer, and hallucinated more than the model with no retrieval at all. Adding retrieval made fabrication worse, because retrieval always returns something. Ask about another manager's proxy voting policy and the index hands over the most similar chunks it holds, which are yours, and the model now has material in front of it that reads like an answer.
That regression is only visible because the plain model call was in the test as a control. An evaluation that compared one RAG system against another would have shown a difference in degree and missed the direction.
BackPro's 98.3% came mostly from one step. After extracting an answer, a separate model call asks whether that answer actually addresses the question. Below a confidence threshold, 0.5 by default, the system returns a structured refusal instead of a low-confidence answer. That is aimed squarely at the case this half of the benchmark produces: retrieval found something that looks right, and the extracted answer does not genuinely address the question.
Scoring, and Why "Hedged" Is Its Own Bucket
An independent model judge at temperature 0.1 saw three things for each item: the question, either the ground truth or the reason the question was unanswerable, and the system's answer. On the unanswerable half it scored Proper Refusal, Hedged or Hallucinated.
Hedged is separate from refused on purpose. A hedge is not a refusal that has been worded politely. It is an answer with a disclaimer attached, and in a DDQ response that is a risk, because a disclaimer is easy to lose when the response is assembled under deadline. Counting hedges as refusals would have lifted every system's refusal rate and told you less.
The same logic separates Incorrect from Hallucinated on the answerable half. Being wrong about a document you retrieved is a different failure from inventing a document, and the fix is different too.
If You Are Building This for a Vendor Evaluation
Use your own documents, not a public corpus. Have a model that has never seen the system under test generate the unanswerable questions, and hold it to the domain plausible constraint. Keep the five categories balanced, and score them separately. Score refusal, hedge and fabrication as three outcomes rather than two. Include a plain model call with no retrieval as a control, because that is the comparison that exposes a retrieval regression.
One limit in ours that you may want to close: our 60 unanswerable questions were generated, not harvested from real inbound questionnaires. Generated questions are balanced and reproducible. Real ones are messier and would probably be harder.
These are March 2026 results against the architecture as it stood then. On that run, across all 120 questions, BackPro's overall hallucination rate was 0.8%, a single borderline case, and it fell on the answerable half rather than this one.
