How to Evaluate SoA Automation: Five Questions Worth Asking Any Vendor
How to evaluate SoA automation before you buy: five questions about templates, deterministic checks, data location and what the software refuses to do.

Every Statement of Advice demo looks the same. A fact find goes in, a document comes out, and the room goes quiet for a moment because the document is long and it arrived in about the time it takes to make coffee.
That moment tells you almost nothing. The demo runs on the vendor's file, with the vendor's template, driven by somebody who knows exactly where the thing breaks and steers around it. What you want to know is what happens on a Tuesday in your practice, with your template, on a file where the client's superannuation balance appears in four different sections and one of them is a projection.
These are the questions that get at that. We sell SoA automation, so read the answers with the discount you would apply to any vendor. The questions themselves work on us as well as on anyone else, which is rather the point of publishing them.
First: does the SoA stay in our template, and can we change it without you?
Underneath every advice generator there is a decision about what a firm's template is. It is either data or it is scenery.
If it is data, the system reads the document the practice already uses and classifies each bracketed marker sitting in the prose. A marker that names a fact on file, a balance or a premium or a fee, is looked up. A marker that decides whether the surrounding paragraph appears at all is evaluated to true or false. Only a marker that is an instruction to whoever writes the section, "explain why this product was chosen", is a writing job. The template is the input.
If it is scenery, the vendor has rebuilt an approximation of your document inside their generator and dressed it in your letterhead. It will look identical in the demo. The difference shows up the first time your compliance committee changes a paragraph, because that change becomes a request on somebody else's roadmap instead of an afternoon's work.
The test is simple enough to run in the meeting. Ask to open the template, change a sentence, and produce the document again. A firm that cannot edit its own advice document without asking permission has handed over something it remains legally accountable for.
Second: which checks are deterministic, and what was each one measured against?
"The AI checks it" is not an answer. It is the absence of one.
The useful split is between checks that are ordinary software, which behave identically every time and can be measured, and checks that are another model call, which cannot be measured the same way because the thing doing the checking is the thing you are worried about. Placeholder scanning, figure comparison, section completeness and disclosure ordering are all in the first category. Judging whether a strategy suits a client is firmly in the second, which is why it stays with the adviser.
The second half of the question is the half vendors skip. A check that was never measured against the failure it exists to catch is a comfort feature.
Ours is the example I would use, because it is the one where we found the answer unflattering. We built a figure check that compared every dollar amount in a draft against the numbers held on the client's file and flagged anything that did not appear there. It ran, it fired, it caught things. Then we measured it against substitution, which is a real figure from the client's record placed in the wrong sentence, and it caught none of them. Not rarely. Never, by construction, because every substituted value is already somewhere in the file, so a check asking "is this number in the file" says yes to exactly the case you wanted it to catch. The rebuild binds each figure to the concept its sentence names.
You do not need a vendor to have a perfect record. You need them to have a number, and to know what it was measured against. Ours is a dev-corpus figure that we will put in front of a firm evaluating us, on the same terms we would want somebody else's, which is with the method attached rather than the headline on its own.
Third: where does the client data sit while the work happens?
"Our cloud is secure" is a different answer from "the work happens in your tenancy", and only one of them survives a licensee questionnaire.
Ours runs inside the customer's own tenancy. The honest qualifier, which belongs in the same breath, is that data residency then follows the subscription the customer brings. It is not a toggle we can flip on request, and any vendor telling an Australian practice that client data categorically never leaves the country should be asked to put the mechanism in writing rather than the assurance.
Ask where a document sits while it is being drafted, where it sits afterwards, and what leaves the tenancy at any point in between. The third part is the one that surprises people.
Fourth: what happens when a check fires?
This is the question nobody asks, and it decides whether the checking is worth anything in practice.
A flag that arrives as a report is work. Somebody opens the report, reads a list of page references, opens the document, finds each sentence, decides what the flag meant, and fixes it. Run that on every file and the review time you saved on drafting comes back to you in a different envelope.
A flag that arrives as a link to the sentence, with a suggested fix attached and an undo behind it, is not work in the same sense. In our editor the adviser jumps straight to the flagged figure's section and asks for the change in chat, and the flag is closed in the document rather than in a spreadsheet next to it.
So ask to see a failed check rather than a successful one. Ask what the reviewer does next, and count the clicks.
Fifth: what does the software refuse to do?
A vendor who will not name a limit has not found the limits yet, and you will find them instead.
The suitability judgement, the review and the sign-off do not move. They are not slow because software has not reached them, they are slow because they are the part where a qualified person takes responsibility for advice given to a retail client. The best interests duty sits in sections 961B to 961J of the Corporations Act, with ASIC RG 175 as the guidance on meeting it, and it attaches to the advice and the adviser rather than to the tool that produced the draft. Automation can make that duty easier to evidence. It cannot carry it.
The same scepticism belongs on the numbers in the brochure, including ours. Frazer Walker, an AFSL holder, put their own reduction in SoA production time at about 60 per cent. That is one firm, their templates, their process, reported by them. It is a reasonable data point and a poor forecast for a practice with different documents, and any vendor quoting a figure like that as an industry norm is quoting it wrong.
What a number like that does tell you is where the hours go when they go. They come out of finding the approved wording and producing the first draft, which are the tasks that repeat identically on every file. They do not come out of the thinking. If a vendor's savings claim is large enough that it must include the thinking, the claim is the answer to the fifth question, and the answer is a bad one.
The shape of a good answer
None of these five questions is about model capability, and that is deliberate. Model capability is the part of the stack that changes every few months and the part every vendor will happily talk about for an hour.
The parts that decide whether an advice document is safe to sign are duller than that and they move slowly: whether the template belongs to you, whether the checks are real and measured, where the data sits, what happens on a failure, and where the software stops. Our own SoA automation is built on answers to those five, which is also why we can be specific about the places it does not reach.
A vendor who answers all five without flinching is not necessarily the right one for your practice. A vendor who cannot answer them has told you something more useful.
