Your real past cases and the decisions your staff made, frozen as the standard the system has to meet.
AI Assurance & Evaluation
Proof of how an AI system behaves — before it goes live, and every month afterwards. We came out of software testing, so we build the evidence first and the agent second.
What we build
The artefacts a risk committee or a regulator asks for. Ask about anything you do not see here.
Bad scans, missing fields, near-miss names, prompt injection — the failures we would rather find than have an auditor find.
Every AI system you run, with its purpose, owner and risk rating. What the CBUAE guidance asks for, and what most banks lack.
Repeatable testing for disparate outcomes on high-impact decisions, designed to give a comparable result next year.
Alerting when live behaviour moves away from the approved baseline, because it will even when nothing was changed.
Every human correction captured and fed back as a new test case, so the system gets harder to fool over time.
Per decision: inputs, rule applied, output and approver, in a form an auditor accepts without a translation layer.
The whole suite re-run on every change of rule, prompt or model version, so you see what moved before it ships.
How much it is allowed to do is your decision, per process.
Nothing starts at the top. A process moves up only when the evidence supports it, and your risk function makes that call.
It drafts. A person writes and commits.
It proposes a decision. A person approves it.
It acts. Anything unusual routes to a human.
It acts inside stated limits. People audit samples.
How an assurance engagement works
Four stages. The first happens before any agent exists.
-
Build the test set from your history
Real past cases and their known-correct outcomes become the specification. This is the work, and it is worth doing before anything is built.
Written first -
Attack it deliberately
Edge cases, malformed input and deliberate attempts to mislead it. A pass rate on happy paths tells you nothing useful.
Red team, not spot check -
Wire it into the pipeline
Every change of rule, prompt or model re-runs the suite. You see the difference before it ships, not after a complaint.
No silent regressions -
Watch it in production
Decisions measured against the approved baseline, with alerts when the shape changes and corrections fed back as tests.
Assurance does not stop at launch
Questions we get asked
No — accountability for third-party systems sits with you under CBUAE guidance, and we test what you run regardless of who wrote it.
Two to four weeks for a first process, most of it spent agreeing what a correct outcome actually is.
Yes. It is yours to re-run whether or not we are still working together.
That is the common case. We start with what it is doing now and establish the baseline from there.
Rarely bought on its own.
Most teams need two or three of these together. Take the layer you are missing, or the whole stack.
Build agents that complete real work end to end — read the case, apply your policy, take the action and record why. Not chatbots that answer questions about the work.
Read moreBuild and modernise web, mobile and internal platforms, wired into the core systems you already run rather than sitting beside them.
Read moreRelease faster with automated regression suites, API and performance coverage, and real-device mobile testing wired into your pipeline.
Read moreAutomate regulatory reporting, financial crime operations and statutory filings — behind maker-checker, with an audit trail a regulator will accept.
Read morePut an agent on the channel your customers already use. Booking, support and follow-up over WhatsApp, handed to a human the moment it should be.
Read moreConnect the tools you already pay for and let agents run the steps between them — on a schedule, on a trigger, or on approval.
Read moreSmall, private AI models that make one decision well — route, classify, triage — with a confidence score your team can trust. Runs inside your network.
Read moreTell us what is slowing you down.
Half an hour, no deck. Describe the process and we will come back with whether it is worth automating, roughly what it would take, and what we would build first — or tell you plainly if it is not a job for us.