Services — 02

AI Assurance & Evaluation

Proof of how an AI system behaves — before it goes live, and every month afterwards. We came out of software testing, so we build the evidence first and the agent second.

What we build

What we build

The artefacts a risk committee or a regulator asks for. Ask about anything you do not see here.

Evaluation suites

Your real past cases and the decisions your staff made, frozen as the standard the system has to meet.

Adversarial test sets

Bad scans, missing fields, near-miss names, prompt injection — the failures we would rather find than have an auditor find.

Model inventories

Every AI system you run, with its purpose, owner and risk rating. What the CBUAE guidance asks for, and what most banks lack.

Bias and fairness testing

Repeatable testing for disparate outcomes on high-impact decisions, designed to give a comparable result next year.

Drift monitoring

Alerting when live behaviour moves away from the approved baseline, because it will even when nothing was changed.

Override logging

Every human correction captured and fed back as a new test case, so the system gets harder to fool over time.

Audit evidence export

Per decision: inputs, rule applied, output and approver, in a form an auditor accepts without a translation layer.

Regression harnesses

The whole suite re-run on every change of rule, prompt or model version, so you see what moved before it ships.

The dial

How much it is allowed to do is your decision, per process.

Nothing starts at the top. A process moves up only when the evidence supports it, and your risk function makes that call.

L1 — Assist

It drafts. A person writes and commits.

L2 — Recommend

It proposes a decision. A person approves it.

L3 — Act with a gate

It acts. Anything unusual routes to a human.

L4 — Act within limits

It acts inside stated limits. People audit samples.

How it runs

How an assurance engagement works

Four stages. The first happens before any agent exists.

  1. Build the test set from your history

    Real past cases and their known-correct outcomes become the specification. This is the work, and it is worth doing before anything is built.

    Written first
  2. Attack it deliberately

    Edge cases, malformed input and deliberate attempts to mislead it. A pass rate on happy paths tells you nothing useful.

    Red team, not spot check
  3. Wire it into the pipeline

    Every change of rule, prompt or model re-runs the suite. You see the difference before it ships, not after a complaint.

    No silent regressions
  4. Watch it in production

    Decisions measured against the approved baseline, with alerts when the shape changes and corrections fed back as tests.

    Assurance does not stop at launch
Common questions

Questions we get asked

We did not build the model. Does that matter?

No — accountability for third-party systems sits with you under CBUAE guidance, and we test what you run regardless of who wrote it.

How long does a test set take?

Two to four weeks for a first process, most of it spent agreeing what a correct outcome actually is.

Do we keep the suite?

Yes. It is yours to re-run whether or not we are still working together.

Can you assure something already live?

That is the common case. We start with what it is doing now and establish the baseline from there.

The other 7

Rarely bought on its own.

Most teams need two or three of these together. Take the layer you are missing, or the whole stack.

AI Agent Development

Build agents that complete real work end to end — read the case, apply your policy, take the action and record why. Not chatbots that answer questions about the work.

Read more
Custom Software Development

Build and modernise web, mobile and internal platforms, wired into the core systems you already run rather than sitting beside them.

Read more
QA & Test Automation

Release faster with automated regression suites, API and performance coverage, and real-device mobile testing wired into your pipeline.

Read more
Compliance & RegTech Automation

Automate regulatory reporting, financial crime operations and statutory filings — behind maker-checker, with an audit trail a regulator will accept.

Read more
WhatsApp AI Assistants

Put an agent on the channel your customers already use. Booking, support and follow-up over WhatsApp, handed to a human the moment it should be.

Read more
Workflow Automation

Connect the tools you already pay for and let agents run the steps between them — on a schedule, on a trigger, or on approval.

Read more
Private Decision Models

Small, private AI models that make one decision well — route, classify, triage — with a confidence score your team can trust. Runs inside your network.

Read more

Tell us what is slowing you down.

Half an hour, no deck. Describe the process and we will come back with whether it is worth automating, roughly what it would take, and what we would build first — or tell you plainly if it is not a job for us.