AI + QA

The CBUAE wants an inventory of every AI system you run. Here’s what that takes.

On 23 February 2026 the Central Bank of the UAE issued its Guidance Note on Consumer Protection and the Responsible Adoption and Use of Artificial Intelligence and Machine Learning by Licensed Financial Institutions. It is guidance rather than a rulebook amendment, which has led some teams to file it under “read later”. That is a mistake — supervisory guidance is what examiners ask about, and three of its expectations take months to satisfy rather than weeks.

The note sets out five principles: governance and accountability, fairness and non-discrimination, transparency and explainability, effective human oversight, and data management and privacy. Most of that is unobjectionable and most institutions will nod along. The work sits in three specific obligations.

1. An inventory of every AI system you run

You are expected to hold a documented inventory of AI and ML systems, each with its purpose and a risk rating. This sounds like a morning’s work and is not, for one reason: nobody knows what counts.

Your credit scorecard is obviously in scope. But what about the fraud engine your card processor runs on your behalf? The CV screening feature inside your HR platform, switched on by default? The chatbot marketing added to the website last year? The forecasting model a treasury analyst built in Python and has quietly run ever since?

Every institution that has done this exercise honestly has found more systems than it expected, and the surprises are rarely in the places with model risk governance already. They are in procurement, HR and marketing — functions that bought a tool with an AI feature and never thought of it as deploying a model.

The inventory is not the hard part. Discovery is. Budget for a genuine sweep across every function that has bought software in the last three years, not a circulated spreadsheet that comes back with four rows.

2. Annual bias testing on anything high-impact

Systems used in high-impact decisions are expected to be tested at least annually — and again whenever they are materially changed — for unintended bias, discriminatory outcomes and model drift.

Read that as a standing obligation rather than a one-off project. An annual test you cannot repeat identically next year is not a test, it is an anecdote. Which means the deliverable is not a report; it is a reusable test set and a documented procedure that produces a comparable result twelve months later, ideally run by someone other than the team that built the model.

This is where most institutions discover the real gap. Testing for disparate outcomes requires a labelled dataset with the protected characteristics you are testing against, and a defined threshold for what counts as a problem. Very few have either. Assembling that — lawfully, with the right approvals — takes longer than running the test.

Start it before you need the result.

3. You are accountable for models you did not build

The expectation of board and senior management accountability is explicit that it extends to third-party systems. There is no carve-out for “the vendor handles that”.

This is the one that changes procurement. If a supplier cannot tell you what their model was trained on, how it is monitored, how often it is retrained, and what happens when it drifts, you cannot evidence oversight of it — and the accountability still sits with your board. Those questions belong in the RFP, not in a remediation exercise eighteen months later.

Expect pushback. Many vendors genuinely cannot answer, and some will treat the questions as intrusive. That is information too.

What this actually looks like in practice

The institutions handling this well are treating it as an operating discipline rather than a compliance exercise, and the sequence tends to be the same:

  • Sweep for systems before writing policy. A governance framework written against an incomplete inventory describes a bank you do not have.
  • Rate by consequence, not by sophistication. A simple rules engine that declines credit is higher impact than a clever model that suggests marketing copy. The regulator cares about the effect on the customer.
  • Build the test set once, reuse it every year. The cost is front-loaded; the annual obligation then becomes routine instead of a fire drill.
  • Write down where a human decides. “Effective human oversight” is not a person with a dashboard. It is a defined point where a decision stops and waits, enforced by the system rather than promised in a document.
  • Keep the evidence as you go. Reconstructing why a model made a decision eight months ago, from logs never designed for it, is far more expensive than logging it at the time.

The uncomfortable part

Nothing above requires new technology. It requires knowing what you are running, being able to show how it behaves, and having somebody accountable when it does not. Institutions that already do this for their credit models will find the extension manageable.

The ones that will struggle are those that adopted AI through product features and pilots rather than through model risk governance — which, candidly, is most of them. The guidance does not create that problem. It just makes it visible on a timetable.

If you are assembling an inventory or standing up repeatable testing for the first time, that is the work our AI Assurance & Evaluation practice does — evaluation suites built from your own cases, adversarial testing, and drift monitoring that produces the same evidence a supervisor is going to ask for.

This article summarises publicly available guidance and is not legal advice. Read the guidance note in full on the CBUAE Rulebook.

Building something that has to hold up?

Bring the process, not a specification. We will tell you honestly whether an agent is the right answer.