Case study — 01 · Banking

Sanctions Screening Agent

Watchlist alerts arrive all day and most of them are nothing. The agent clears the obvious ones, and sends every case it is not sure about to a person with the evidence already gathered.

The problem

What it was like before

Sanctions screening produces a queue that never empties. The overwhelming majority of alerts are false matches — a name that resembles a listed one, a date of birth that half agrees — and each still has to be opened, checked against the underlying record, dispositioned and evidenced.

The cost is not the decision. It is that an analyst spends the day on alerts that were never going to be true matches, and the ones that matter wait behind them.

How it looks

What the analyst sees

The row that matters is the third. Rules and model disagreed, so nothing cleared automatically and it went to a person with the evidence already attached.

Alert queuetoday · 1,284 processed
AHMED RAZADOB partial · list: UN 12670.62Cleared
M. A. KHANName only · list: OFAC SDN0.55Cleared
S. HUSSAINName + DOB · list: UN 12670.91EscalatedRules say hit, model says no. Disagreement always goes to a person — evidence attached.
FATIMA N.Transliteration · list: EU CFSP0.48Cleared
TARIQ M.Name + nationality · list: OFAC SDN0.87Escalated
What we built

What it does

Alert ingestion

Watchlist alerts are collected continuously rather than pulled in batches, so the queue reflects the position now.

Rule engine, then a model

Deterministic rules run first and decide what they can. The model is asked only about what the rules could not settle, which keeps the reasoning auditable where it can be.

Evidence assembled before escalation

When a case goes to a human it arrives with the match, the underlying record and the reason for doubt already attached.

Disagreement is the designed path

Where rules and model disagree, nothing is auto-cleared. That case is escalated by definition.

Read-only oversight

A dashboard Compliance can watch without operating, so nobody has to chase a queue to know where it stands.

A daily report

Branded, scheduled, and the same every day — which is what makes it usable as evidence rather than as a status update.

Where it got to

In production

AutoFalse matches cleared without a person
HumanEvery disagreement escalated
DailyBranded compliance report
Our view

The interesting decision was what not to automate

It would have been straightforward to let the model disposition everything and report a confidence score. We did not, because a confidence score is not an audit trail and a regulator does not accept one.

Rules decide what rules can decide. The model handles the residue. Where they disagree, a person decides — and that disagreement is treated as signal about the rules rather than as noise to be suppressed.

Built with

Tech stack

  • Risk Nucleus
  • Playwright
  • Claude
  • Node.js
  • Prisma
  • PostgreSQL
  • Docker
The work behind it

Delivered under Compliance & RegTech Automation and AI Agent Development.

More work

Other things we have built.

Banking

Law-Enforcement Request Intake

Requests arrive as letters with scanned attachments. The agent reads the mailbox, OCRs the attachments, verifies identifiers against core banking, classifies the intent and routes it — with an audit trail behind every step.

Read it
Banking

Internal Policy Assistant

Staff ask a policy question in plain language and get an answer drawn strictly from approved documents, with the source named. Nothing leaves the bank, and the assistant says so when the documents do not answer.

Read it
Compliance

CTR to goAML Conversion

A compliance team spent hours a week turning Currency Transaction Reports into goAML XML by hand. Parser, XML builder, schema validator and audit trail, behind a interface officers actually use.

Read it
Compliance

Regulatory Reporting Engine

A registry of every periodic return a bank owes. The engine pulls the data, builds each report against the mandated template, then schedules and submits it behind a maker-checker step.

Read it
Product

Autonomous Social Media Agent

A ReAct loop that researches, reads its own history to avoid repeating itself, writes a strategy memo, critiques its own drafts, publishes, and learns from every rejection.

Read it
Product

Multi-Agent Marketing System

Two agents on one platform with a person approving everything before it ships. One researches and takes a position; the other plans a week of posts reviewed as a batch.

Read it
Healthcare

WhatsApp Booking Agent

Patients book on the number a clinic already advertises, at any hour, in the language they normally write. Only genuinely free slots are offered, and anything clinical goes to a person immediately.

Read it

Have something like this?

Describe the process and we will come back with whether it is worth automating, roughly what it would take, and what we would build first — or tell you plainly if it is not a job for us.