A general-purpose model returns prose. A queue, a router or a case system needs one value from a fixed list, in a field, every time.
Private Decision Models
AI that decides, not chats. Fine-tuned on your data, deployed inside your network, every decision returned with a probability.
Need an AI that completes a whole workflow? See AI Agent Development. Decision models do one thing: make a single, auditable call with a confidence score.
Why a general model is the wrong tool here
Sending customer records to an external API is a data-residency question before it is a technical one, and in a regulated institution that question is usually the end of the conversation.
If every answer arrives with the same apparent certainty, you cannot tell which ones a person needs to see. The probability is what makes automation safe to switch on.
One input, one decision, one number
The model reads an input — an email, a ticket, an alert, a document — and scores a fixed set of options that you define. It returns the option it selected and how confident it is. Decisions above your threshold are actioned automatically. Everything below it goes to a person.
The same input, two kinds of output
| Field your system needs | General-purpose LLM | Decision model |
|---|---|---|
| Team | Buried in a paragraph of prose | Card Operations · 93% |
| Reversal | Implied, never stated as a value | Yes · 75% |
| Urgency | “Today” mentioned in passing | Today · 92% |
An illustrative example. The values and probabilities are made up to show the shape of the output, not to claim an accuracy — what a model achieves depends entirely on your data, and we baseline before we promise anything.
Decisions banks and fintechs make thousands of times a week
Which department owns it, what priority it carries, which category it belongs to.
Complaint type, severity, and the SLA tier the clock should run against.
Request type, urgency, and the team that has to action it.
Close, escalate, or needs more information — advisory only. Your rules remain authoritative and the disposition stays with the analyst.
What the person on WhatsApp or chat actually wants, before anyone reads the thread.
Which document this is, for KYC and onboarding, so the right checks run against it.
Eight steps, one decision at a time
We do not start by training anything. We start by finding out whether the decision is worth automating and whether your data can support it.
-
Discovery
Pick one decision with clear business value. Volume, cost of getting it wrong, and who makes it today.
One decision, not a programme -
Data audit
Assess the historical labelled data you already hold — tickets, cases, alerts — and say plainly whether there is enough of it.
-
Label schema
Define the option list with your team, and the edge cases that sit between options. This is where most of the disagreement surfaces.
-
Baseline
Measure how accurate the current process is, human or rule-based. Without this number there is nothing to improve on.
Measure before you build -
Fine-tune
Train a small open model on your data, inside your environment.
-
Calibration and evaluation
Verify that “90% confident” really means roughly 90% correct, then set the thresholds for what is automated and what goes to a person.
A probability that means something -
Shadow mode
The model runs alongside your people with no live impact. You compare its calls against theirs for as long as you need to.
-
Go-live and monitoring
Drift tracking, periodic retraining, and an audit trail of every decision the model made and why.
It runs inside your network
The decision service sits beside your source systems, behind your own perimeter. Nothing is sent to an external API, and the hardware required is modest — this is a small model doing one job, not a frontier model.
Why bring this to us
Led by a founder with 18 years in quality engineering across fintech and regulated industries.
Models are tested like software: holdout sets, calibration checks and regression runs before anything is switched on, and after every retrain.
Documentation, validation evidence and a human in the loop by design — because the model has to survive a review, not just a demo.
You own the weights. The model is yours to keep, host, retrain or walk away with.
Questions we get asked
How much data do we need?
Fewer examples than people expect for a decision with a small number of options, and more than people expect once the option list grows or the classes are unbalanced. The data audit at step two answers this for your case specifically, before anyone commits to a build.
What accuracy should we expect?
We will not give you a number before seeing your data, and you should be wary of anyone who does. It depends on how consistent your historical labels are, how distinct the options are, and how much ambiguity sits between them. We measure your current baseline first, and we only automate above thresholds you have agreed.
Does any data leave our environment?
No. Training and inference both run inside your network. That is the point of the approach rather than a configuration option, and it is why a small fine-tuned model is the right tool here instead of a hosted frontier model.
What hardware is required?
Modest. These are small models making one decision, not general assistants, so a single GPU is usually sufficient for inference and fine-tuning is a short job rather than a standing cost. We size it against your actual volume during discovery.
How does this fit our model risk governance?
It is designed around it. You get the label schema, the baseline, the evaluation results, the calibration evidence and the threshold rationale as documentation, plus an audit trail of every decision in production. Human review below threshold is built into the flow rather than offered as a setting.
How long does a pilot take?
That depends on how quickly the data audit and label schema can be agreed, which is usually a question of your team's availability rather than ours. Shadow mode then runs for as long as you need to trust the result — we would rather it ran longer than shorter.
Start with one decision.
Prove it in shadow mode. Scale from there.
Rarely bought on its own.
Most teams need two or three of these together. Take the layer you are missing, or the whole stack.
Build agents that complete real work end to end — read the case, apply your policy, take the action and record why. Not chatbots that answer questions about the work.
Read moreProve how an AI system behaves before it goes live, and keep proving it afterwards. Evaluation suites, adversarial testing and drift monitoring built from your own cases.
Read moreBuild and modernise web, mobile and internal platforms, wired into the core systems you already run rather than sitting beside them.
Read moreRelease faster with automated regression suites, API and performance coverage, and real-device mobile testing wired into your pipeline.
Read moreAutomate regulatory reporting, financial crime operations and statutory filings — behind maker-checker, with an audit trail a regulator will accept.
Read morePut an agent on the channel your customers already use. Booking, support and follow-up over WhatsApp, handed to a human the moment it should be.
Read moreConnect the tools you already pay for and let agents run the steps between them — on a schedule, on a trigger, or on approval.
Read moreTell us what is slowing you down.
Half an hour, no deck. Describe the process and we will come back with whether it is worth automating, roughly what it would take, and what we would build first — or tell you plainly if it is not a job for us.