A number is not an explanation
A risk score can help order work, but the analyst’s real questions begin after the number appears: What changed? Which transactions or customer facts contributed? Is the source current? What information is missing? What action does policy require? If the interface cannot answer those questions, the score adds opacity rather than clarity.
Explainability should be designed for the decision and the user. A model developer may need feature distributions and validation results. An AML analyst needs understandable drivers tied to customer records, transactions, screening results, and policy. A reviewer needs to reconstruct what the analyst saw and why the outcome followed.
NIST’s AI Risk Management Framework distinguishes explainability—how a system produced an output—from interpretability—what the output means in its intended context. Both matter. “Transaction velocity contributed 18 points” describes mechanics; “four outward transfers occurred within two hours after an unusual incoming payment” gives operational meaning.
Use reason codes that lead to evidence
A reason code should be specific enough to investigate. “Behavioural risk” is too broad. “New-counterparty count exceeded the customer baseline” is better, especially when it links to the relevant counterparties, period, baseline, and threshold. The analyst should be able to move from driver to evidence without searching several systems.
Good explanations separate observed facts from interpretation. “Five transfers totalling €42,000 were sent within 90 minutes” is an observation. “Rapid dispersal inconsistent with expected activity” is an interpretation based on customer context and policy. Showing both helps the reviewer challenge a faulty baseline or supply a legitimate explanation.
Data quality belongs beside the driver. If 30 percent of counterparties lack a business category, say so. If device data was unavailable for the period, the interface should not imply a clean signal. Explanations are incomplete when uncertainty is hidden.
A simple illustrative calculation
Assume a fictional institution uses a 0–100 policy score with four components. Customer-profile risk has a maximum of 25 points, geographic exposure 20, observed behaviour 35, and screening or external context 20. For one synthetic SME, the assigned component values are 10, 5, 24, and 0. The total is 39.
The arithmetic is simply 10 + 5 + 24 + 0 = 39. The institution’s illustrative policy maps 0–29 to low, 30–59 to medium, and 60–100 to high, so the result falls in the medium band. The behavioural component is the primary driver because rapid new-counterparty payments differ from the stated profile.
These weights and bands are teaching assumptions, not a recommendation. A real institution must define factors and thresholds for its risks, data, controls, and legal obligations, then validate their performance. The score is not a 39 percent probability of crime. It is a policy output used to organise review.
The explanation should include component values, contributing events, missing data, calculation time, and policy version. If an analyst later confirms that a supplier migration explains the new counterparties, the record can preserve the original score and document the reviewed outcome.
Overrides need structure
Human review is not a decorative approval button. Analysts need authority to correct data, record legitimate explanations, and escalate when the model understates concern. At the same time, an undocumented override defeats consistency.
Require a reason category, a short rationale, and links to supporting evidence. Distinguish correcting a source record from accepting a known risk under policy. Capture who made the decision and when, and preserve the original output. Higher-impact exceptions may require a second review, depending on the institution’s control framework.
Override analysis is also feedback. If analysts repeatedly lower scores triggered by seasonal payroll, the model may need a calendar or customer-specific expectation. If they repeatedly elevate cases involving a signal the model ignores, the control design may have a gap. Feedback should prompt investigation; it should not be fed automatically into a model without governance.
Build a defensible decision record
A useful record connects input to outcome: customer snapshot, relevant transactions, external results, reason codes, missing information, model or policy version, analyst notes, requested evidence, and the final action. It should distinguish an alert from a case. An alert identifies an event or pattern for review; a case groups related alerts and evidence into an investigation.
Documentation should be proportionate. Copying every available field creates noise. Preserve the material evidence and the logic that moved the review forward. Jurisdiction-specific obligations still apply: for example, SAMA’s rulebook addresses documenting risk assessments and keeping records in a retrievable and auditable manner. Firms should interpret and implement applicable requirements with qualified legal and compliance advice.
The practical test is straightforward: can another trained reviewer understand what the system observed, what it did not know, what the analyst concluded, and why? If so, the score supports a defensible process. If not, more model complexity will not solve the documentation problem.
Primary sources
These sources support the regulatory and standards-related statements in this article. They do not endorse Merid or certify any product.
- AI Risks and TrustworthinessUS National Institute of Standards and Technology (NIST), AI RMF 1.0 resources
- Section 1: ML/TF Risk AssessmentSaudi Central Bank (SAMA) Rulebook
- Section 6: Record KeepingSaudi Central Bank (SAMA) Rulebook