Governance Row
Insights
MAS AI Governance

How to Assess the Risk of an AI Use Case

Last updated: /6 min read

Ask a compliance officer how to assess AI risk under Singapore's incoming rules and the answers vary. Some describe four dimensions. Some reach for reputational impact, or criticality, or a firm-wide risk score. The proposed AI Risk Management Guidelines are more precise than that, and the precision matters, because the assessment is the mechanism the entire framework depends on.

To assess an AI use case's risk under MAS's proposed Guidelines, rate it on three dimensions: impact, complexity and reliance (paragraph 3.10). Each use case is scored both before controls are applied (inherent) and after (residual) under paragraph 3.9, the rating is recorded with its rationale, and it is reviewed periodically and reassessed when something significant changes. That rating then decides how much governance, control and oversight the use case receives.

MAS's proposed Guidelines rate each AI use case on three dimensions: impact, complexity and reliance. Not four, not a single institutional score. This is how the regulator expects a licensed firm to decide how much governance each use of AI actually needs.

Why the rating is the hinge

MAS could have asked every firm for one rating, or sorted AI into fixed tiers. It did neither. The Guidelines assess materiality per use case, set out across paragraphs 3.8 to 3.11 of the proposed text, and that choice runs through everything that follows. The rating a use case receives decides the depth of controls applied to it, the rigour of its validation, and what reaches the board. Rate a use too low and the controls that follow are too thin for the risk it carries. Rate it without recording why, and the number cannot be defended when a supervisor asks.

The rating is not a filing exercise. It is the decision that allocates a firm's governance effort to where the risk actually sits.

The three dimensions of AI risk materiality

Impact asks what happens if the AI is wrong. A model that flags transactions for human review carries different impact from one that declines a customer outright. The question is the consequence of error: customers affected, money moved, obligations breached, decisions made that are hard to reverse.

Complexity asks how hard the AI is to understand, validate and explain. A rules-based tool whose logic a reviewer can read end to end is low on this axis. A large language model, or a system whose behaviour shifts with its inputs, is high. Complexity is what makes a use case difficult to test and difficult to account for, independent of how much damage it could do.

Reliance asks how much the firm leans on the output, and how much a human still decides. An analyst who treats an AI summary as one input among several relies on it lightly. A workflow that acts on the AI's output with no human in the loop relies on it heavily. Reliance captures how much real control the firm has surrendered to the system.

The three are assessed together. A use case can be high on one axis and low on another, and the combination is what produces its rating. A tool with severe impact but heavy human oversight is a different risk from the same tool running autonomously. The proposed text names these three dimensions at paragraph 3.10, as the axes a materiality assessment should minimally cover.

The three dimensions a risk materiality assessment must minimally cover (paragraph 3.10).
DimensionThe question it asksParagraph 3.10 wording
ImpactWhat happens if the AI is wrong?The potential consequences of a failure, malfunction or poor performance, for the firm and its customers.
ComplexityHow hard is the AI to understand, validate and explain?Arising from the nature of the AI technology, the novelty of its application, or the data it uses.
RelianceHow much does the firm lean on the output, and how much does a human still decide?The level of autonomy, the degree of human involvement or oversight, and the availability of alternatives.

Inherent, then residual

A use case is first rated on its inherent risk: what it carries before any controls are applied. Controls are then put in place, and what remains is the residual risk. Paragraph 3.9 of the proposed text asks for both. The inherent rating shows why a use case was treated seriously; the residual shows what the firm did about it and what exposure it accepts. A rating that records only the final number, with no trace of the reasoning between the two, is a number without a defence.

Why per use case, not per firm

This is the point most firms get wrong, and it is worth stating plainly. The same institution can run an internal meeting-note summariser and a customer-facing autonomous workflow at the same time. Rated at the firm level, one of those disappears into the average. Rated per use case, each carries the governance its own risk demands: the summariser stays light, the autonomous workflow is controlled deeply.

Materiality follows the use, not the institution. A firm does not have a risk rating. Its individual uses of AI do, and the firm's job is to hold all of them at once, each at the right depth.

A rating is not a one-time act

An AI use case does not hold still. A tool adopted cautiously as a pilot becomes embedded in daily work; an analyst's occasional reference becomes the desk's default. As reliance deepens, the rating that was accurate at launch is no longer accurate now.

The Guidelines expect assessments to be kept current: reviewed periodically, as paragraph 3.9 asks, and reassessed when something significant changes, which paragraphs 4.24 and 4.25 treat as its own trigger. A use case rated once at go-live and never revisited is not a governed use case, however careful that first assessment was. The record has to move as the use moves.

What a defensible assessment looks like

The test underneath all of this is simple: can the firm show it. A rating that exists as a number in a spreadsheet, with no recorded rationale, no history, and no evidence of the controls behind it, answers the question badly. A rating scored on the three dimensions, with the reasoning recorded at the time, the inherent-to-residual path visible, and the reassessments dated, answers it well.

The difference is not the number. It is whether the number can be defended the day a supervisor asks how it was reached. That record cannot be assembled retroactively, which is why the assessment is worth doing properly from the first use case, not reconstructed under inspection.

Governance Row runs this as a live engine: every AI use rated on impact, complexity and reliance, the rationale recorded, the history kept, and controls proportionate to each rating attached with their evidence. From register to regulator, the assessment stays current and stays producible.

For the inventory the assessment sits on, see our guide to the eleven fields a working AI inventory needs. For how the assessment fits the wider framework, see our map of the full Singapore AI governance stack.

Frequently asked questions

What are the three dimensions of AI risk materiality under the MAS Guidelines?
Impact, complexity and reliance. Paragraph 3.10 of the proposed Guidelines lists these three as the dimensions a risk materiality assessment should minimally cover: impact is the consequence if the AI is wrong, complexity is how hard it is to understand and validate, and reliance is how much the firm leans on its output. Firms may add further dimensions, but these three are the floor.
Is AI materiality assessed per firm or per use case?
Per use case. Paragraphs 3.8 to 3.11 set out an assessment performed for each AI use case, system or model rather than a single firm-wide rating, so the same institution can hold a lightly governed internal tool and a deeply controlled customer-facing workflow at once. A separate, aggregate view of a firm's overall AI risk exposure exists, but it does not replace the per-use-case rating.
What is the difference between inherent and residual AI risk?
Inherent risk materiality is what a use case carries before controls are applied; residual is what remains after them. Paragraph 3.9 asks firms to take both into account and to ensure the residual materiality sits within the firm's risk appetite before deployment. Recording only the final number, with no trace of the reasoning between the two, leaves a rating that cannot be defended.
How often must an AI risk assessment be updated?
There is no fixed interval, but assessments are not a one-time act. Paragraph 3.9 requires the methodology and each use case's assessment to be regularly reviewed, and paragraphs 4.24 and 4.25 treat significant changes to the AI or its operating environment as a trigger for re-assessment or re-validation. A rating set once at go-live and never revisited is not, in the Guidelines' terms, a governed use case.
Does a rule-based tool count as AI under the MAS Guidelines?
Not on its own. The proposed Guidelines scope AI to models and systems that learn or infer from inputs to generate outputs such as predictions, content, recommendations or decisions, and expressly exclude calculators or tools whose outputs rest solely on predefined programming logic or rules. A deterministic rules engine is out of scope, while a machine-learning model or a large language model is in.
What does a defensible AI risk assessment look like?
One the firm can show. A defensible assessment scores the use case on impact, complexity and reliance, records the rationale at the time, keeps the inherent-to-residual path visible, and dates each reassessment, so the firm can answer how a rating was reached the day a supervisor asks. A number in a spreadsheet with no recorded reasoning, history or evidence does not meet that bar.