Risk and compliance
The risk comes from the consequence of a decision, and the decision can be questioned long after it was made. An auditor or an examiner asks which rule made it, which version was in force, and what data it saw. You need an answer for every decision.
You have run this discipline for twenty years
Write a candidate rule, test it against history, run it beside the rule in force, compare, get it approved, promote it, and show an examiner every step. Model risk management has worked this way since long before anyone deployed an agent.
What has changed is that an agent now makes the decision, and supervisory guidance on model risk was written for models that score. The Federal Reserve's SR 26-2 states in footnote 3 that generative and agentic AI models "are not within the scope of this guidance". What that means for a bank.
We bring the practice your risk team already runs to the layer where the agent acts.
The part your framework cannot reach
A decision here runs on one of three things: a rule, a model's judgment, or a person. You already know how to govern the first and the third.
Model validation stays with your model risk team. Swiftward makes the model's output an input to a rule.
A classifier or a judge model produces a signal: a number, a label, a yes or no. A rule acts on it with a threshold, a combination, or a route to a person, and that rule is deterministic, versioned and frozen. The rule is what decided, and the record keeps the signal with the value it had, so "what did it see when it decided" always has an answer.
How much determinism you get is your choice, made per rule rather than per product.
| You send | The rule does | What you get |
|---|---|---|
| everything on the event, scores included | reads only what arrived | full determinism: the record holds every input the rule read |
| the event alone | fetches a score, a classifier, or a model through a function you write | the latest score from any source you can call; a second call may return a different answer, so the record, not a rerun, proves the decision |
| a mix, which is most systems | both, rule by rule | full determinism where you need it, and the latest data everywhere else |
That choice is a materiality judgment, the kind your risk function already makes about its own models, and here it is recorded. For a decision an examiner will read back to you, put every input on the event; a rule that decides whether a support reply sounds rude does not need that.
What you can test before promoting
| The change | How you validate it |
|---|---|
| A threshold on a model score | backtest the candidate against the recorded signals: same model outputs, new rule, every decision that flips |
| Two competing thresholds | champion and challenger on a deterministic A/B split, or the challenger in shadow, deciding nothing |
| Swapping the classifier itself | run it in shadow beside the classifier you use today and compare the decisions the two produce, not the scores |
| The classifier's own accuracy and drift | yours. That is model validation, and it belongs to your model risk function |
For that last row we give you the evidence: every signal value recorded beside the decision it produced, across your whole history, exportable.
An examiner asking "how do you govern the generative part" gets a straight answer: it is an input, and
- here is what it said on every request;
- here is the versioned rule that acted on it;
- here is what changed the last time we moved that rule.
What the discipline is made of
Controls. Business rules for thresholds, counters and external scores. Spend and loop limits. Authorization for the chain where every action is permitted on its own and the sequence is not. Registries and documentation for what exists and what governs it.
System. Safe change — the backtest against your own recorded history is the center of this discipline. Human review for approvals and referrals. Audit and evidence for the record that has to survive being questioned.