LuxeDetect
A stepped column of machined dark bronze, seven tiers deep, standing on a black floor. A single unbroken band of blue cobalt light runs around it below the fifth tier and spills onto the floor; the rest of the column is dark.

Model Risk Management for AI Assistants: Who Tests the Answer?

Written by: Melat Tadesse

Published: September 2026

A model risk function can produce the validation file for every model it owns. It usually cannot produce a record of the sentence an AI assistant wrote to a client that morning.

Most of the published discussion stops at scope. Is the assistant a model? Does it belong in the inventory? Those questions have answers now, and they differ by jurisdiction.

What the published treatments stop short of is the consequence for the tier.

The decision that binds comes after them. OSFI Guideline E-23 expects an independent assessment of a model's conceptual soundness and performance. On our reading, that does not settle what performance means for the sentence a client reads.

Scope the assistant in and the obligation arrives without a method. Tier it down and the institution has tiered it down because the system resists assessment. That is the one justification proportionality does not offer, and it has a date on it. May 1, 2027.

Key Takeaways

  • Model risk management governs a population of models. Its unit is the model, not the answer.
  • The frameworks have split. SR 26-2 puts generative and agentic AI models out of scope. OSFI Guideline E-23 writes no exclusion and takes effect May 1, 2027.
  • Scope is the question everyone has answered. What binds is the assessment a scoped-in assistant owes.
  • Tiering a system down because its outputs resist assessment is the one justification the guideline does not offer.
  • A generated answer cannot be reproduced on demand. The check applied to it can be defined in advance, versioned and logged.

What Model Risk Management Is, and Whether It Covers AI

Model risk management is how a regulated institution governs the models behind its decisions. It identifies, inventories, tiers, validates and monitors them. The unit of governance is the model.

The unit that reaches the client is the answer. The distance between the two is this article.

Whether it covers AI depends on which supervisor you answer to.

In the United States it does not cover the generative kind. Footnote 3 to the SR 26-2 attachment states that "Generative AI and agentic AI models are novel and rapidly evolving. As such, they are not within the scope of this guidance." It adds the qualifier that matters: "the principles described in this guidance apply to traditional statistical and quantitative models and non-generative, non-agentic AI models."

In Canada, OSFI Guideline E-23 contains no such exclusion. It defines a model as "An application of theoretical, empirical, judgmental assumptions or statistical techniques, including AI/ML methods, which processes input data to generate results." No carve-out follows, for generative AI, agentic AI or large language models.

Read the absence carefully. The verified fact is that no exclusion was written, not what OSFI intended. Both instruments are supervisory guidance, not binding law.

Why the Split Matters Now

SR 26-2 was issued April 17, 2026, jointly by the Federal Reserve, the OCC and the FDIC, with the OCC publishing it as Bulletin 2026-13. It supersedes SR 11-7 and SR 21-8.

It is "expected to be most relevant to banking organizations with over $30 billion in total assets regulated by the Federal Reserve."

Where did the excluded systems go? To the institution itself.

The same footnote assigns them to the organization's own "risk management and governance practices." It is an assignment, not a gap.

The Federal Reserve's 2026 SR letter index, checked October 6, 2026, carries no AI-specific supervisory guidance and no amendment to SR 26-2. OCC and FDIC bulletins were not checked.

OSFI Guideline E-23 was published September 11, 2025. It takes effect May 1, 2027 and reaches every federally regulated financial institution. Proportionality turns on size, strategy, risk profile, "nature, scope, and complexity of operations," and interconnectedness.

It was written with AI/ML models in view, contemplating "models characterized by dynamic self-learning and autonomous decision-making." Its explainability expectations vary with a model's autonomy.

That asymmetry is well covered. Yields.io set out the divergence in June 2026. It is the setup, not the finding.

OSFI E-23 and SR 26-2, Side by Side

Are generative and agentic AI in scope?

E-23 (Canada): No exclusion is written. The model definition expressly includes AI/ML methods.

SR 26-2 (United States): Expressly out of scope.

What decides how much assessment the system owes?

E-23 (Canada): The inherent risk tier the institution assigns it.

SR 26-2 (United States): Not applicable. The system is out of scope.

What does independent assessment require?

E-23 (Canada): Independent of development, validating that models are properly specified, working as intended and fit for purpose. No exemption by model type.

SR 26-2 (United States): Validation that models perform as expected, including reliability and limitations.

What does the record contain?

E-23 (Canada): Evidence about the model. Nothing takes an individual answer as its unit of account.

SR 26-2 (United States): Evidence about the model. Excluded systems fall to the organization's own practices.

Where the Gap Sits in a Working Day

What follows is a composite scenario, drawn from patterns common across regulated financial institutions. It names no organization.

A relationship manager opens a client record and asks the assistant for a draft reply. It generates language grounded on retrieved record data, with platform controls running underneath: data masking, toxicity scoring, prompt defense. The manager skims the draft. The skim is discretionary.

Then it sends.

The platform logs the prompt, the response and its trust signals. It logs no verdict on whether the answer was right or on-brand, because there was no fixed reference to assess it against.

That is the gap. E-23 asks for an independent review that a model is properly specified, working as intended and fit for purpose. At the level of one customer-facing sentence there is nothing to review against.

The consequence does not stay small. A statement about a hardship program, a fee or a timeline reaches a client who acts on it, and remediation runs across every matching recipient.

Accountability here has already been tested. In Moffatt v. Air Canada, 2024 BCCRT 149, decided February 14, 2024, a British Columbia small claims tribunal held the airline responsible for its chatbot's wrong bereavement-fare answer.

The decision is persuasive, not binding precedent. The Air Canada chatbot case is covered in its own article.

The audit position is the awkward part. The function can produce a complete record of the model and none of the answer. A banking chatbot meets the same wall: a transcript records what was said, not whether it was right.

What Model Validation and the Platform Controls Cover, and Where Each Stops

Each of these controls is real, and each was built for a different question.

Model validation. It governs model performance. SR 26-2 describes validation as work that "evaluates whether models perform as expected and includes an assessment of a model's reliability and its limitations." It assumes a system that can be run again and compared, and it does not reach the customer-facing sentence.

The Einstein Trust Layer. It governs data handling and screens for toxicity, with zero data retention, dynamic grounding, prompt defense, data masking, toxicity scoring and an audit trail. It reduces real classes of risk.

The documentation describes no component that checks a statement against the institution's published position, approved tone or compliance language. Salesforce's Agentforce documentation assigns that work to the human reviewer, noting it is "important to make sure that LLM-generated responses intended for external audiences are accurate and helpful, and that they align with your company's values, voice, and tone."

The human skim at send. It is the most flexible control and the least recordable, and a supervisor cannot inspect it afterwards. The same gap appears in outbound marketing, where approval names a template, not the message that sends.

Then the control a senior model risk manager reaches for next. E-23 is principles-based and proportional. Tier the assistant low, document the explainability limits, and the obligation is discharged. That is the strongest objection to everything above.

Model Risk Tiering: The Decision That Actually Binds

Proportionality sets the depth of independent assessment. It does not decide whether an output was ever assessed.

E-23 expects institutions to have "a process to independently assess conceptual soundness and performance of models," carried out independently of model development. It expects the inventory to hold every model carrying "non-negligible inherent model risk," rated by tier, with no exemption by model type.

A system writing to clients at volume is not self-evidently negligible, so the assistant goes in the inventory and the tier decides how much assessment it owes.

Here the method runs out. On our reading, E-23 does not define what performance means for a single customer-facing answer.

Conceptual soundness can be argued for the model and performance described in aggregate. Neither tells a reviewer whether Tuesday's sentence matched the institution's position.

Work is under way. Published proposals cover behavioral and semantic validation, prompt-variance testing, human-alignment proxies and second-model judges, and some are already in products.

The narrow claim is not that a generated answer cannot be assessed, only that no method is yet settled or supervisorily recognized in the way E-23 asks.

So the tier becomes the decision. Rate the assistant high and it owes an assessment with no settled method. Rate it low and the justification has to be written down.

So read one. No head of model risk writes that the system is difficult to assess. They write that the assistant is low inherent risk because a relationship manager reviews every draft before it sends.

That justification has a problem before the evidence question arrives. E-23 rates models on inherent risk, and says it does not expect residual model risk to drive the primary governance and oversight of models. A reviewer at the end of the process is a mitigant, on the residual side of that line.

It is also the one control here that leaves no record. The skim is discretionary and unlogged.

And the formal basis was never there. The guideline gives five bases for proportionality, and difficulty of assessment is not among them.

There is a third option: assess the output rather than the model. A generated answer cannot be reproduced on demand, and validation as both frameworks describe it assumes it can be. The check applied to that answer can be defined in advance, versioned, applied to every output and logged.

The May 2027 Tiering Review: Five Questions

  • Are the customer-facing AI assistants inside the model inventory, under each framework the institution answers to?
  • If they are in scope under E-23, which independent assessment method applies?
  • If an assistant is tiered low, what does the justification say, and would it survive a supervisor reading it back?
  • What evidence exists at the level of an individual answer, rather than the model?
  • Is there a documented standard for what "correct and on-brand" means for generated text, and does anyone own it?

Where LuxeDetect™ Sits

The enterprise stack still lacks a consistent, independent control layer that verifies AI-generated language against the brand's own standard before release. LuxeDetect™ is AI Brand Integrity Infrastructure: the independent, deterministic brand-integrity release gate across the enterprise AI stack.

LuxeDetect™ independently evaluates every AI-generated output before it goes live, scoring alignment with brand tone, style, and standards against a benchmark the brand owns. It detects misalignment, flags inconsistencies, and intercepts brand-damaging content before it reaches the public.

Each in-scope AI-generated or AI-assisted customer-facing output is evaluated against a brand-owned, version-controlled LF1000 Brand Benchmark and either approved, routed for review, or intercepted before release. Alignment is measured with the LuxeFactor™ methodology.

The standard does not belong to the model, the platform or LuxeDetect™. It belongs to the brand. LuxeDetect™ turns that standard into something enforceable.

The agencies gave a different reason for the exclusion. They called these models novel and rapidly evolving. We read the obstacle underneath as the same one: a generated answer resists reproduction, and validation assumes it.

The boundary matters. LuxeDetect™ does not make an institution compliant with E-23 or SR 26-2, and it does not perform model validation.

Frequently Asked Questions

What Is Model Risk Management?

Model risk management is the discipline governing the models behind a financial institution's decisions, from credit and pricing to capital. OSFI Guideline E-23 sets it out in Canada and SR 26-2 in the United States. Its unit is the model, which is why a customer-facing assistant sits awkwardly inside it.

Does Model Risk Management Cover AI?

It depends on the supervisor. SR 26-2 puts generative and agentic AI models outside its scope. E-23 writes no exclusion, and its model definition expressly includes AI/ML methods. An institution answering to both can find one assistant in scope in Canada and out of scope in the United States.

What Is OSFI E-23?

OSFI Guideline E-23 is Canada's model risk management guideline. It was published September 11, 2025, takes effect May 1, 2027 and applies to every federally regulated financial institution. It is guidance rather than binding law.

Is an AI Assistant a Model Under OSFI E-23?

If it meets the definition, yes. E-23's definition expressly includes AI/ML methods, with no exclusion for generative or agentic systems. An assistant with non-negligible inherent model risk belongs in the inventory, rated by tier.

Does the Einstein Trust Layer Close This Gap?

No, and it is not built to. It governs data handling and screens for toxicity inside Salesforce. Salesforce's own documentation assigns the question of values, voice and tone to the human reviewer, the same control a low-tier justification leans on.

How Do You Validate a Generative AI Model?

Not the way either framework describes, since both assume a system that can be run again and compared. Proposals include behavioral and semantic validation, prompt-variance testing, human-alignment proxies and second-model judges. None is yet settled or supervisorily recognized, which is why the obligation lands on the tiering decision.

How Is AI Content Risk Different from Model Risk?

Model risk asks whether a model's calculation is accurate and stable over time. AI content risk asks whether the language a customer sees is correct, on-brand and compliant. A model can pass every validation test and its output can still misstate the institution's own policy.

What Your AI-Generated Communications Would Have to Show

Between now and May 1, 2027 the question is not which supervisor is right. It is what the institution writes down when it tiers its assistant, and whether it can show what an AI-written sentence was checked against before a client read it.

In a LuxeDetect™-governed workflow, no in-scope output reaches the public without evaluation.

See how a release gate evaluates an AI answer before it sends.

Request early access to LuxeDetect™ and test each answer your AI assistant gives a customer.

AI content risk is not solved after publication; it has to be controlled before release.

Melat Tadesse

Head of Content

Melat Tadesse leads content at Luxe Factor Intelligence Corp. (LFIC), where she helps define the standards that protect brand voice in the AI era. Her work focuses on the failure modes of AI-generated content, how voice gets diluted at scale, and what it takes to turn brand identity into enforceable rules. She brings a sharp editorial lens shaped by years of studying high-standard brands and the systems behind what makes them unmistakable.