Back to Blogs

Model Risk Management for Generative AI: How Banks Can Extend SR 11-7 to LLMs

Transparent AI Enterprises

Date

September 3, 2026

Type

Share

Every regulated bank already has a model risk management function. It was built around SR 11-7, the supervisory guidance that has governed how institutions validate and monitor models since 2011, long before a large language model could draft a credit memo. 

Now generative AI is landing in those same workflows, and the model risk team is being asked a question the original guidance never anticipated: how do you validate a model that is non-deterministic, has no fixed specification, and was trained by someone else? 

This article maps the principles of SR 11-7 onto generative AI, shows where classic model risk management breaks, and lays out the validation evidence a regulator or an internal audit committee will expect before an LLM touches a regulated decision.

Key Insights

This article is written for the people who own model risk in a regulated institution: heads of model risk, model validators, chief risk officers, and the compliance and audit teams who have to sign off before generative AI reaches a customer or a regulator. By the end you will have a clear read on whether SR 11-7 covers your AI, where the guidance strains against generative systems, and a checklist for validating an LLM that holds up to effective challenge.

What Is Model Risk Management, and What Does SR 11-7 Require?

Model risk management is the set of policies, controls, and validation practices an institution uses to limit the losses that arise when a model produces wrong outputs or is used incorrectly. SR 11-7 is the supervisory guidance that defines the expectation in US banking. It was issued jointly by the Board of Governors of the Federal Reserve System and the Office of the Comptroller of the Currency in April 2011 (the OCC published it as Bulletin 2011-12), and it remains the reference standard supervisors apply.

SR 11-7 defines a model as a quantitative method that applies statistical, economic, financial, or mathematical techniques to turn input data into a quantitative estimate. It rests on three pillars: robust model development, implementation, and use; an effective validation framework; and sound governance, including a model inventory and clear ownership. The phrase that carries the most weight in practice is an effective challenge: a credible, independent review with the competence and standing to find a model’s flaws and the authority to force changes.

Read that definition again with generative AI in mind. An LLM that drafts a regulatory filing, scores a loan application, or summarizes a contract is turning input data into an output that informs a decision. The label on the technology is new. The supervisory logic that captures it is not.

Does SR 11-7 Apply to Generative AI?

Yes, when the output informs a decision. Supervisors have been consistent that model risk principles are technology-neutral: what matters is the use, not the math. If a generative model contributes to a credit decision, a capital calculation, a fraud review, or customer-facing guidance, it falls within the model risk perimeter and inherits the SR 11-7 expectations. The harder question is not whether the guidance applies, but how to satisfy it when the model behaves nothing like a logistic regression.

SR 11-7 expectationClassic statistical modelGenerative AI reality
A documented model specification to validateFixed equations and variablesA prompt, a retrieval pipeline, and a foundation model, often third-party, with no closed-form spec
Effective, independent challengeRe-derive and benchmark the mathRequires testing non-deterministic outputs and tracing each one back to its sources
Ongoing monitoring against a benchmarkTrack performance drift over timeMonitor hallucination and faithfulness, plus prompt and model-version changes that silently alter behavior
A complete model inventoryOne registered model per useMust inventory prompts, base models, agents, and data sources as connected components

Where Generative AI Breaks Classic Model Risk Management

Three properties of generative systems collide with assumptions SR 11-7 takes for granted.

  1. Outputs are non-deterministic. The same prompt can produce different answers. Validation built on reproducing a fixed output has to shift to testing distributions of behavior across representative cases, which is a different statistical exercise than re-deriving an equation.
  2. There is no fixed specification. A traditional model is its equations. A generative system is a prompt, a retrieval layer, a base model, and sometimes a chain of agents, any of which can change. A vendor updating a foundation model behind an API can alter your model’s behavior with no change on your side, which is a validation and change-management problem at once.
  3. The core model is usually someone else’s. Most institutions consume foundation models they did not train and cannot fully inspect. SR 11-7 already addresses vendor models and asks for the same validation rigor applied to internal ones, but a closed foundation model limits how deeply a validator can look, which raises the importance of output-level testing and traceability.

On top of all three sits AI hallucination. Stanford RegLab research measured general-purpose models hallucinating on 69 to 88 percent of legal queries, and OpenAI’s own April 2025 system card reported its o3 reasoning model fabricating on 33 percent of one factual benchmark, higher than the model before it. For a validator, that is not a quality footnote. It is the central risk to be measured and controlled.

A Validation Checklist for LLMs Under SR 11-7

Validating a generative model is the same discipline applied to new material. Work the three pillars, adapted for LLMs.

The Real Bottleneck: Effective Challenge Needs Traceability

Of all the SR 11-7 expectations, effective challenge is the one generative AI strains hardest. An independent validator cannot meaningfully challenge an output that arrives with no indication of where it came from. If the model says a covenant threshold is 3.0 times EBITDA, the validator needs to see the document and passage that produced that figure, not re-read the credit file to check. Without traceability, effective challenge degrades into spot-checking, which is exactly the control weakness supervisors look for.

This is the gap Seekr was built to close. 

SeekrFlow traces every output back to the specific source documents and training data that shaped it, scores how much each source influenced the answer, and logs full execution traces across agent workflows, which gives a validator the evidence that effective challenge requires. 

SeekrGuard adds model evaluation and certification-readiness before deployment, so a model enters production with a documented, defensible risk profile rather than an unmeasured one. The result is an LLM that fits the SR 11-7 process instead of fighting it: a model you can validate, challenge, monitor, and explain to a supervisor with evidence rather than assurances.

An Honest Limitation

SR 11-7 was written in 2011 and does not mention generative AI, so applying it requires judgment, and supervisors are still developing detailed expectations for these systems. Mapping the guidance to LLMs is an act of reasonable interpretation, not a literal reading, and institutions should expect the specifics to keep evolving through examiner feedback and newer guidance. Traceability and validation reduce model risk and make it defensible, but they do not eliminate it; generative models remain probabilistic, and the highest-stakes decisions still need a human accountable for the call. The point is not a risk-free LLM. It is a model risk you can measure, challenge, and stand behind.

Summary. Model risk management under SR 11-7 applies to generative AI whenever an LLM informs a decision. The guidance strains because outputs are non-deterministic, there is no fixed specification, and the base model is usually third-party, with hallucination as the central risk. A workable program validates the whole system on representative cases, measures a monitored hallucination rate, and, above all, gives independent validators the source-level traceability that effective challenge depends on.

Frequently Asked Questions

What is model risk management?

Model risk management is the discipline of identifying, measuring, and controlling the losses that occur when a model produces incorrect outputs or is used inappropriately. In banking it covers model development, independent validation, and ongoing governance, and in the United States it is governed by the supervisory guidance known as SR 11-7.

What is SR 11-7?

SR 11-7 is supervisory guidance on model risk management issued jointly by the Federal Reserve and the OCC in 2011 (OCC Bulletin 2011-12). It sets expectations for model development and use, independent validation with effective challenge, and governance including a model inventory, and it is the reference standard US bank examiners apply to model risk.

Does SR 11-7 apply to generative AI and LLMs?

SR 11-7 applies to generative AI whenever the model informs a business decision, because the guidance is technology-neutral and defined by use rather than technique. An LLM that scores credit, drafts regulatory content, or summarizes documents for a decision falls within the model risk perimeter and inherits the same validation and monitoring expectations as a traditional model.

What is effective challenge in model risk management?

Effective challenge is the SR 11-7 requirement for credible, independent review of a model by people with the competence, standing, and authority to find its flaws and force changes. For generative AI, effective challenge depends on traceability, because a validator cannot challenge an output without seeing the sources and reasoning that produced it.

How do you validate a large language model?

You validate an LLM by treating the prompt, retrieval, foundation model, and agent steps as one model, testing it on a representative set of real in-domain cases, measuring accuracy, faithfulness, and a hallucination rate, and capturing the sources behind each output so an independent validator can verify it. Validation is ongoing, with re-checks whenever prompts, retrieval, or the base model change.

Why is generative AI harder to validate than traditional models?

Generative AI is harder to validate because its outputs are non-deterministic, it has no fixed equation to inspect, and the core model is usually a third-party system that can change without notice. These properties break validation methods designed to reproduce a single fixed output, so the work shifts toward testing behavior across cases and tracing outputs to sources.

How does model risk management relate to the NIST AI RMF and the EU AI Act?

SR 11-7, the NIST AI Risk Management Framework, and the EU AI Act converge on the same controls: governance, validation, monitoring, and traceability. SR 11-7 is US banking supervisory guidance, the NIST AI RMF is a voluntary framework, and the EU AI Act is binding law with high-risk obligations applying from December 2027 after the 2026 Digital Omnibus deferral. Evidence built for one, especially validation and traceability, supports the others.

What is cost per defensible output (CPDO)?

Cost per defensible output is total AI spend divided by the number of outputs that survive verification. It turns model accuracy into a unit economic a risk or finance team can track over time, and it tends to fall when source traceability makes validation and effective challenge faster.

How much of your AI is trusted?

Before the next validation cycle, estimate what unverifiable AI output is already costing.

Calculate your Cost Per Defensible Output in just a couple minutes.

8-Content CTA BG-1440×642@2x