Back to Blogs
How to Evaluate an AI Governance Platform for a Bank When SR 26-2 Leaves Generative AI Out of Scope
If the checklist on your desk is based on SR 11-7, it is based on guidance that has been withdrawn. On 17 April 2026 the Federal Reserve, the FDIC and the OCC issued revised Supervisory Guidance on Model Risk Management, published by the Federal Reserve as SR 26-2 and by the OCC as Bulletin 2026-13. In the agencies’ own words it supersedes and replaces SR 11-7 from 2011 and the 2021 interagency statement covering BSA and AML systems.
It also carries a footnote the 2011 letter never could. Generative AI and agentic AI models sit outside its scope, while its principles do apply to traditional statistical and quantitative models and to AI models that are neither generative nor agentic. The same passage says a bank’s own risk management and governance practices should guide the determination of appropriate governance and controls for tools the document does not cover.
So if you are buying for large language models and agents, this is not a compliance exercise, and an offer of compliance with SR 26-2 for those models is an offer that guidance does not support. It is a checklist for filling the gap the agencies handed back to you, built from the model risk principles worth borrowing for the job. Ten weighted criteria follow, so you can evaluate an AI governance platform on evidence rather than on claims.
What SR 26-2 changed
Beyond SR 11-7 and SR 21-8, the OCC’s version also rescinds the Model Risk Management booklet of the Comptroller’s Handbook.
Expectations now scale: the guidance is expected to be most relevant to banking organizations with over 30 billion dollars in total assets, and below that line the agencies point to internal practices appropriate for size and risk profile, while noting it may still be relevant where model risk exposure is significant.
Two further changes matter when you evaluate an AI governance platform.
The definition of a model narrowed
A model is now a complex quantitative method, system or approach applying statistical, economic or financial theories to turn input data into quantitative estimates. The text expressly excludes simple arithmetic such as spreadsheet calculations, and deterministic rule-based processes and software with no such theory underpinning their design or use. The input, processing and reporting framing practitioners knew from SR 11-7 is not in the new text.
It sets no enforceable standards
The guidance says so directly: it does not set out enforceable standards or prescriptive requirements, and non-compliance with it will not by itself draw supervisory criticism. Supervisory action can still follow from violations of law or from unsafe and unsound practices resulting from inadequate model risk management. Weigh any vendor claim of compliance against that.
The OCC bulletin also says the agencies plan to issue a request for information covering model risk management and banks’ use of AI, including generative and agentic systems. That gathers views rather than setting rules, so it would not by itself change the scope of SR 26-2.
The generative AI carve-out, and what it puts back on you
The scope footnote is short and consequential. In the agencies’ words, “Generative AI and agentic AI models are novel and rapidly evolving.” On that basis they are placed outside the guidance, with responsibility for governing uncovered tools handed to the bank’s own practices.
That sorts your estate into three buckets. Traditional statistical and quantitative models are covered, and are what an existing model risk framework is already shaped around. AI models that are neither generative nor agentic are covered too, provided they also meet the narrower definition of a model, which is how a gradient-boosted fraud or credit model normally lands inside the guidance. Generative and agentic AI models sit outside it.
The instinct is to read the carve-out as a reprieve. For a buyer it is better read as a transfer of work. Where guidance applies, you can assess adequacy against its text. Where no guidance applies, you must assess it against the standards your institution has set, in front of reviewers who did not set them. Only evidence makes that argument hold, which is why an AI governance platform earns or loses its place on the evidence criteria rather than on its dashboards.
Section VII, and why generative systems strain it
What the section treats as sound practice
Section VII of SR 26-2 covers vendor and other third-party products, and now stands alone rather than sitting inside the validation section. It acknowledges that proprietary components may mean you never see the underlying code, data or methodology, and holds that the principles of model risk management apply anyway.
The sound practices it describes: validating vendor products, by internal or outside parties; understanding the vendor model’s conceptual soundness, design, development data and performance; ongoing monitoring and outcome analysis of whether it stays accurate, fit for purpose and reliable.
Customizations are documented, justified and evaluated as part of validation. These are sound practices rather than requirements, so your evaluation framework must determine how much weight to give them. And if the product is a generative or agentic system, Section VII does not govern it either.
What it no longer gives you
The 2011 letter expected banks to maintain contingency plans in case a vendor model became unavailable. That expectation is not in the 2026 text, and neither is anything on change notification or on the supply chain behind your vendor.
Those protections now come from the June 2023 interagency guidance on third-party relationships and from the contract you negotiate. If nobody owns that seam on your side, it stays open.
Four properties that make this hard
Behavior is not deterministic, so assessing whether a system performs as intended is statistical rather than a check against a specification.
The model is not only the weights: change the retrieval index, the system prompt, the temperature or the tool list and behavior changes, and those artifacts tend to move on an engineering cadence rather than a validation one.
A third party can change your model, because a foundation model provider updates or deprecates on its own schedule. And there is no closed-form specification, so conceptual soundness has to be argued from evaluation design.
The weighted evaluation checklist
Read the middle column as borrowed reasoning rather than citation. For generative and agentic systems none of it is a requirement. These are the model risk principles that transfer best to the problem.
The weights sit where validation actually fails during a review, on evidence, reproducibility and change control, rather than on dashboards and integrations. Adjust them before you see the demos, and record why. A weighted evaluation checklist reweighted after the finalists present is a preference dressed as a method.
| Criterion | Weight | Why it carries this weight | The question that exposes a weak vendor | What an adequate answer sounds like |
|---|---|---|---|---|
| Evaluation evidence you can reproduce | 18% | Conceptual soundness turns on developmental evidence. For a generative system that evidence is the evaluation, and one nobody can rerun is an opinion. | Show me a completed evaluation. Can my team rerun it next quarter without your staff in the room? | They open a stored run and show the dataset version, evaluator definitions, model version, sampling settings and raw per-item results, then export the lot. |
| Independent challenge without vendor mediation | 15% | Effective challenge means critical analysis by objective experts with the expertise, independence and standing to force a change. | Which parts of this evaluation can my second line run and interpret with no vendor involvement? | The platform is the instrument and nothing more: your reviewers write the criteria, run the tests and own the conclusions, with export for internal audit. |
| Testing on your data and your criteria | 13% | Outcome analysis for a vendor model asks whether it stays accurate, fit for purpose and reliable in your use. A vendor benchmark score is not evidence about your book. | Take our credit policy and our declined-application file. How long until that is an evaluation set, and who touches the data? | A concrete path from your own documents to custom evaluators and test data, with residency and access controls named rather than implied. |
| Inventory, versioning and change control | 12% | An effective inventory carries enough to understand model risk individually and in aggregate. For generative systems that means reaching prompts, retrieval sources and tool definitions. | A prompt changed on Tuesday. Where does that show up, what evidence is now stale, and who is told? | Every component versioned, changes attributable to a person, and superseded evidence flagged as no longer current rather than quietly retained. |
| Ongoing monitoring and benchmarking | 12% | Monitoring is described as evaluating whether a model still performs as expected as products, clients, data and conditions change, with benchmarking a practical route to conceptual soundness. | What tells us the model has drifted, as opposed to telling us traffic went up? | Quality metrics tied to the original criteria, with alternatives scored side by side so a challenger model is a fact rather than a hunch. |
| Attribution and traceability of outputs | 10% | Documentation supports continuity and the tracking of recommendations, responses and exceptions, all of which depend on reconstructing why an output happened. | For this answer, which source passages drove it, and can I show that to a reviewer a year from now? | Per-output attribution and confidence signals retained with the record, not regenerated by a model that has since changed. |
| Human review, override and escalation | 7% | Governance depends on clearly defined roles and accountability across the lifecycle, including who may intervene and on what authority. | Show me where a reviewer disagreed with the system and what happened next. | Review steps sit inside the workflow, overrides are logged with a reason, and disagreement rates reach the model owner. |
| Model supply chain transparency | 6% | SR 26-2 is silent here, so this is carried by third-party risk management and your contract. A model you cannot name is a dependency you cannot monitor. | Name every model, host and sub-processor in the path of a production request, and tell me how we learn when one changes. | They put the list in writing, commit to notification in the contract, and let you pin or approve model versions rather than inheriting updates. |
| Customization, contingency and evidence portability | 4% | Adjustments made to fit your needs are documented, justified and evaluated as part of validation. Continuity and exit are yours to negotiate. | If we tune this to our policy, what does the evidence trail look like, and what do we keep if we terminate? | Customizations are versioned and testable, and results, attribution records and audit logs can be exported in an open format, as required by the contract. |
| Deployment topology and data residency | 3% | Controls over confidential data and jurisdiction are a precondition to running any evaluation on real material, not a differentiator. | Can this run where our data has to stay, and does anything leave that boundary during evaluation? | A specific answer about on-premises, private cloud or sovereign deployment, with the exceptions named rather than glossed. |
Scoring it so the number means something
Score each criterion 0 to 3 against evidence shown, not claims made. Zero is no capability, one is a roadmap item or manual workaround, two is working with a named gap, three is demonstrated end to end on something resembling your data. A three requires an artifact you take away, not a screen you watched.
Set knockout floors where a low score cannot be compensated: scoring 0 or 1 on reproducible evidence, on independent challenge, or on change control should disqualify, however strong the total, because those three decide whether your file exists. Run the same written scenario across every vendor, sent before the demos. If each vendor picks the demo, you are comparing sales scripts.
Three questions that separate an AI governance platform from a dashboard
Does it tell you whether the system is right, not only whether it is running? Can it produce a defensible, reproducible comparison between two candidate models against criteria you wrote? Can it hand a reviewer evidence that stands on its own if the vendor is not reachable?
If the answers are uptime, latency and a login, you are looking at monitoring rather than governance. Put those three in the AI governance RFP alongside your scoring weights, and ask for one complete redacted evaluation package from another regulated customer.
Answers that should end the evaluation
“We are SR 11-7 compliant.” SR 11-7 was superseded in April 2026, and SR 26-2 states plainly that it sets no enforceable standards. There is nothing here for a product to be compliant with, and nobody that would confer it.
“We handle validation for you.” Validation can be performed by internal or outside parties, so an external validator is legitimate. Ask what the finding is worth when the party producing it is the party selling the model.
“We use the latest model automatically.” Convenient in a consumer product, harder to defend in a regulated one. Ask what happens to evidence gathered on the previous version, and whether you can pin or approve a version before it reaches production.
Where this sits in your model risk program
An AI governance platform is a control instrument, not a control. Whoever owns the model in your structure still owns it, and whoever challenges it still challenges it. SR 26-2 is deliberately neutral on how you draw those lines: it ties the quality of validation to the rigor and effectiveness of the review rather than to where reviewers sit, and notes that internal audit generally evaluates whether practices are effective rather than repeating the validation itself.
Tier your own bar by materiality, since materiality is the guidance’s organizing idea and transfers cleanly to systems it does not cover. Give the full weighted checklist and the knockout floors to a use case touching credit decisions, financial crime alerts or customer disclosures; an internal drafting assistant that touches none of those usually warrants a lighter pass. The judgment should be based on where the output goes, not on what the tool is called. If you run model risk management software, check whether it also produces the evaluation evidence a generative system needs, because that is the gap you are buying to close.
Put the checklist to work
SeekrGuard independently evaluates AI models and agents, so teams can test candidates on their own data and criteria, turn the results into a use-case-specific risk picture, and reach a deployment decision they can trust and defend, with results that are observable and reproducible. SeekrFlow brings model-agnostic explainability, using data attribution, confidence scoring and prompt comparisons to validate accuracy while keeping governance and control over deployed AI.
Book a demo and bring your own use case. We are glad to apply the checklist to your use case with you, whichever platform you ultimately choose.
Frequently Asked Questions
Is SR 11-7 still in effect?
No. On 17 April 2026 the Federal Reserve, the FDIC and the OCC issued revised Supervisory Guidance on Model Risk Management, published as SR 26-2 and OCC Bulletin 2026-13. It supersedes and replaces SR 11-7 from April 2011 and SR 21-8, the 2021 interagency statement covering BSA and AML systems, and the OCC’s version also rescinds the Model Risk Management booklet of the Comptroller’s Handbook. SR 26-2 is expected to be most relevant to banking organizations with over 30 billion dollars in total assets, and it states that it does not set forth enforceable standards or prescriptive requirements.
Does SR 26-2 cover generative AI?
No. The guidance describes generative AI and agentic AI models as novel and rapidly evolving and places them outside its scope. It adds that a bank’s own risk management and governance practices should guide the controls for tools it does not cover, and that its principles do apply to traditional statistical and quantitative models and to AI models that are neither generative nor agentic. In practice a gradient-boosted fraud model that meets the guidance’s definition of a model is in scope, while a retrieval-augmented assistant is not, and what counts as adequate for the second is worked out inside your institution.
What should an AI governance RFP include?
An AI governance RFP should ask for evidence rather than capabilities. Request the vendor’s view of which of your systems count as models and why. Request one complete redacted evaluation package from a regulated customer, including criteria, dataset description and raw per-item results. Request a written model supply chain with a change-notification commitment, a statement of which artifacts your own reviewers can produce without vendor involvement, and the export format for every record you would need after termination. Publish your scoring weights so vendors invest effort where your risk actually is.
Which part of SR 26-2 applies to vendor platforms?
Section VII, on vendor and other third-party products, which now stands alone rather than sitting inside the validation section. Its sound practices are validation by internal or outside parties, understanding the vendor model’s conceptual soundness, design, development data and performance, and ongoing monitoring and outcome analysis. Note the scope limit: if the vendor product is a generative or agentic system, Section VII does not govern it either, and those practices are principles worth borrowing rather than rules you are held to.
How should a smaller bank adjust the weights?
Start from the applicability statement, because scalability is the point of the revision. SR 26-2 is aimed primarily at organizations above 30 billion dollars in total assets, and below that it points to internal practices appropriate for their size and risk profile, while noting it may still be relevant where model risk exposure is significant. Keep the knockout floors and simplify the rest. If you decide an independent challenge is warranted for a use case, the useful question is what makes it affordable rather than what lets you drop it: favor an AI governance platform whose stored evaluations your own reviewers can rerun and read without specialist tooling.