Back to Blogs
Build or Buy an AI Governance Layer? A Decision Framework for Enterprise Platform Teams
Should you build AI governance or buy an AI governance platform? Build the parts whose requirements your own business writes: policy thresholds, evaluation cases, and the integrations into your systems. Consider buying the parts whose requirements regulators, standards bodies, and auditors write and revise: the evidence trail, control mapping, and reporting. Treat the decision as a boundary to draw, component by component. Our advice: plan as if accountability stays with you on both sides of it.
If your engineers can build the layer, treat the build vs buy AI governance question as one about ownership. Which parts encode judgments only your business can make? And which parts follow a specification written outside your company?
Key insights
- Choosing between an in-house build and an AI governance platform? Build the parts whose requirements you write, consider buying the parts whose requirements someone outside your company writes and revises, and plan for accountability to stay with your organization either way.
- In this guide we use our own working breakdown of an AI governance layer, not a standard: inventory, policy, evaluation and testing, runtime controls, evidence and audit trail, and reporting. Decide build or buy for each of the six separately.
- On April 17, 2026, the Federal Reserve issued SR 26-2, which “supersedes and replaces” SR 11-7. The attached interagency text says generative AI and agentic AI models “are not within the scope of this guidance,” and the agencies plan a request for information on banks’ use of AI.
- The European Commission now lists 2 December 2027 as the start date for high-risk rules in Annex III areas, after an amending regulation entered into force on 27 July 2026.
- Buying can move work to a vendor. Plan to stay responsible for the outcome: US banking guidance says model risk principles “remain applicable” to vendor products.
What Does an AI Governance Layer Consist Of?
Before you price either option, name the parts. Compare a build estimate and a vendor quote only after both cover the same scope. The NIST AI Risk Management Framework organizes AI risk work into four functions: govern, map, measure, and manage. It calls govern “a cross-cutting function” that runs through the other three. In this piece, a governance layer means the software and process you use to carry those functions out for the models and agents you run.
The table below is that working breakdown. Read the right-hand column as the first draft of your decision.
| Component | What we suggest it produce | Who writes the specification |
| Inventory | A current registry of the models, agents, and use cases you run, with version, owner, and business purpose | You |
| Policy | The thresholds, prohibited uses, and escalation paths your business is willing to defend | You |
| Evaluation and testing | Repeatable tests with recorded thresholds and pass or fail evidence per version | Shared |
| Runtime controls | Blocking, flagging, and monitoring of live outputs against stated policy, inside your latency budget | Shared |
| Evidence and audit trail | A record, with attribution for each output, that a third party can read after the builders have moved on | Outside |
| Reporting | Audit-ready evidence, arranged by the control framework a reviewer names | Outside |
The rows marked “You” are yours whichever way this goes. Look hardest at the rows marked “Outside.” Outside bodies revise those requirements, as the timeline further down shows.
Inventory
Start with the list. US banking supervisors call it “common industry practice” to keep “a comprehensive set of information for models under development or in use.” Federal agencies have a parallel duty: OMB Memorandum M-25-21 tells them to keep updating their “annual AI use case inventory.”
Plan for upkeep as well as the first build of the registry. Ask how your inventory would find a system that a team adopted without telling anyone (shadow AI detection). Does a vendor’s inventory discover models across your clouds? Would a built one cover your internal platforms better?
Policy
In this breakdown, policy is the content of governance: what is allowed, at what threshold, and who gets called when a line is crossed. Write it yourself, since it records what your business will defend. Then ask where a product would keep it: in a format you control, or inside its configuration screens?
Evaluation and Testing
Evaluation, as we use the word, turns a policy into a test a model can pass or fail. We suggest you own the test cases, because they should reflect how your business could lose money or harm a customer. Scope the runner that records a reproducible result per version as an engineering task.
Then budget for upkeep. Ask who re-checks your evaluators when your models and prompts change, and how often. Our guide on how to evaluate a fine-tuned LLM covers the metrics side.
Runtime Controls
We use runtime controls to mean the checks in the request path that can block or flag an output and feed monitoring. Ask three things before you own this component. How much of your latency budget may each check spend? What should happen when a guardrail fails: block traffic, or let it through? And who gets paged? If you run agents, also ask what a control can do before a multi-step agent acts. We cover that case in how to audit multi-step AI agents.
Evidence and Audit Trail
Build this component for an outside reader. The European Commission’s summary of the AI Act lists “logging of activity to ensure traceability of results” among the obligations that high-risk systems will be subject to starting on 2 December 2027. Design your audit trail to answer a specific question about one output on a past date: which model version, which policy, which sources, which approval. Put attribution here too, because a reviewer may ask why the system said what it said. Our explainable AI guide goes deeper.
Then test it with someone who did not build the system and has no reason to trust it. Would they accept the record without your team in the room?
Reporting
Reporting, in this breakdown, arranges the evidence by whichever control framework your reviewer names. One auditor may want it mapped to ISO/IEC 42001, which ISO describes as specifying requirements for “establishing, implementing, maintaining, and continually improving” an AI management system. Another may ask for the NIST functions. A bank examiner may want model risk language. Ask whether one evidence store can feed all three layouts. Plan to update each layout when its framework changes.
The Rules an In-House Build Has to Track: Nine Changes Since April 2025
Before you price an in-house build, list the documents your specification depends on. Then count how often they changed since April 2025. Who on your team noticed?
US Banking: SR 26-2 Supersedes SR 11-7
On April 17, 2026, the Federal Reserve, the OCC, and the FDIC issued revised model risk management guidance. The Federal Reserve’s SR 26-2 says the new text “supersedes and replaces” SR 11-7, its 2011 letter. Separately, OCC Bulletin 2026-13 rescinds the OCC’s own 2011 bulletin.
Three details matter here. First, the agencies set out “a risk-based approach” and say the text “does not set forth enforceable standards or prescriptive requirements.” Second, they say generative AI and agentic AI models “are not within the scope of this guidance.” If your bank uses tools the text does not cover, a footnote says your own risk management and governance practices “should guide the determination of appropriate governance and controls” for them. Third, the agencies “plan to issue in the near future a request for information” on banks’ use of AI.
So if your bank mapped its controls to SR 11-7 sections, re-map them against the new text. We wrote about what that means for a purchase in how to evaluate an AI governance platform for a bank.
Europe: The High-Risk Dates Moved
The European Commission’s AI Act page reports that an amending regulation entered into force on 27 July 2026. Under it, rules for high-risk systems in the Annex III areas “will apply from 2 December 2027.” For AI built into regulated products under Annex I, the Commission gives 2 August 2028. Earlier in the same window, obligations for general-purpose AI models became applicable on 2 August 2025.
If you built toward an earlier high-risk date, check which of your internal deadlines should move.
US Federal and State Rules: Replaced, Delayed, Replaced Again
On April 3, 2025, OMB issued M-25-21, which “rescinds and replaces” M-24-10 on agency use of AI. The same day it issued M-25-22, which does the same to M-24-18 on AI acquisition.
Colorado changed its AI law twice in that window. On August 28, 2025, the governor signed SB25B-004, which extended the effective date of the state’s 2024 AI law. On May 14, 2026, the governor signed SB26-189, which repeals those provisions and reenacts them with new requirements for automated decision-making technology. If you wrote controls against the first law, give them a second reading.
Standards: Two More Dates
NIST describes its AI RMF as “intended for voluntary use.” On April 7, 2026, NIST released a concept note for a new profile on trustworthy AI in critical infrastructure. ISO published ISO/IEC 42005, its guidance on AI system impact assessments, in May 2025.
| Date | What changed | Rule-maker |
| April 3, 2025 | M-25-21 rescinds and replaces M-24-10 | OMB |
| April 3, 2025 | M-25-22 rescinds and replaces M-24-18 | OMB |
| May 2025 | ISO/IEC 42005 on AI system impact assessment published | ISO |
| 2 August 2025 | Obligations for general-purpose AI models become applicable | EU |
| August 28, 2025 | SB25B-004 signed, extending the 2024 law’s effective date | Colorado |
| April 7, 2026 | Concept note for a critical infrastructure profile | NIST |
| April 17, 2026 | SR 26-2 supersedes and replaces SR 11-7 | Federal Reserve |
| May 14, 2026 | SB26-189 signed, repealing and reenacting the 2024 provisions | Colorado |
| 27 July 2026 | Amending regulation enters into force; high-risk dates move | EU |
That is nine dated changes. Does your build business case have a line for regulatory and compliance rework? Budget an in-house AI governance layer as an ongoing cost for as long as your models run.
Five Criteria That Decide Build vs Buy AI Governance
1. Rule Velocity
How often does the specification for this component change, and who tells you when it does? Ask who sets the pace. For your own policy, you may. For a reporting layout, check whether ISO, NIST, or a supervisor does when it revises its requirements. We weight high rule velocity that you do not control heavily toward buying.
2. Evidence Burden
Ask who reads the output and what happens if they do not believe it. Score a dashboard for your own engineering leadership low. Score an audit file for a regulator, an external auditor, or opposing counsel high, because that reader did not build the system. If only your team can interpret the record, it fails this test.
3. Differentiation
Does this component show up in a decision a customer makes? If your governance layer is a reason prospects choose you, that is a real argument for building it. Check the claim first: did governance win any recent deal, or only clear a security review? Treat a condition of sale as a candidate to buy, and a reason for the sale as a candidate to build.
4. Staffing Depth in Year Three
Write the build case against year three as well as year one. Ask who maintains the layer if the engineers who designed it move on. Name that owner now.
5. Cost of the Worst Failure
Name the specific failure and the specific bill. Where the worst case is a halted deployment or a contract that does not renew, stop comparing build cost with license cost. Compare each option against the failure you are paying to avoid.
What Does an In-House AI Governance Layer Really Cost?
We are not going to quote a price. Treat any published build-cost range with caution unless it explains what it includes and how it was calculated, and price the build with your own loaded engineering cost. Pick a horizon long enough for sustaining work to show up; criterion four suggests looking at least as far as year three. The worksheet gives you eight lines and the question that sizes each one.
| Line item | How to compute it over your horizon | The question that sizes it |
| Initial build | Engineers x months to first accepted evidence x loaded monthly cost | Is the milestone “a reviewer accepts the record” or “the service is deployed”? |
| Sustaining engineering | Fractional headcount x months in horizon x loaded monthly cost | Is this line in the plan, with a name next to it? |
| Rule tracking and rework | Hours per month reading regulators and standards, plus rework per change, x months in horizon x blended rate | Who is assigned to notice a change like the nine above? |
| Evaluator upkeep | Re-validation cycles per year x years in horizon x hours per cycle x blended rate | How often do your models, prompts, and evaluators change? |
| Audit and review support | Reviews per year x years in horizon x hours per review x blended rate | Did you count the engineers pulled in to reconstruct evidence? |
| On-call for runtime controls | Rotation size x months in horizon x on-call cost, plus incident hours | Who gets paged when a guardrail blocks production traffic? |
| Storage and retention | Retention period x volume x unit cost, at your required durability | How long must the audit trail outlive the application logs? |
| Deferred roadmap | Sustaining headcount expressed as features not shipped | What does the product team give up to fund the rest? |
The Cost Lines Platform Teams Forget
Check four of those lines twice, because no invoice may arrive to remind you of them. Maintenance against changing rules: ask what each row in the timeline table would have cost you in reading, interpreting, and rebuilding. Evaluator upkeep: plan to re-validate your tests when the model changes. Audit support: count the engineer time spent answering the auditor, as well as the auditor’s fee. On-call: treat a control in the request path as a production service with a pager.
Run the Same Lines Against the Buy Option
Go through the worksheet again for the buy option, and check each line before you delete it. Swap initial build for license cost. Keep an integration line, and ask each vendor how many hours a comparable customer spent. Name the person who will administer the product. Ask whether rule tracking moves to the vendor, and get the answer into the contract.
Do not assume responsibility moves with the work. If you are a bank, supervisors note that you “may not receive from the vendor the underlying code, data, or methodology” and still hold that “the principles of model risk management remain applicable.” They call “the validation of vendor products, either by internal or outside parties” an “important element of model risk management.” So add one line to your buy column: the cost of validating the platform itself.
When Build Wins
Build can be the right call, and a decision framework should say so. We would want these four conditions to hold together.
- You write your own requirements, and they are stable. No external supervisor, standards body, or customer contract writes your control specification.
- Customers see the layer and pay for it. You can point to revenue that turns on it.
- Sustaining headcount is funded for the life of the models. Named people, in a plan that survives a bad quarter.
- One team owns a small estate. Ask how many models you run, how alike they are, and whether one team owns them. Then ask whether one in-house layer could keep up if that estate grows across business units and has to scale.
There is a fifth case: a deployment boundary no vendor you have evaluated can operate inside. Confirm it before you assume it. Ask each vendor of AI governance solutions on your list whether its product runs on-premises or air-gapped. Our piece on air-gapped AI requirements lists what to check.
When Buy Wins
Buying fits when those conditions invert. Outsiders write your requirements and revise them. Governance gates your deals but closes none. You cannot promise sustaining capacity past the next reorganization. Your model estate is already mixed and growing across teams.
Compare total cost, as you would in any build vs buy software decision. Then add a second cost for governance: proving the capability worked, on a date in the past, to someone who was not there. Add it to any build vs buy AI decision an outside reviewer will read.
The Hybrid Path: Buy the Evidence and Evaluation Layer, Build the Integrations
The useful output of a build vs buy review can be a boundary instead of a single verdict. Here is the one we recommend you test first.
Buy the evidence spine and the test runner: the tamper-evident record, the attribution trail, the control mapping, stored results per version, and reports your external reviewer accepts.
Build the policy content and the integrations: your thresholds, your prohibited uses, your escalation paths, your evaluation cases, and the connectors into your CI/CD pipeline, identity system, ticketing, and data platforms.
When the Hybrid Path Fits
- Your scoring worksheet splits: evidence and reporting lean buy, policy and inventory lean build.
- You have platform engineers who can own integrations but nobody to spare for rule tracking.
- You run models from more than one provider and want one evidence format across them.
Skip the hybrid if the product cannot export policy and evidence in a format you control. Skip it too if one lightly configured product covers your rows. Count the coordination cost of splitting ownership before you commit.
If you want to see how one vendor answers, Seekr states that SeekrGuard “independently evaluates AI models and agents for organizations using AI in high-stakes environments,” with teams testing candidates “on their own data and criteria.” Seekr also says SeekrFlow offers “data attribution, confidence scoring, and prompt comparisons,” and can run on your own compute, on-premises, or fully air-gapped. Put the questions below to us as you would to anyone else.
Staffing and Time to First Audit: Questions to Answer Before You Commit
We will not give you a headcount or a duration, because both depend on your estate. These questions produce your own numbers.
- Which skills does your build need (platform, testing, security, risk and compliance), and which of those people already work full time on another roadmap?
- Who provides independent review? The banking guidance describes “effective challenge” as critical analysis by “objective experts” with “sufficient independence.” If your builders also validate, who challenges them?
- When is your next external review, customer audit, or certification assessment? Count backward from that date.
- What is the first artifact the reviewer will ask for, and which component produces it?
- For the buy option: how long did comparable customers take to hand an auditor their first report? Ask for references, and call them.
What to Ask an AI Governance Platform Vendor, and What to Ask Your Own Team
Use the same questions on both sides. Treat a build proposal as a bid from an internal vendor, and hold it to the bar you set for AI governance tools on the market.
| Topic | Ask the vendor | Ask your own team before building |
| Scope | Which of the six components do you cover? | Which of the six are in the estimate, and which are “phase two”? |
| Rule changes | How soon after SR 26-2 did your control mappings reflect it? | Who read SR 26-2 the week it came out, and what did we change? |
| Agents | What evidence do you record for a multi-step agent run? | Do we record tool calls and intermediate steps, or only final outputs? |
| Evaluation | Can we bring our own test cases and reproduce a result later? | Who re-validates our evaluators when the model changes? |
| Evidence | Show us a report an external auditor accepted without your staff present. | Has anyone outside the team read our audit record cold? |
| Deployment | Does it run on-premises or air-gapped, where our data has to stay? | Can we operate it in each environment our models run in? |
| Exit | In what format do policy and evidence leave your system? | If we replace this later, what do we keep? |
Treat a vague answer as a finding, from either side. And know the difference between a monitoring product and a governance one before you compare them. Ask a monitoring product what a model did. Ask an AI governance platform who allowed it and under which rule.
Exit and Lock-In: How Do You Leave Either Option?
Check lock-in on both sides. Ask what would tie you to a vendor. Ask, too, who could maintain an in-house layer if the engineers who designed it move on.
The federal acquisition memo is a useful model even if you never sell to government. OMB M-25-22 tells agencies to “pay careful attention to vendor sourcing, data portability, and long-term interoperability to avoid significant and costly dependencies on a single vendor.” A footnote adds “open and standard data formats and application programming interfaces (APIs).” Borrow that language for your own contract.
Exit Questions for a Bought Platform
- Can we export our policy definitions in a documented, machine-readable format?
- Can we export the full evidence history, with timestamps and version references, and read it without your software?
- Who owns our test cases, results, and any evaluator tuned on our data?
Exit Questions for an In-House Build
- Is the audit record in a standard format, or one only our code can parse?
- How many people can explain the design today, and is it written down for a successor?
- If we later buy, which parts of the build survive as integrations?
Ask how long your reviewers expect evidence to be kept, and whether that is longer than you will keep the system that produced it. Keep policy content in your own definitions and insist on portable evidence, so that a platform change is closer to an integration than to a rewrite.
A Scoring Worksheet You Can Fill In
Plan one session with the platform lead and the risk owner in the room. Score every one of the six components against the five criteria, from 1 to 5. A 1 means the rule is yours, only your team reads the output, the component wins no deals, nobody owns it in year three, and the worst failure is an internal complaint. A 5 means the opposite on each.
| Component | Rule velocity | Evidence burden | Differentiation | Year-three owner | Worst failure | Lean |
| Inventory | ||||||
| Policy | ||||||
| Evaluation and testing | ||||||
| Runtime controls | ||||||
| Evidence and audit trail | ||||||
| Reporting |
How to Read Your Scores
- Lean toward buying the component when rule velocity, evidence burden, and worst failure score high together.
- Lean toward building it when differentiation is high and the year-three owner is named and funded.
- Both sides high? That row is a hybrid candidate: buy the engine, build the content.
- If the year-three owner scores 1, we would not build that row, whatever the other columns say.
Write “build,” “buy,” or “hybrid” in the last column. If the rows split, treat the split as your architecture.
Talk Through Your Build vs Buy Decision
Book a consultation with an AI expert. Share your challenges and objectives, and our team will connect to explore solutions and walk you through a live demo of SeekrFlow.
Sources
- Federal Reserve, SR 26-2 (April 17, 2026)
- Federal Reserve, FDIC, and OCC, interagency model risk text attached to SR 26-2
- OCC Bulletin 2026-13
- European Commission, AI Act page and application timeline
- OMB Memorandum M-25-21 (April 3, 2025)
- OMB Memorandum M-25-22 (April 3, 2025)
- NIST, AI Risk Management Framework
- NIST, AI RMF Core
- ISO/IEC 42001:2023
- ISO/IEC 42005:2025
- Colorado General Assembly, SB25B-004
- Colorado General Assembly, SB26-189
FAQ
What is an AI governance platform?
As we use the term, it is software that keeps the registry of models and agents, applies policy, runs evaluations, enforces controls on live outputs, and stores the evidence a reviewer will ask for. Check each product you shortlist against the six components.
Should we build or buy AI governance?
Decide per component. Build what encodes your own judgment, such as policy and test cases. Consider buying what an outside party specifies and revises, such as evidence formats and control mapping.
What is the difference between AI governance software and model monitoring tools?
Ask each product two questions. Does it watch how a deployed model behaves? Does it also record which policy applied, who approved the model, and what evidence supports that approval? In this guide, the first job is monitoring and the second is governance.
How much does an AI governance platform cost compared with building one?
Be wary of any published average that does not explain what it includes and how it was calculated. Use the eight-line worksheet with your own loaded costs over your planning horizon, and add a validation line to the buy column.
How long does an in-house build take with a strong platform team?
Pick the milestone before you estimate. Deploying the service is one. A reviewer outside your team accepting the record is another. Estimate to the second milestone.
Can we build part of an AI governance layer and buy the rest?
Yes. One workable split keeps policy content and integrations in-house and buys evidence, evaluation, and reporting. Before you choose it, check that the bought product exports policy and evidence in a format you control.
What happens to an in-house governance layer when the rules change?
Assign someone to notice the change, interpret it, and rebuild against it. In 2026, the Federal Reserve issued SR 26-2, which “supersedes and replaces” SR 11-7, and an amending regulation moved the EU high-risk dates. If you build in-house, plan for that work on your own roadmap.
Does buying an AI governance platform transfer accountability to the vendor?
Plan as if it does not. For banks, US supervisors say model risk management principles “remain applicable” to vendor products, and they describe validating those products as an important element. If you are outside banking, ask your own reviewer who they will hold responsible.
Accelerate your path to AI impact
Book a consultation with an AI expert. We’re here to help you speed up your time to AI ROI.
Request a demo