Last updated: September 3, 2026
Implementing the NIST AI Risk Management Framework means turning four voluntary functions (Govern, Map, Measure and Manage) into a repeatable operating routine: a governance structure with real authority, an inventory of every AI system with a named owner, documented risk analysis on the systems that matter, and monitoring that continues after launch. Most organizations reach a defensible baseline in one to two quarters by starting with Govern and Map on a handful of high exposure systems rather than attempting all 72 subcategories at once.
This guide is part of our AI governance guide. If you are still deciding whether the framework fits your organization, start with our explainer on what the AI RMF is and come back here for the rollout.
Key takeaways
- AI RMF 1.0 contains four functions, 19 categories and 72 subcategories. NIST expects you to build a profile, meaning a selected subset, not to implement everything.
- Govern is cross cutting and is the usual failure point. Without a named decision maker who can say no, the rest becomes documentation theater.
- The AI system inventory is the hard prerequisite. You cannot map risk on systems nobody has found yet, and unsanctioned tools are where the surprises live.
- Measure only works if your metrics are repeatable. Pick a small set you can rerun on every model or prompt change.
- The framework is not certifiable. If you need proof for customers or regulators, pair it with ISO 42001 or with EU AI Act obligations.
What does implementing the AI RMF actually require?
AI RMF 1.0 was published in January 2023 as NIST AI 100-1. It is voluntary, free and sector agnostic. There is no auditor, no certificate and no deadline. That is the appeal and also the difficulty: nothing external tells you when you are finished, so scope has to come from you.
The structure is straightforward. Govern is the cross cutting function with 6 categories covering policy, accountability, workforce, and third party risk. Map (5 categories), Measure (4 categories) and Manage (4 categories) are applied at the level of an individual AI system. Across the four functions sit 72 subcategories, each one a statement of an outcome rather than a prescribed control. The companion AI RMF Playbook, hosted in the NIST Trustworthy and Responsible AI Resource Center, suggests concrete actions for each subcategory, and it is more useful during implementation than the framework document itself.
Everything the framework asks for ties back to seven characteristics of trustworthy AI: valid and reliable, safe, secure and resilient, accountable and transparent, explainable and interpretable, privacy enhanced, and fair with harmful bias managed. When you are deciding whether a control is worth building, ask which of the seven it protects. If the answer is none of them, it is probably paperwork.
Step 1: Set scope and stand up Govern
Start with a written scope statement covering which business units, which system types and which risk tier you are addressing first. A scope that says all AI everywhere produces a program that finishes nothing.
Then build the Govern layer, which is mostly three artifacts. First, an AI policy that states permitted and prohibited uses, who approves new systems, and what happens when someone deploys without approval. Second, a decision forum, whether that is an existing risk committee with an added agenda item or a dedicated AI review board. Third, a RACI that names the accountable executive for AI risk. In practice, GRC or risk owns the framework, data science owns technical evaluation and legal advises on use restrictions. Programs that hand the entire framework to the data science team stall, because Govern requires authority that team does not have.
Step 2: Build the AI system inventory
Map begins with context, and context begins with knowing what exists. Build an inventory that records, for each system: the business purpose, the owner, whether it was built or bought, the model or vendor behind it, the data it touches, whether it makes or influences decisions about people, and whether outputs reach customers.
Three sources will find most of what you have. Procurement and expense records surface purchased tools. Identity and SSO logs surface applications employees signed into. Engineering repositories and cloud billing surface internally built systems and API usage. Expect the first pass to find two to three times more AI usage than leadership assumed, mostly embedded features in software you already bought.
Rank the inventory by exposure rather than by build date. Systems that affect employment, credit, healthcare, education, insurance or access to services belong at the top, along with anything that generates content shown to customers. Those are also the categories regulators care about, which is why the same list drives your work under the EU AI Act if you operate in Europe.
Step 3: Run Map on the systems that matter
For each priority system, work through the Map categories with the people who built and use it. The questions that produce real answers are practical: what decision does this system influence, who is affected when it is wrong, what does wrong look like, what is the fallback, and what were the training data and its known limits. Document the answers in a short system record rather than a long report. Two pages that stay current beat forty pages written once.
Map also asks you to characterize impacts on individuals, groups and society. Teams often skip this as vague. A better framing is to name the three worst plausible outcomes for a person on the receiving end, and note whether anyone would find out if it happened. If the answer is no, you have discovered a monitoring gap before it becomes an incident.
Step 4: Choose measurements you can repeat
Measure carries 22 subcategories, more than any other function, and it is where implementations get expensive. The discipline is to select a small set of tests tied to the risks you actually mapped, then run them on a schedule and on every material change.
For most enterprise systems that means four things: a performance baseline on representative data, a fairness check across the groups your system affects, a robustness or red team exercise proportionate to exposure, and a record of human oversight, meaning how often reviewers overrode the system and whether those overrides were correct. Generative systems need an additional layer. NIST published the Generative AI Profile as NIST AI 600-1 in July 2024, and it names risks conventional model validation misses, including confabulation, information integrity, data privacy leakage, intellectual property exposure and value chain opacity. If you deploy generative features, use that profile as your test list.
Step 5: Treat, monitor and close the loop
Manage turns findings into decisions: accept, mitigate, transfer or stop. Each priority system should end this step with a treatment decision recorded by a named person and a date, plus post deployment monitoring, an incident path for AI specific failures, and a decommissioning trigger. Manage also expects a plan for third party risk, which loops back to the vendor questions you defined in Govern.
The loop matters more than the artifacts. A system whose risk record was updated after its last model change is governed. A system with a beautiful assessment from eighteen months ago is not.
How long does implementation take?
| Phase | Typical duration | What good looks like at the end |
|---|---|---|
| Scope and Govern | 4 to 6 weeks | Policy approved, accountable owner named, review forum meeting |
| Inventory | 3 to 6 weeks, overlapping | Every AI system recorded with an owner and an exposure tier |
| Map on priority systems | 6 to 10 weeks | System records for the top tier, impacts characterized |
| Measure | Ongoing, first cycle 6 to 8 weeks | A repeatable test set running on a schedule |
| Manage and monitoring | Ongoing | Treatment decisions recorded, monitoring and incident path live |
A mid sized company that sequences this way is usually at a defensible baseline in one to two quarters. Full coverage across every system and all four functions is a multi year program, which is why exposure based sequencing beats function by function completion.
What evidence should you keep?
Because there is no AI RMF audit, the evidence question is really about who else will ask. Customers send AI questionnaires. Insurers ask about model risk. Regulators in Europe ask for technical documentation. Internal audit asks whether the policy is followed.
The set that satisfies all of them is small: the approved policy with version history, the inventory with owners and dates, system records for priority systems, evaluation results with dates and the model version tested, treatment decisions with approvers, and evidence that monitoring ran. If you later pursue certification, that same set maps closely onto what ISO 42001 requires, and our ISO 42001 checklist shows where the overlap sits.
Frequently asked questions
Do we have to implement all 72 subcategories?
No. NIST designed the framework around profiles, meaning a selected subset appropriate to your context, sector and risk tolerance. Selecting a subset and documenting why is the intended use. Attempting all 72 in one program is the most common reason implementations stall.
Can we be certified against the NIST AI RMF?
No. There is no accredited certification scheme and no auditor. You can self assess, commission a third party readiness review, or use AI RMF as the design basis for a certifiable management system. Any vendor advertising NIST AI RMF certification is describing its own assessment.
Where should a company with no AI governance program start?
Inventory first, then Govern. You can write a policy in a week, but you cannot scope it sensibly until you know what you are governing. Running both in parallel, with the inventory informing the policy, is the fastest honest path.
Does implementing AI RMF help with the EU AI Act?
Partly. The evidence overlaps well, particularly risk management, data governance, technical documentation and human oversight. It does not satisfy the law by itself, because the AI Act adds legally defined obligations, conformity assessment and registration that a voluntary framework does not address. Our EU AI Act compliance checklist covers what the regulation adds.
Is AI RMF 1.0 still current in 2026?
Yes, with a revision underway. NIST has stated that AI RMF 1.0 is being revised as part of the White House AI Action Plan, and in April 2026 released a concept note for an AI RMF profile on trustworthy AI in critical infrastructure. The four function structure is expected to persist, so inventory, evaluation and monitoring work you do now will carry over.
From framework to evidence
The difficult part of AI RMF implementation is not understanding the four functions. It is keeping the inventory accurate, the evaluations current and the approvals traceable while the AI estate keeps growing. Teams that manage this in spreadsheets usually find the record is stale by the time someone asks for it. Compyl gives GRC teams one place to hold the AI inventory, run the risk workflow, collect evaluation evidence automatically and reuse it across frameworks, so an ISO 42001 certification effort or an EU AI Act response starts from work you have already done rather than from a blank page. If you want to see how that looks against your current AI estate, we can walk through it.
