A workable first milestoneChoose three real AI use cases and run them through Govern, Map, Measure and Manage. Improve the workflow from evidence before expanding the inventory.
GovernPolicy, risk appetite, roles and oversight.
MapContext, impact, actors and system boundaries.
MeasureEvaluation methods, limits and confidence.
ManagePriorities, treatment, monitoring and response.

Govern: decide who may say yes

Govern establishes the conditions for responsible decisions. Define an executive sponsor, business-use owner, technical owner and independent reviewer. Write down which uses are prohibited, which require enhanced review and who may accept residual risk.

Keep the policy short enough to use. Put the operational detail in standards for data, model sourcing, evaluation, human oversight, logging and incident handling. A policy should establish direction; it should not try to be the engineering manual.

  • Publish an AI-use intake route
  • Define risk tiers and approval thresholds
  • Assign provider and model inventory ownership
  • Create an exception expiry date

Map: describe the system people actually use

Map the full service, not a box labelled “LLM.” Include data sources, retrieval, prompts, model endpoints, guardrails, user groups, plug-ins, tools, downstream decisions and support processes. Identify people affected even when they never interact with the system.

Capture assumptions that could break: the provider will not train on prompts, retrieved documents belong to the requester, tool calls are reversible or a human will notice a bad answer. These assumptions become test cases.

  • Draw trust boundaries and data flows
  • Identify foreseeable misuse and affected groups
  • Record external dependencies and contractual claims
  • Name the conditions that would make the use case unacceptable

Measure: test claims with evidence

Measurement is broader than model accuracy. Test prompt injection, sensitive disclosure, insecure output handling, cross-user retrieval, excessive agency, fallback behaviour and monitoring coverage. Record both the result and the limits of the test.

Use a repeatable evaluation set for the most important tasks. Include ordinary user behaviour, hostile instructions and operational failures such as timeouts or unavailable tools. A percentage without test context is not assurance.

  • Maintain versioned evaluation cases
  • Set launch thresholds before seeing results
  • Review false positives and false negatives
  • Retest material changes

Manage: connect findings to release and operations

Prioritise treatment based on impact and likelihood, then make the decision visible. Some risks need technical controls; others need a narrower use case, stronger review, user communication or a decision not to launch.

After launch, watch for changes in data, models, users and connected actions. Feed incidents and overrides back into the evaluation set. The framework should become a loop, not a gate passed once.

  • Assign every finding an owner and due date
  • Use staged rollout and transaction limits
  • Define suspend and rollback criteria
  • Report unresolved risk to the person authorised to accept it

A practical 30-day field plan

Week one — Govern. Policy, risk appetite, roles and oversight. Put one accountable owner in the room, agree which business service or decision is in scope, and record the assumptions the team is making. Resist the urge to begin with a technology purchase; the first deliverable is a shared view of the problem and the authority to change it.

Week two — Map. Context, impact, actors and system boundaries. Walk through the current process with the people who operate it. Compare the written design with real access paths, data flows, exceptions and on-call practice. Mark every point where an owner is missing or where the team cannot produce evidence that a control works.

Week three — Measure. Evaluation methods, limits and confidence. Choose a narrow pilot that can be observed safely. Define the expected result, the rollback path and the person who may accept a trade-off. Capture operational friction as product feedback; controls that are difficult to use will eventually be bypassed.

Week four — Manage. Priorities, treatment, monitoring and response. Review the pilot with engineering, operations, security and the service owner. Close urgent gaps, assign longer work to a funded backlog and set the next evidence review. The month should end with a repeatable operating rhythm, not a one-time presentation.

Evidence worth keeping

Good evidence is understandable outside the team that created it. Keep a concise record that connects the decision, owner, technical implementation and observed result. Screenshots can support evidence, but configuration, logs, test output and approved records are stronger because another person can reproduce or challenge them.

  • Publish an AI-use intake route — owner, current state, last validation and any open exception
  • Draw trust boundaries and data flows — owner, current state, last validation and any open exception
  • Maintain versioned evaluation cases — owner, current state, last validation and any open exception
  • Assign every finding an owner and due date — owner, current state, last validation and any open exception
  • Decision log showing who approved residual risk and when it will be reviewed
  • Test or exercise result with the actual outcome, not only a pass label

Questions for the leadership review

Use these questions to keep the discussion connected to operating risk rather than tool activity. A useful answer names a person, a service and evidence.

  • Who is accountable for the ai framework outcome when teams disagree about delivery and risk?
  • Which critical service or customer promise would be affected by a failure in this area?
  • What evidence would tell us the design is working in production today?
  • Which exception creates the largest concentration of access, dependency or recovery risk?
  • What would the team contain first, and how would it restore a trustworthy service?
  • Which improvement can be completed in the next 30 days without waiting for a large programme?

Common questions

Do we need to implement every AI RMF subcategory?

No. Use the framework to select outcomes that fit the use case, risk tolerance and obligations, and document that scope.

How long should a first implementation take?

A small cross-functional team can test an initial workflow on a few use cases in weeks; organisation-wide maturity is an ongoing programme.