Defense & Governance · 6 min

Governance and Compliance: NIST AI RMF, the EU AI Act, and Provenance

A working model of AI governance: the NIST RMF as an operating loop, the EU AI Act's risk tiers and obligations, and the provenance artifacts that make a system auditable.

Governance is the part of the stack engineers reflexively file under "someone else's problem." That instinct is expensive. The two documents that now shape how AI ships both assume the same thing about you: that you can show your work. One is the NIST AI Risk Management Framework, voluntary and US-origin, but the de facto vocabulary everyone reaches for. The other is the EU AI Act (Regulation (EU) 2024/1689), legally binding with fines that scale to global revenue. If your team can't produce a record of what a model is, what went into it, and how you decided it was safe enough to deploy, you don't have a governance gap. You have an evidence gap, and this lesson is about closing it.

NIST AI RMF as an operating loop

The AI RMF 1.0 (published 26 January 2023) organizes everything around four functions: Govern, Map, Measure, Manage. Reading them as a flat checklist is the common mistake. Three of them form a loop, and the fourth wraps around it.

        ┌─────────────────────────────────────────┐
        │  GOVERN  (culture, policy, accountability)│
        │  cross-cutting — sits above everything     │
        │  ┌───────┐   ┌─────────┐   ┌────────┐     │
        │  │  MAP  │──▶│ MEASURE │──▶│ MANAGE │──┐  │
        │  │context│   │ metrics │   │controls│  │  │
        │  └───────┘   └─────────┘   └────────┘  │  │
        │      ▲                                 │  │
        │      └────────── feedback ─────────────┘  │
        └─────────────────────────────────────────┘

Map establishes context: who the stakeholders are, where the system boundary sits, what could go wrong and to whom. Measure turns those mapped risks into something quantified, whether benchmarks, red-team results, drift metrics, or error rates broken out by subgroup. Manage allocates resources against what you measured: which risks you mitigate, which you accept, which kill the project. Govern is the cross-cutting function, the only one that spans the whole organization. It is the policies, roles, and accountability that make Map→Measure→Manage a repeatable process rather than a one-off heroic effort before launch.

The framework ships with a companion Playbook of suggested actions per subcategory, and, for anyone shipping LLMs, the Generative AI Profile (NIST AI 600-1, July 2024). That profile enumerates twelve risks specific to generative systems, including CBRN uplift, confabulation, data privacy, information integrity, and value-chain integration, then maps concrete actions back onto the four functions.

Defender: Treat Measure as a continuous signal, not a launch gate. A jailbreak eval that passed on deployment day tells you nothing about the model after a silent provider-side weight update. Wire your red-team suite into CI and alert on regression. That is Measure feeding Manage feeding back into Govern.

Nobody certifies you against the RMF; there is no auditor who stamps "RMF-compliant." Its value is the shared vocabulary and the forcing function. If you can't fill in Map for a system, you don't understand that system well enough to ship it.

The EU AI Act's risk tiers

Where the RMF is a process, the Act is a product-safety law with teeth. It sorts systems by use, not by model architecture.

UNACCEPTABLE   social scoring, manipulative/subliminal,
(prohibited)   real-time biometric ID in public (narrow carve-outs)
               → banned outright.  Applied 2 Feb 2025.
─────────────────────────────────────────────────────────────
HIGH-RISK      Annex III (use-based: hiring, credit, biometrics,
               critical infra, education) + Annex I (safety
               components of regulated products)
               → conformity assessment, risk-mgmt system,
                 data governance, logging, human oversight,
                 technical documentation, post-market monitoring
─────────────────────────────────────────────────────────────
LIMITED        chatbots, emotion recognition, deepfakes
               → transparency: tell people it's AI / synthetic
─────────────────────────────────────────────────────────────
MINIMAL        spam filters, game AI → no mandatory obligations

Two things engineers get wrong here. First, most obligations attach to high-risk, and "high-risk" is about deployment context, not how clever the model is. A boring gradient-boosted model screening résumés is high-risk. A frontier LLM writing marketing copy is not. Same model class, different tier depending on where you point it.

Second, general-purpose AI (GPAI) models sit on a parallel track. Under Article 53, every GPAI provider placing a model on the EU market must keep technical documentation per Annex XI, give downstream integrators enough information to meet their obligations, maintain a policy to comply with EU copyright law (including the text-and-data-mining opt-out from Directive (EU) 2019/790), and publish a "sufficiently detailed summary" of training content using the AI Office's template. Free and open-source models get a narrow exemption from the first two duties, but only if they are not classified as systemic-risk. That classification triggers at a bright line: training compute above 10²⁵ FLOPs (Article 51), which pulls in extra evaluation and incident-reporting duties.

The penalties are built to hurt. Up to €35M or 7% of global annual turnover, whichever is higher, for prohibited-practice violations; €15M or 3% for most other obligations; €7.5M or 1% for supplying incorrect information. SMEs pay the lower of the two figures rather than the higher. For a large firm, "whichever is higher" means the percentage is the real number.

On timing: the Act entered into force 1 August 2024, prohibitions applied February 2025, and GPAI obligations followed in August 2025. The Annex III high-risk deadline, originally 2 August 2026, was pushed to 2 December 2027 by the Digital Omnibus package adopted in 2026 (a 16-month deferral), with product-embedded high-risk systems under Annex I following later. That is relief on the calendar, not on the substance. Build as if the clock is still running, because documentation lead time is measured in quarters.

Builder: Fine-tune and redistribute a GPAI model and you may become a provider in the Act's sense, inheriting Article 53 duties. Modifying an open model is not a free pass; it can move you up the obligation stack.

Provenance: the evidence layer

Both regimes converge on one demand: an auditable record. Three artifacts carry it.

Model cards (Mitchell et al., 2019) document intended use, training data at a high level, evaluation results disaggregated by relevant subgroup, and known limitations. Datasheets for datasets (Gebru et al., 2021) do the same one layer down: provenance, collection process, consent, known biases. Together they answer "what is this and what went into it."

Data lineage is the runtime version, the traceable chain from a training or RAG corpus back to its sources with the transforms in between. It is what lets you answer a copyright takedown or a GDPR erasure request without guessing.

AI/ML-BOM extends the software-bill-of-materials idea to models. Standard SBOM formats now carry ML extensions: CycloneDX has an ML-BOM, and SPDX 3.0 added an AI profile. An AI-BOM enumerates the base model, datasets, dependencies, and their licenses, so when a base checkpoint turns up poisoned or its license gets revoked, you can query "which of our systems include this?" in seconds instead of a week of archaeology.

AI-BOM  ──▶ base model (weights, license, source, systemic-risk?)
        ──▶ datasets   (datasheet ref, lineage, consent basis)
        ──▶ deps       (tokenizer, guardrail model, versions)
        ──▶ evals      (model card ref, red-team date, results)
Researcher: Provenance is also an integrity control. The 2024 discovery of malicious pickle-serialized models on public hubs is exactly the supply-chain threat that an AI-BOM plus signed lineage is meant to catch. It is the same threat model as a compromised npm package, applied to weights. For output provenance (is this media synthetic?), C2PA content credentials are the emerging input-complementary standard.

The through-line: governance work you can't produce as an artifact didn't happen, as far as an auditor or a court is concerned. Wire the RMF loop into your engineering process, know which Act tier each deployment lands in, and generate provenance records as a build output rather than a scramble before review. For where this connects downstream, see the module lesson on incident response and post-market monitoring; the Manage function and the Act's post-market obligations are the same muscle, exercised after launch.

Sources