What is governed inference?

· 3 min read · kvrun team

Governed inference is running AI models so that every request is subject to rules you can enforce and prove. That covers who may call the model, where prompts and context are stored, how long anything is kept, how it is erased, and what evidence is produced along the way. It turns inference from a black box into a system you can audit.

Most AI governance work today happens on paper: policies, model cards, risk registers. Those matter, but they describe intentions. Governed inference is about the runtime, where a prompt containing a customer's personal data actually meets a GPU.

Why inference needs governing at all

A model call looks stateless from the outside: text in, text out. Inside the serving engine it is anything but. The prompt is tokenised, turned into attention keys and values for every layer, and held in GPU memory as the KV cache while the answer is generated. On long-context work that cache can be tens of gigabytes per request, and it is a faithful representation of everything the user sent.

In a typical serving stack that memory is treated as disposable plumbing. It is unencrypted, it can be shared between users through prefix caching, it is never logged, and "deleting" it means waiting for it to be overwritten. For a regulated business that is a gap: the most sensitive copy of the data is the one nobody governs.

The five controls of governed inference

ControlQuestion it answersWhat good looks like
IsolationCan one customer's data reach another?Separate keys and memory per tenant, including cached context
EncryptionWhat does someone with memory access see?Context sealed at rest in every tier, not just on the wire
ResidencyWhere does the data physically live?Policy that pins compute and storage to a region
ErasureCan we prove it is gone?Deletion by key destruction, with a certificate
EvidenceCan we show an auditor?An append-only log mapped to the controls they test

Each control has to hold on the hot path without making the model noticeably slower. If governance costs 30% of throughput, teams switch it off. The engineering challenge is keeping the overhead in the low single digits.

Governed inference versus "private AI"

Private AI usually means one of two things: the model runs on hardware you control, or the provider promises not to retain or train on your data. Both help. Neither, on its own, gives you isolation inside a shared serving engine, erasure you can prove, or evidence you can hand over. Governed inference includes privacy and adds the proof.

Who needs it

  • Financial services, legal and healthcare teams handling personal or privileged data under UK GDPR and sector rules.
  • SaaS companies serving many customers from one deployment, where cross-tenant leakage would be a breach.
  • Public sector and suppliers with residency requirements, for example keeping data in the UK.
  • Anyone preparing for the EU AI Act who needs records of how models were evaluated and operated.

How kvrun implements it

kvrun runs open models on dedicated GPUs behind a governed KV cache. Every block of context is encrypted with a key unique to the workspace and session, tenants never share cached blocks, and erasing a session destroys its keys so every copy on every storage tier becomes unreadable. Allocations, access and erasure are logged, and the evidence maps to GDPR, SOC 2 and ISO 27001 controls. See the governed inference overview for detail.

Questions

Is governed inference a standard?
Not yet. It is a practical term for applying enforceable controls to model serving, as opposed to governance documents that sit beside it.
Does governed inference slow models down?
It should not do so noticeably. kvrun's early benchmarks show under 3% added time to first token. Treat any figure as workload-specific and measure your own.
Can I get governed inference with closed models?
Only as far as the provider exposes controls. With open models on dedicated compute, you can enforce all five.