Governed inference

Inference you can prove, not just promise.

kvrun treats the KV cache, the working copy of every prompt, as governed data: encrypted, isolated per tenant, erasable on request and logged on every access. Without slowing your models down.

Every request checked, sealed and logged
Time-to-first-token overhead
<3%
Keys, cache and GPU
Per tenant
Provable erasure of a session
<60s
Audit evidence export
1 click

The problem governed inference solves

On long-context work, the KV cache holds 60 to 80% of inference memory, and it is a faithful copy of what users sent. In most serving stacks it is unencrypted, can be shared between tenants through prefix caching, is never logged, and cannot be provably deleted. It is the most sensitive copy of your data, and the least governed.

Five controls, on every request

ControlWhat kvrun does
EncryptionEvery KV block sealed with AES-256-GCM under a key unique to the workspace and session
IsolationSeparate key hierarchies per tenant; prefix caching never crosses tenants
ErasureDestroying a session's key makes every copy, on every tier, unreadable at once, with a certificate
ResidencyPolicy-as-code decides where context may live; policies ship through Git, not engine rebuilds
EvidenceAllocation, access, movement and erasure written to an append-only log you can export

More context on the same GPU

Governance and capacity come from the same design. Sealed blocks can move safely from GPU memory to host memory, local disk and object storage, so a deployment holds far longer contexts than its VRAM alone allows. Early internal benchmarks show up to 4× the context per 80 GB GPU.

Evidence for the frameworks you answer to

FrameworkRequirementEvidence produced
UK GDPR Art. 17Right to erasureErasure certificate per session
UK GDPR Art. 25Data protection by designEncryption and isolation on by default
SOC 2 CC6.1Logical accessPer-tenant key isolation
SOC 2 CC7.2MonitoringContinuous audit log
ISO 27001Secure disposalCryptographic erasure
EU AI ActEvaluation recordsAn evaluation on every fine-tune, kept with the weights

kvrun produces evidence that maps to these controls. Certification of your organisation remains between you and your auditor.

Questions

What is governed inference?
Running models so every request is under enforceable, provable controls: isolation, encryption, residency, erasure and evidence. Read the full explainer.
Which serving engines does it work with?
The governance layer sits between the serving engine and cache storage, and has been benchmarked with vLLM, SGLang, TensorRT-LLM and llama.cpp.
How fast is erasure?
Erasing a session destroys its keys, which takes effect immediately across every tier. The certificate is issued in under a minute.

Governed, private, open models

Pick a model. Leave with an endpoint.