Fine-tuning open models on private data, in the UK

· 1 min read · kvrun team

Fine-tuning adapts an existing open model to your domain, tone or task using your own examples. With LoRA and QLoRA it takes minutes to hours on a single GPU, not weeks on a cluster. The harder part is doing it with sensitive data: keeping the dataset in the UK, tied to one model, and evaluated in a way you can show later.

LoRA or QLoRA?

LoRAQLoRA
What trainsSmall adapter matrices; the base stays frozenThe same adapters, over a 4-bit quantised base
MemoryBase model at full precisionMuch lower, so bigger models fit smaller GPUs
SpeedFaster per stepSlower per step, cheaper per run
Use whenThe model already fits its GPUYou want a larger model than the GPU would otherwise hold

Keeping private data private

  • Minimise first. Remove fields the task does not need, and pseudonymise identifiers.
  • One dataset, one model. Your data should train your adapter and nothing else, never a shared base model.
  • Short-lived compute. The training job should start, finish and shut down, so nothing lingers on a machine afterwards.
  • UK storage. Keep the dataset and the resulting weights in the same UK region as inference.

Evaluate every run, automatically

A fine-tune you cannot explain later is a liability. Hold out a split of the data, score the tuned model against it, and keep the result with the weights. kvrun does this on every run, turning eval loss and perplexity into a plain verdict and writing a record mapped to the EU AI Act articles it answers to.

Serve it

Once training finishes, the tuned model deploys behind the same kind of private, OpenAI-compatible endpoint as any catalog model, with its own governed KV cache.

Related: private LLM hosting in the UK and the UK GDPR checklist.