AI inference glossary

The terms that come up when you run language models in production, defined in a sentence or two.

AI inference
Running a trained model on new input to produce an output. For language models, generating a response to a prompt, token by token. Read more
Context window
The maximum number of tokens a model can consider in one request, prompt and answer together.
Cryptographic erasure
Deleting data by destroying the key that decrypts it, so every copy becomes unreadable at once wherever it is stored.
Data residency
A requirement that data is stored and processed in a particular country or region, such as the UK.
Decode
The second phase of LLM inference, where tokens are generated one at a time, each step reading the full KV cache.
Dedicated GPU endpoint
A model served on a GPU reserved for one customer, rather than shared hardware billed per token. Read more
Governed inference
Inference where every request runs under enforceable, provable controls: isolation, encryption, residency, erasure and evidence. Read more
KV cache
The attention keys and values a transformer stores for every token it has processed, so each new token does not recompute the whole prompt. It grows with context length and is a working copy of the prompt. Read more
LoRA
Low-rank adaptation: fine-tuning a small set of adapter weights while the base model stays frozen. Read more
Open-weight model
A model whose trained weights are published, so it can be run on your own or rented hardware. Examples include Qwen, Gemma and Mistral.
OpenAI-compatible API
An endpoint that accepts the same requests as OpenAI's chat completions API, so existing SDKs work by changing the base URL and key.
Policy-as-code
Governance rules written as versioned code and enforced automatically, instead of documents applied by hand.
Prefill
The first phase of LLM inference, where the whole prompt is processed in parallel and the KV cache is built.
Prefix caching
Reusing cached KV blocks when requests begin with the same tokens. Fast, but a leakage risk if blocks are shared across tenants.
QLoRA
LoRA over a base model quantised to 4 bits, so larger models can be fine-tuned on smaller GPUs.
Restricted transfer
Under UK GDPR, sending personal data to a country outside the UK. It needs adequacy regulations or safeguards such as the IDTA. Read more
Tenant isolation
Guaranteeing that one customer's data, including cached context, can never be read or reused by another.
Throughput
Total tokens generated per second across every request on a GPU. It sets the cost per token.
Time to first token (TTFT)
How long a request waits before the first token of the answer arrives. Dominated by prefill on long prompts.
UK-based AI inference
Inference whose compute, storage and operator are all in the United Kingdom, so prompts and context never make a restricted transfer. Read more