How It Works
Prompt caching is automatic — there is nothing to enable and no API changes are required:- The inference engine (vLLM / SGLang) keeps the KV-cache of recently processed prompt prefixes inside the TEE.
- When a new request arrives whose prompt starts with a cached prefix, only the new suffix is computed.
- Cached input tokens are billed at the model’s discounted cache-read rate instead of the regular input rate.
/v1/models response (pricing.input_cache_read).
Tracking Cache Hits
The number of input tokens served from cache is reported in theusage object of every response:
Chat Completions (/v1/chat/completions):
/v1/responses):
cached_tokens is 0 (or prompt_tokens_details is null). Repeat the request — or send another request sharing the same prefix — and cached_tokens reflects the reused portion.
The Responses API is stateless, so your application sends the relevant history with each turn. Keeping that caller-managed history in a stable order creates the repeated prefix that prompt caching can reuse. See Stateless Responses.
Getting the Most Out of the Cache
- Put stable content first. The cache matches prefixes: keep your system prompt, instructions, and reference documents at the start of the message list, and the parts that change (the user’s latest question) at the end.
- Keep prefixes byte-identical. Any change at position N invalidates the cache from that point on — even reordered JSON keys or a timestamp in the system prompt.
- Short prompts may not be cached. There is a minimum prefix length (one engine-internal cache block) before caching kicks in; very short prompts are always computed in full.
- The cache is per model deployment. Cache state lives inside each model’s TEE and is evicted over time as capacity is needed. Frequent traffic against the same prefix keeps it warm.
PrivacyThe cache never leaves the TEE — cached prefixes are part of the same hardware-isolated inference state as the rest of your request and are inaccessible to NEAR or the infrastructure provider.