> ## Documentation Index
> Fetch the complete documentation index at: https://docs.near.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Prompt Caching

> How NEAR AI Cloud caches prompt prefixes to reduce latency and cost

NEAR AI Cloud automatically caches prompt prefixes on TEE-hosted models. When consecutive requests share a common prefix — a long system prompt, conversation history, few-shot examples, or a large document — the cached portion is reused instead of being recomputed, which lowers both latency and cost.

## How It Works

Prompt caching is **automatic** — there is nothing to enable and no API changes are required:

1. The inference engine (vLLM / SGLang) keeps the KV-cache of recently processed prompt prefixes inside the TEE.
2. When a new request arrives whose prompt starts with a cached prefix, only the new suffix is computed.
3. Cached input tokens are billed at the model's discounted **cache-read** rate instead of the regular input rate.

Per-model cache-read pricing is listed on the [Models page](/cloud/models) and in the [`/v1/models`](https://cloud-api.near.ai/v1/models) response (`pricing.input_cache_read`).

## Tracking Cache Hits

The number of input tokens served from cache is reported in the `usage` object of every response:

**Chat Completions** (`/v1/chat/completions`):

```json theme={"dark"}
{
  "usage": {
    "prompt_tokens": 614,
    "prompt_tokens_details": { "cached_tokens": 576 },
    "completion_tokens": 5,
    "total_tokens": 619
  }
}
```

**Responses API** (`/v1/responses`):

```json theme={"dark"}
{
  "usage": {
    "input_tokens": 3302,
    "input_tokens_details": { "cached_tokens": 1216 },
    "output_tokens": 133,
    "output_tokens_details": { "reasoning_tokens": 96 },
    "total_tokens": 3435
  }
}
```

On the first request with a given prefix, `cached_tokens` is `0` (or `prompt_tokens_details` is `null`). Repeat the request — or send another request sharing the same prefix — and `cached_tokens` reflects the reused portion.

The Responses API is stateless, so your application sends the relevant history with each turn. Keeping that caller-managed history in a stable order creates the repeated prefix that prompt caching can reuse. See [Stateless Responses](/cloud/guides/stateless-responses).

## Getting the Most Out of the Cache

* **Put stable content first.** The cache matches *prefixes*: keep your system prompt, instructions, and reference documents at the start of the message list, and the parts that change (the user's latest question) at the end.
* **Keep prefixes byte-identical.** Any change at position *N* invalidates the cache from that point on — even reordered JSON keys or a timestamp in the system prompt.
* **Short prompts may not be cached.** There is a minimum prefix length (one engine-internal cache block) before caching kicks in; very short prompts are always computed in full.
* **The cache is per model deployment.** Cache state lives inside each model's TEE and is evicted over time as capacity is needed. Frequent traffic against the same prefix keeps it warm.

<Note>
  **Privacy**

  The cache never leaves the TEE — cached prefixes are part of the same hardware-isolated inference state as the rest of your request and are inaccessible to NEAR or the infrastructure provider.
</Note>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.