> ## Documentation Index
> Fetch the complete documentation index at: https://docs.near.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Usage Reporting API

Use the Usage Reporting API when you need automated, read-only usage and cost reporting for a NEAR AI Cloud organization. Reporting tokens are separate from inference API keys: they can read usage reports for one organization, but they cannot call model inference endpoints or manage organization resources.

## Reporting Tokens

Organization owners and admins can create reporting tokens with the management API. Use your user access token for management calls:

```bash theme={"dark"}
curl https://cloud-api.near.ai/v1/organizations/11111111-1111-4111-8111-111111111111/reporting-tokens \
  -X POST \
  -H "Authorization: Bearer $ACCESS_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
    "name": "finance usage export",
    "expires_at": "2026-12-31T23:59:59Z"
  }'
```

The create response returns the reporting token once:

```json theme={"dark"}
{
  "id": "22222222-2222-4222-8222-222222222222",
  "organization_id": "11111111-1111-4111-8111-111111111111",
  "name": "finance usage export",
  "token": "rpt-0123456789abcdef0123456789abcdef",
  "token_prefix": "rpt-01234567",
  "created_by_user_id": "33333333-3333-4333-8333-333333333333",
  "created_at": "2026-07-09T12:00:00Z",
  "expires_at": "2026-12-31T23:59:59Z",
  "last_used_at": null,
  "scope": "usage:read"
}
```

Store the token securely after creation. List calls show only non-secret metadata:

```bash theme={"dark"}
curl https://cloud-api.near.ai/v1/organizations/11111111-1111-4111-8111-111111111111/reporting-tokens \
  -H "Authorization: Bearer $ACCESS_TOKEN"
```

`last_used_at` is an approximate audit timestamp. To avoid writing on every report request, Cloud API refreshes it at most once every 15 minutes.

Revoke a token when an integration no longer needs it:

```bash theme={"dark"}
curl https://cloud-api.near.ai/v1/organizations/11111111-1111-4111-8111-111111111111/reporting-tokens/22222222-2222-4222-8222-222222222222 \
  -X DELETE \
  -H "Authorization: Bearer $ACCESS_TOKEN"
```

## Authentication

Use reporting tokens only with the reporting endpoints:

```bash theme={"dark"}
Authorization: Bearer rpt-0123456789abcdef0123456789abcdef
```

Reporting tokens are scoped to exactly one organization and the `usage:read` scope. A token from another organization returns `403 Forbidden`; an invalid, expired, or revoked token returns `401 Unauthorized`.

## Export Usage Rows

`GET /v1/organizations/{org_id}/usage/export` returns cursor-paginated usage rows ordered by newest first.

```bash theme={"dark"}
curl "https://cloud-api.near.ai/v1/organizations/11111111-1111-4111-8111-111111111111/usage/export?start_time=2026-07-01T00:00:00Z&end_time=2026-07-09T00:00:00Z&source=all&limit=2" \
  -H "Authorization: Bearer rpt-0123456789abcdef0123456789abcdef"
```

```json theme={"dark"}
{
  "data": [
    {
      "source": "inference",
      "id": "44444444-4444-4444-8444-444444444444",
      "created_at": "2026-07-08T16:20:00Z",
      "workspace_id": "55555555-5555-4555-8555-555555555555",
      "api_key_id": "66666666-6666-4666-8666-666666666666",
      "total_cost_nano_usd": 1280000,
      "inference": {
        "model": "zai-org/GLM-5.1-FP8",
        "inference_type": "chat_completion",
        "input_tokens": 1200,
        "output_tokens": 320,
        "cache_read_tokens": 0,
        "total_tokens": 1520,
        "input_cost_nano_usd": 480000,
        "output_cost_nano_usd": 800000,
        "total_cost_nano_usd": 1280000,
        "response_id": "77777777-7777-4777-8777-777777777777",
        "inference_id": "88888888-8888-4888-8888-888888888888"
      }
    },
    {
      "source": "service",
      "id": "99999999-9999-4999-8999-999999999999",
      "created_at": "2026-07-08T16:19:30Z",
      "workspace_id": "55555555-5555-4555-8555-555555555555",
      "api_key_id": "66666666-6666-4666-8666-666666666666",
      "total_cost_nano_usd": 250000,
      "service": {
        "service_name": "web_search",
        "quantity": 1,
        "total_cost_nano_usd": 250000,
        "inference_id": "88888888-8888-4888-8888-888888888888"
      }
    }
  ],
  "next_cursor": "opaque_cursor_value"
}
```

Pass `next_cursor` as `cursor` to fetch the next page. Cursors are opaque and URL-safe; do not parse or construct them in your integration. Keep the time range and all other filters unchanged while following a cursor. The effective time window remains stable across pages, including when a time bound is omitted. For repeatable jobs and auditable exports, set explicit `start_time` and `end_time` values and reuse them on every page.

## Summarize Usage

`GET /v1/organizations/{org_id}/usage/summary` returns totals and standard breakdowns for the same filters.

```bash theme={"dark"}
curl "https://cloud-api.near.ai/v1/organizations/11111111-1111-4111-8111-111111111111/usage/summary?start_time=2026-07-01T00:00:00Z&end_time=2026-07-09T00:00:00Z&source=all" \
  -H "Authorization: Bearer rpt-0123456789abcdef0123456789abcdef"
```

```json theme={"dark"}
{
  "source": "all",
  "start_time": "2026-07-01T00:00:00Z",
  "end_time": "2026-07-09T00:00:00Z",
  "totals": {
    "request_count": 42,
    "service_usage_count": 9,
    "input_tokens": 81200,
    "output_tokens": 18400,
    "cache_read_tokens": 2100,
    "total_tokens": 101700,
    "inference_cost_nano_usd": 73500000,
    "service_cost_nano_usd": 2250000,
    "total_cost_nano_usd": 75750000
  },
  "by_workspace": [
    {
      "workspace_id": "55555555-5555-4555-8555-555555555555",
      "request_count": 42,
      "service_usage_count": 9,
      "total_cost_nano_usd": 75750000
    }
  ],
  "by_api_key": [
    {
      "api_key_id": "66666666-6666-4666-8666-666666666666",
      "request_count": 42,
      "service_usage_count": 9,
      "total_cost_nano_usd": 75750000
    }
  ],
  "by_model": [
    {
      "model": "zai-org/GLM-5.1-FP8",
      "request_count": 42,
      "input_tokens": 81200,
      "output_tokens": 18400,
      "cache_read_tokens": 2100,
      "total_tokens": 101700,
      "total_cost_nano_usd": 73500000
    }
  ],
  "by_service": [
    {
      "service_name": "web_search",
      "usage_count": 9,
      "quantity": 9,
      "total_cost_nano_usd": 2250000
    }
  ],
  "by_day": [
    {
      "day": "2026-07-08",
      "request_count": 42,
      "service_usage_count": 9,
      "input_tokens": 81200,
      "output_tokens": 18400,
      "cache_read_tokens": 2100,
      "total_tokens": 101700,
      "inference_cost_nano_usd": 73500000,
      "service_cost_nano_usd": 2250000,
      "total_cost_nano_usd": 75750000
    }
  ]
}
```

## Filters

Both reporting endpoints accept these filters:

| Parameter | Description |
| - | - |
| `start_time` | Inclusive RFC3339 timestamp. If omitted, defaults to 366 days before the effective `end_time`. |
| `end_time` | Inclusive RFC3339 timestamp. Defaults to the request time and must be greater than or equal to `start_time`. |
| `source` | `all`, `inference`, or `service`. Defaults to `all`. |
| `workspace_id` | Limit results to one workspace. |
| `api_key_id` | Limit results to one API key. |
| `model` | Limit inference rows or summaries to one model name. |
| `inference_type` | Limit inference usage to `chat_completion`, `chat_completion_stream`, `image_generation`, `image_edit`, `audio_transcription`, `rerank`, `score`, `embedding`, or `privacy_classify`. |
| `service_name` | Limit platform service usage, such as `web_search`. |

`/usage/export` also accepts:

| Parameter | Description |
| - | - |
| `limit` | Page size. Defaults to `100`; maximum is `1000`. |
| `cursor` | Opaque cursor from the previous page. |

If neither time bound is supplied, the effective window is the 366 days ending at the request time. With only `end_time`, the window is the preceding 366 days. With only `start_time`, the window ends at the request time and returns `400 Bad Request` if it exceeds 366 days.

Invalid timestamps, an end time before the start time, a range above 366 days, an unsupported `source` or `inference_type`, a malformed cursor, or a `limit` outside `1` to `1000` return `400 Bad Request`.

Reporting endpoints can return `429 Too Many Requests` when rate or concurrency limits are reached, and `504 Gateway Timeout` when a report exceeds its execution deadline. Retry these responses with exponential backoff and jitter.

## Sources and Cost Units

`source=all` includes both inference usage and platform service usage. Use `source=inference` to report only model requests, or `source=service` to report only platform services.

Reporting endpoints return JSON. v1 does not provide native CSV responses; convert the exported rows in your reporting system if you need CSV. Summary attribution is available by workspace and API key, not by individual human user. Usage from people who share an API key cannot be separated by user.

v1 reports billing usage and cost. It does not expose time to first token (TTFT), inter-token latency (ITL), or other inference performance telemetry.

Cost fields ending in `_nano_usd` are integers in nano-USD, where `1,000,000,000` nano-USD equals `$1.00`. Treat these integer fields as the source of truth and format display USD values in your reporting system if needed.

`cache_read_tokens` reports cached prompt tokens when they are recorded. `cache_read_cost_nano_usd` is optional and is omitted when Cloud API has not persisted a separate cache-read cost split. Do not interpret an omitted cache-read cost as zero cost; use `total_cost_nano_usd` as the authoritative row cost.

For v1 reporting, `inference.inference_id` is the customer-facing request correlation id for inference rows and for service rows tied to an inference request. Reporting responses intentionally exclude upstream `provider_request_id` values and provider-attribution fields.

Cloud API enforces a maximum 366-day query range per request. Reporting returns records retained in Cloud API usage tables for your organization. It is not a payment ledger or archival warehouse, and there is no separate reporting TTL/SLA beyond the records currently retained in those usage tables. If you need durable historical archives, export or sync usage regularly.

## Per-Request Costs with an Inference API Key

Inference API keys can read the cost of individual requests with `POST /v1/billing/costs`, without a reporting token. This is the per-request cost signal available to data-plane integrations, for example programmatic burn monitoring.

Every successful `/v1/chat/completions` and `/v1/messages` response carries its billing request ID in the `inference-id` response header, for both streaming and non-streaming requests:

```bash theme={"dark"}
curl -sD - https://cloud-api.near.ai/v1/chat/completions \
  -H "Authorization: Bearer $NEAR_AI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "zai-org/GLM-5.1-FP8", "messages": [{"role": "user", "content": "Hello"}], "max_tokens": 32}' \
  -o /dev/null | grep -i inference-id
```

```
inference-id: 02958e3f-c812-57a3-b6e7-387dd0aba823
```

Do not use the `x-request-id` response header: it is a transport correlation ID for support and debugging, matches no usage record, and `/v1/billing/costs` reports `costNanoUsd: 0` for it. The `inference-id` value is deterministic — UUIDv5 of the response body `id` in the DNS namespace — so it can also be recomputed later from stored responses (in Python: `uuid.uuid5(uuid.NAMESPACE_DNS, body_id)`).

Query costs for up to 10,000 request IDs at a time:

```bash theme={"dark"}
curl https://cloud-api.near.ai/v1/billing/costs \
  -X POST \
  -H "Authorization: Bearer $NEAR_AI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"requestIds": ["02958e3f-c812-57a3-b6e7-387dd0aba823"]}'
```

```json theme={"dark"}
{
  "requests": [
    {
      "requestId": "02958e3f-c812-57a3-b6e7-387dd0aba823",
      "costNanoUsd": 59000
    }
  ]
}
```

Request IDs that match no usage record for your organization are still returned, with `costNanoUsd: 0`, and the response then carries a `warning` field explaining where the correct IDs come from. Usage for a just-finished request can take a few seconds to become visible, so re-query briefly before treating a zero as authoritative. The same identifier appears as `inference.inference_id` in the usage export above, which lets you reconcile per-request checks against full reporting exports.

## Privacy Exclusions

Usage reporting responses do not include prompt text, model output text, uploaded file contents, OAuth/session material, raw API keys, stored token digests, upstream provider identifiers, or provider-attribution fields. Reporting tokens also cannot be used for chat completions, responses, files, workspace mutation, or reporting-token management.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.