Reporting Tokens
Organization owners and admins can create reporting tokens with the management API. Use your user access token for management calls:last_used_at is an approximate audit timestamp. To avoid writing on every report request, Cloud API refreshes it at most once every 15 minutes.
Revoke a token when an integration no longer needs it:
Authentication
Use reporting tokens only with the reporting endpoints:usage:read scope. A token from another organization returns 403 Forbidden; an invalid, expired, or revoked token returns 401 Unauthorized.
Export Usage Rows
GET /v1/organizations/{org_id}/usage/export returns cursor-paginated usage rows ordered by newest first.
next_cursor as cursor to fetch the next page. Cursors are opaque and URL-safe; do not parse or construct them in your integration. Keep the time range and all other filters unchanged while following a cursor. The effective time window remains stable across pages, including when a time bound is omitted. For repeatable jobs and auditable exports, set explicit start_time and end_time values and reuse them on every page.
Summarize Usage
GET /v1/organizations/{org_id}/usage/summary returns totals and standard breakdowns for the same filters.
Filters
Both reporting endpoints accept these filters:/usage/export also accepts:
If neither time bound is supplied, the effective window is the 366 days ending at the request time. With only
end_time, the window is the preceding 366 days. With only start_time, the window ends at the request time and returns 400 Bad Request if it exceeds 366 days.
Invalid timestamps, an end time before the start time, a range above 366 days, an unsupported source or inference_type, a malformed cursor, or a limit outside 1 to 1000 return 400 Bad Request.
Reporting endpoints can return 429 Too Many Requests when rate or concurrency limits are reached, and 504 Gateway Timeout when a report exceeds its execution deadline. Retry these responses with exponential backoff and jitter.
Sources and Cost Units
source=all includes both inference usage and platform service usage. Use source=inference to report only model requests, or source=service to report only platform services.
Reporting endpoints return JSON. v1 does not provide native CSV responses; convert the exported rows in your reporting system if you need CSV. Summary attribution is available by workspace and API key, not by individual human user. Usage from people who share an API key cannot be separated by user.
v1 reports billing usage and cost. It does not expose time to first token (TTFT), inter-token latency (ITL), or other inference performance telemetry.
Cost fields ending in _nano_usd are integers in nano-USD, where 1,000,000,000 nano-USD equals $1.00. Treat these integer fields as the source of truth and format display USD values in your reporting system if needed.
cache_read_tokens reports cached prompt tokens when they are recorded. cache_read_cost_nano_usd is optional and is omitted when Cloud API has not persisted a separate cache-read cost split. Do not interpret an omitted cache-read cost as zero cost; use total_cost_nano_usd as the authoritative row cost.
For v1 reporting, inference.inference_id is the customer-facing request correlation id for inference rows and for service rows tied to an inference request. Reporting responses intentionally exclude upstream provider_request_id values and provider-attribution fields.
Cloud API enforces a maximum 366-day query range per request. Reporting returns records retained in Cloud API usage tables for your organization. It is not a payment ledger or archival warehouse, and there is no separate reporting TTL/SLA beyond the records currently retained in those usage tables. If you need durable historical archives, export or sync usage regularly.
Per-Request Costs with an Inference API Key
Inference API keys can read the cost of individual requests withPOST /v1/billing/costs, without a reporting token. This is the per-request cost signal available to data-plane integrations, for example programmatic burn monitoring.
Every successful /v1/chat/completions and /v1/messages response carries its billing request ID in the inference-id response header, for both streaming and non-streaming requests:
x-request-id response header: it is a transport correlation ID for support and debugging, matches no usage record, and /v1/billing/costs reports costNanoUsd: 0 for it. The inference-id value is deterministic — UUIDv5 of the response body id in the DNS namespace — so it can also be recomputed later from stored responses (in Python: uuid.uuid5(uuid.NAMESPACE_DNS, body_id)).
Query costs for up to 10,000 request IDs at a time:
costNanoUsd: 0, and the response then carries a warning field explaining where the correct IDs come from. Usage for a just-finished request can take a few seconds to become visible, so re-query briefly before treating a zero as authoritative. The same identifier appears as inference.inference_id in the usage export above, which lets you reconcile per-request checks against full reporting exports.