Skip to main content

Install

The SDK is available on PyPI. The package requires Python 3.12 or later.

Send and verify a completion

Save this example as chat.py and run it with python chat.py:
chat.py
The client connects to https://cloud-api.near.ai/v1 by default. Reuse a client for multiple requests; the async context manager closes its HTTP connections when finished.

E2EE

End-to-end encryption is disabled by default. e2ee=True encrypts supported Chat fields to a verified model key and decrypts the response automatically. In Python, model attestation requires catalog metadata with providerType="vllm" and attestationSupported=True. Other models, including Chutes-hosted 3P Confidential TEE models, are verified at the Gateway only. The SDK calls this Incognito mode, which is separate from the Incognito privacy tier in the model catalog. E2EE and model deployment policies reject these models before sending Chat. Use e2ee=False for Gateway-only Chat without a claim about model TEE execution.

Verification

Before sending Chat, the client verifies the Gateway and, for supported NEAR TEE models, every returned model report. Gateway TLS identity verification is enabled by default and subsequent requests are pinned to that identity. Failed required checks stop the request. Response-signature verification is explicit: call await client.verify_response(completion.id) after receiving the completion. It verifies the exact request and response bytes retained by the client against the relevant verified evidence.
  • provider_tee: a verified model signer signed the exact request and response bytes.
  • gateway: a verified Gateway signer signed those bytes; this does not prove that an attested model generated them.
Successful deployment checks are cached for 60 minutes by default. Set attestation_cache_time_to_live_ms=0 to verify deployments for every request. Response records also expire after 60 minutes by default; verify each response before its record expires. To verify a deployment before the first Chat request, call await client.verify(model). This shares Chat’s attestation cache but does not verify a future response signature. For application-owned deployment policies and standalone evidence verification, see the Python verification guide and Verification.

Streaming

Consume the entire stream before calling verify_response(). This example buffers the text until verification succeeds; use it inside the client context above:

Use the OpenAI SDK

For an existing AsyncOpenAI integration, use InferenceClient.http_client as its transport. Configure the API key on the inference client, which authenticates Chat, evidence, and signature requests.
chat_openai.py
Run it with python chat_openai.py. This transport supports Chat Completions, including streaming; it does not support the Responses API.

OHTTP

Set ohttp=True when constructing the client to encapsulate the Chat HTTP exchange to the attested Gateway:
OHTTP requires the default Ed25519 algorithm and cannot be used with signing_algo="ecdsa". E2EE is configured independently. Missing or invalid Gateway OHTTP evidence blocks the Chat request.
OHTTP applies only to Chat. Attestation and signature requests remain regular HTTPS. The endpoint must expose /ohttp at the same origin. Outer authorization headers and the client’s network address are not hidden by OHTTP.

Examples and reference