> ## Documentation Index
> Fetch the complete documentation index at: https://docs.near.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Create chat completion

> Generate AI model responses for chat conversations. Supports both streaming and non-streaming modes.
OpenAI-compatible endpoint.



## OpenAPI

````yaml /api-reference/openapi.json post /v1/chat/completions
openapi: 3.1.0
info:
  title: NEAR AI Cloud API
  description: >-
    NEAR AI Cloud API for private AI model inference and organization
    administration.
  contact:
    name: NEAR AI Team
    email: support@near.ai
  license:
    name: MIT
  version: 1.0.0
servers:
  - url: https://cloud-api.near.ai
    description: NEAR AI Cloud
security:
  - session_token: []
  - api_key: []
tags:
  - name: Chat
    description: Chat completion endpoints for AI model inference
  - name: Images
    description: Image generation endpoints
  - name: Audio
    description: Audio transcription endpoints
  - name: Rerank
    description: Document reranking endpoints
  - name: Score
    description: Text similarity scoring endpoints
  - name: Privacy
    description: Privacy classification (PII span detection) endpoints
  - name: Models
    description: Public model catalog and information
  - name: Responses
    description: >-
      Stateless response inference (`store: false` only). Raw request/response
      content, response items, and history are not persisted. Clients must
      include any prior context in each request. Every successful Responses
      inference makes exactly one Chat Completions call. Only custom `function`
      tools are supported. They are client-managed: Cloud returns
      `function_call` items but never executes them; a later `store: false`
      request replays the individual call (the raw item from output is accepted)
      with its matching `function_call_output`, alongside caller-managed message
      history and the same function tool definitions. The minimal replay path
      also accepts assistant `message` text parts of type `output_text`, but not
      reasoning or arbitrary full `response.output` items. Server-executed tools
      (`web_search`, `web_context_search`, `file_search`, `code_interpreter`,
      `computer`, and remote `mcp`) and image-generation/editing models are
      rejected. The separate `POST /mcp` endpoint continues to expose its
      `web_search` tool independently of Responses; use `/v1/images/*` for image
      generation/editing. Existing completed-response gateway attestation is
      preserved best-effort: when the signature write succeeds, `GET
      /v1/signature/resp_*` retrieves signatures over SHA-256 request/response
      digests, never raw content. Interrupted streams create no `resp_*`
      attestation record or legacy disconnect fallback. Conversations, response
      history, and file input are rejected.
  - name: Organizations
    description: Organization management
  - name: Organization Members
    description: Organization member and invitation management
  - name: Workspaces
    description: Workspace and API key management
  - name: Users
    description: User profile and token management
  - name: Invitations
    description: Token-based invitation handling
  - name: Usage
    description: Usage tracking and billing information
  - name: Reporting
    description: Read-only customer usage reporting
  - name: Billing
    description: Billing costs endpoint (HuggingFace integration)
  - name: Staking Farm
    description: House of Stake farm credit configuration and synchronization
  - name: Health
    description: Health check endpoints
  - name: Attestation
    description: Attestation and verification endpoints
  - name: Gateway
    description: Model gateway integration endpoints
  - name: Admin
    description: Administrative endpoints (admin access required)
  - name: Services
    description: Public platform services (e.g. web_search pricing)
paths:
  /v1/chat/completions:
    post:
      tags:
        - Chat
      summary: Create chat completion
      description: >-
        Generate AI model responses for chat conversations. Supports both
        streaming and non-streaming modes.

        OpenAI-compatible endpoint.
      operationId: chat_completions
      requestBody:
        content:
          application/json:
            schema:
              $ref: '#/components/schemas/ChatCompletionRequest'
        required: true
      responses:
        '200':
          description: Completion generated successfully
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ChatCompletionResponse'
        '400':
          description: Invalid request parameters
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ErrorResponse'
        '401':
          description: Invalid or missing API key
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ErrorResponse'
        '402':
          description: Insufficient credits
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ErrorResponse'
        '429':
          description: >-
            Rate limited or overloaded — retry with backoff. Check error.type:
            rate_limit_exceeded or service_overloaded.
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ErrorResponse'
        '500':
          description: Server error
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/ErrorResponse'
      security:
        - api_key: []
components:
  schemas:
    ChatCompletionRequest:
      type: object
      required:
        - model
        - messages
      properties:
        frequency_penalty:
          type:
            - number
            - 'null'
          format: float
        max_tokens:
          type:
            - integer
            - 'null'
          format: int64
        messages:
          type: array
          items:
            $ref: '#/components/schemas/Message'
        model:
          type: string
        'n':
          type:
            - integer
            - 'null'
          format: int64
        presence_penalty:
          type:
            - number
            - 'null'
          format: float
        stop:
          oneOf:
            - type: 'null'
            - $ref: '#/components/schemas/StopSequences'
              description: >-
                OpenAI `stop` accepts either a single string or an array of
                strings.
        stream:
          type:
            - boolean
            - 'null'
        temperature:
          type:
            - number
            - 'null'
          format: float
        top_p:
          type:
            - number
            - 'null'
          format: float
      additionalProperties: {}
    ChatCompletionResponse:
      type: object
      required:
        - id
        - object
        - created
        - model
        - choices
        - usage
      properties:
        choices:
          type: array
          items:
            $ref: '#/components/schemas/ChatChoice'
        created:
          type: integer
          format: int64
        id:
          type: string
        model:
          type: string
        object:
          type: string
        usage:
          $ref: '#/components/schemas/CompletionUsage'
      additionalProperties:
        description: >-
          Unknown top-level fields from the provider (e.g. system_fingerprint,
          prompt_logprobs).

          Re-emitted when serializing so we do not drop them.
    ErrorResponse:
      type: object
      required:
        - error
      properties:
        error:
          $ref: '#/components/schemas/ErrorDetail'
    Message:
      type: object
      required:
        - role
      properties:
        content:
          oneOf:
            - type: 'null'
            - $ref: '#/components/schemas/MessageContent'
        name:
          type:
            - string
            - 'null'
        role:
          type: string
        tool_call_id:
          type:
            - string
            - 'null'
        tool_calls:
          type:
            - array
            - 'null'
          items:
            $ref: '#/components/schemas/ToolCall'
    StopSequences:
      oneOf:
        - type: string
        - type: array
          items:
            type: string
      description: 'OpenAI `stop`: either a single string or an array of strings.'
    ChatChoice:
      type: object
      required:
        - index
        - message
      properties:
        finish_reason:
          type:
            - string
            - 'null'
        index:
          type: integer
          format: int64
        message:
          $ref: '#/components/schemas/Message'
    CompletionUsage:
      type: object
      description: >-
        Usage for chat/completions endpoints.

        Serializes as prompt_tokens, completion_tokens, prompt_tokens_details,
        completion_tokens_details, total_tokens.
      required:
        - prompt_tokens
        - completion_tokens
        - total_tokens
      properties:
        completion_tokens:
          type: integer
          format: int32
        completion_tokens_details:
          oneOf:
            - type: 'null'
            - $ref: '#/components/schemas/OutputTokensDetails'
        prompt_tokens:
          type: integer
          format: int32
        prompt_tokens_details:
          oneOf:
            - type: 'null'
            - $ref: '#/components/schemas/InputTokensDetails'
        total_tokens:
          type: integer
          format: int32
    ErrorDetail:
      type: object
      required:
        - message
        - type
      properties:
        code:
          type:
            - string
            - 'null'
        message:
          type: string
        param:
          type:
            - string
            - 'null'
        type:
          type: string
    MessageContent:
      oneOf:
        - type: string
        - type: array
          items:
            $ref: '#/components/schemas/MessageContentPart'
      description: Content can be text or array of content parts
    ToolCall:
      type: object
      required:
        - id
        - type
        - function
      properties:
        function:
          $ref: '#/components/schemas/FunctionCall'
        id:
          type: string
        thought_signature:
          type:
            - string
            - 'null'
          description: |-
            Gemini-3 thought_signature. The client must echo this verbatim on
            the next turn or Gemini rejects the request with
            "Function call is missing a thought_signature".
        type:
          type: string
    OutputTokensDetails:
      type: object
      required:
        - reasoning_tokens
      properties:
        reasoning_tokens:
          type: integer
          format: int64
    InputTokensDetails:
      type: object
      required:
        - cached_tokens
      properties:
        cached_tokens:
          type: integer
          format: int64
    MessageContentPart:
      oneOf:
        - type: object
          required:
            - text
            - type
          properties:
            cache_control:
              description: >-
                Anthropic prompt-caching breakpoint (`{"type":"ephemeral"}`),
                forwarded

                verbatim (#666). Without this field serde drops the breakpoint
                at

                deserialization, so it never survives the re-serialize in

                `From<Message> for ChatMessage` to reach the Anthropic
                converter.

                Stripped before any non-Anthropic upstream sees it (see

                `inference_providers::strip_cache_control`). Omitted on the wire
                when

                absent so the common (uncached) request is byte-identical to
                before.
            text:
              type: string
            type:
              type: string
              enum:
                - text
        - type: object
          required:
            - image_url
            - type
          properties:
            cache_control:
              description: |-
                Prompt-caching breakpoint on an image content part (#666). Same
                verbatim forwarding and same non-Anthropic stripping as `Text`.
            detail:
              type:
                - string
                - 'null'
            image_url:
              $ref: '#/components/schemas/MessageImageUrl'
            type:
              type: string
              enum:
                - image_url
        - type: object
          required:
            - input_audio
            - type
          properties:
            input_audio:
              $ref: '#/components/schemas/MessageInputAudio'
            type:
              type: string
              enum:
                - input_audio
        - type: object
          required:
            - audio_url
            - type
          properties:
            audio_url:
              $ref: '#/components/schemas/MessageAudioUrl'
            type:
              type: string
              enum:
                - audio_url
        - type: object
          required:
            - video_url
            - type
          properties:
            type:
              type: string
              enum:
                - video_url
            video_url:
              $ref: '#/components/schemas/MessageVideoUrl'
        - type: object
          required:
            - file_id
            - type
          properties:
            file_id:
              type: string
            type:
              type: string
              enum:
                - file
      description: >-
        Content part (text, image, audio, video, file)

        Supports both OpenAI format (input_audio) and vLLM format (audio_url,
        video_url)
    FunctionCall:
      type: object
      required:
        - name
        - arguments
      properties:
        arguments:
          type: string
        name:
          type: string
    MessageImageUrl:
      oneOf:
        - type: string
        - type: object
          required:
            - url
          properties:
            url:
              type: string
    MessageInputAudio:
      type: object
      description: 'OpenAI format: input_audio with data + format'
      required:
        - data
      properties:
        data:
          type: string
        format:
          type:
            - string
            - 'null'
    MessageAudioUrl:
      oneOf:
        - type: string
        - type: object
          required:
            - url
          properties:
            url:
              type: string
      description: 'vLLM format: audio_url with url field'
    MessageVideoUrl:
      oneOf:
        - type: string
        - type: object
          required:
            - url
          properties:
            url:
              type: string
      description: 'vLLM format: video_url with url field'
  securitySchemes:
    session_token:
      type: http
      scheme: bearer
      bearerFormat: JWT
      description: >-
        JWT access token for user authentication (Authorization: Bearer
        <jwt_token>). Create via POST /users/me/access_tokens.
    api_key:
      type: http
      scheme: bearer
      bearerFormat: api_key
      description: 'API key for programmatic access (Authorization: Bearer sk-<api_key>)'

````

This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.