Skip to main content
LibreChat can use NEAR AI Cloud through a custom OpenAI-compatible endpoint in librechat.yaml. Keep the endpoint name unique, store the API key in .env, and mount the config file when running LibreChat with Docker.

Prerequisites

  • A running LibreChat installation.
  • A NEAR AI Cloud API key from the NEAR AI Cloud Dashboard.
  • Access to the LibreChat project directory that contains .env and docker-compose.yml.
Do not put a real API key in librechat.yaml. Reference an environment variable from .env instead.

Base URL

Use the NEAR AI Cloud gateway base URL:
Do not append /chat/completions to baseURL. LibreChat appends the chat-completions path when it calls the OpenAI-compatible API.

Model ID

Use the NEAR AI Cloud gateway model ID:
Check Model Discovery and Refresh before adding newer model IDs.
Retired NEAR AI model IDs are kept as aliases onto their successor, so an older ID such as z-ai/glm-5.2 still resolves today. Prefer the canonical ID from /v1/models; an alias can be repointed without notice.

Configure

Create or edit librechat.yaml in the LibreChat project root:
librechat.yaml
The custom endpoint name nearai is intentionally not a built-in LibreChat endpoint name. Omit the provider field to get LibreChat’s OpenAI-compatible custom endpoint path, which is what this guide configures.

Reasoning models

GLM 5.3 Flash is a reasoning model. It returns its thinking in a reasoning_content field, both in complete responses and as streamed deltas. The customParams block above tells LibreChat where to find that field so the thinking renders as a reasoning block instead of being discarded. Omit it and you lose the reasoning trace.

Title generation

titleModel is called once per conversation. Because GLM 5.3 Flash is a reasoning model, each title costs a short reasoning pass. Set titleModel: 'current_model' to follow whichever model the conversation uses, or point it at a cheaper non-reasoning model from /v1/models if title cost matters at your volume. Add the key to .env in the same project root:
.env
For Docker deployments, mount librechat.yaml into the API container. If you do not already have an override file, copy LibreChat’s example:
Then make sure the override includes this mount:
docker-compose.override.yml
Restart LibreChat after changing librechat.yaml, .env, or the Docker mount:
For non-Docker local installs, place librechat.yaml in the project root next to .env, then restart the LibreChat backend process.

Refresh models

The primary refresh path is manual because this guide keeps a short, explicit model list:
  1. Run curl https://cloud-api.near.ai/v1/models.
  2. Copy the new model’s exact id.
  3. Add it under models.default in librechat.yaml.
  4. Restart LibreChat so the endpoint selector reads the updated config.
LibreChat’s custom endpoint reference also documents models.fetch: true for OpenAI-compatible custom endpoints. NEAR AI Cloud serves GET /v1/models without authentication, so fetching works:
Fetching returns every model on the gateway, 50+ of them, including embedding and reranker models that are not chat models, so stay with the manual list if you want a short endpoint selector. Either way, keep models.default populated: LibreChat falls back to it when fetching is slow or fails.

Quick test

Before debugging LibreChat, verify the same key, base URL, and model with curl:
Keep max_tokens generous on reasoning models. GLM 5.3 Flash spends its first tokens on reasoning_content, so a tight limit returns "content": null with "finish_reason": "length" and looks like a failure when the request actually succeeded. If curl fails, fix the NEAR AI Cloud key, model ID, or network path before changing LibreChat settings.

Troubleshooting

Sources Checked

Sources checked on 2026-09-22: