librechat.yaml. Keep the endpoint name unique, store the API key in .env, and mount the config file when running LibreChat with Docker.
Prerequisites
- A running LibreChat installation.
- A NEAR AI Cloud API key from the NEAR AI Cloud Dashboard.
- Access to the LibreChat project directory that contains
.envanddocker-compose.yml.
librechat.yaml. Reference an environment variable from .env instead.
Base URL
Use the NEAR AI Cloud gateway base URL:/chat/completions to baseURL. LibreChat appends the chat-completions path when it calls the OpenAI-compatible API.
Model ID
Use the NEAR AI Cloud gateway model ID:Retired NEAR AI model IDs are kept as aliases onto their successor, so an older ID such as
z-ai/glm-5.2 still resolves today. Prefer the canonical ID from /v1/models; an alias can be repointed without notice.Configure
Create or editlibrechat.yaml in the LibreChat project root:
librechat.yaml
nearai is intentionally not a built-in LibreChat endpoint name. Omit the provider field to get LibreChat’s OpenAI-compatible custom endpoint path, which is what this guide configures.
Reasoning models
GLM 5.3 Flash is a reasoning model. It returns its thinking in areasoning_content field, both in complete responses and as streamed deltas. The customParams block above tells LibreChat where to find that field so the thinking renders as a reasoning block instead of being discarded. Omit it and you lose the reasoning trace.
Title generation
titleModel is called once per conversation. Because GLM 5.3 Flash is a reasoning model, each title costs a short reasoning pass. Set titleModel: 'current_model' to follow whichever model the conversation uses, or point it at a cheaper non-reasoning model from /v1/models if title cost matters at your volume.
Add the key to .env in the same project root:
.env
librechat.yaml into the API container. If you do not already have an override file, copy LibreChat’s example:
docker-compose.override.yml
librechat.yaml, .env, or the Docker mount:
librechat.yaml in the project root next to .env, then restart the LibreChat backend process.
Refresh models
The primary refresh path is manual because this guide keeps a short, explicit model list:- Run
curl https://cloud-api.near.ai/v1/models. - Copy the new model’s exact
id. - Add it under
models.defaultinlibrechat.yaml. - Restart LibreChat so the endpoint selector reads the updated config.
models.fetch: true for OpenAI-compatible custom endpoints. NEAR AI Cloud serves GET /v1/models without authentication, so fetching works:
models.default populated: LibreChat falls back to it when fetching is slow or fails.
Quick test
Before debugging LibreChat, verify the same key, base URL, and model with curl:max_tokens generous on reasoning models. GLM 5.3 Flash spends its first tokens on reasoning_content, so a tight limit returns "content": null with "finish_reason": "length" and looks like a failure when the request actually succeeded.
If curl fails, fix the NEAR AI Cloud key, model ID, or network path before changing LibreChat settings.
Troubleshooting
Related guides
Sources Checked
Sources checked on 2026-09-22:- LibreChat Custom Endpoints
- LibreChat Custom Config
- LibreChat Custom Endpoint Object Structure
librechat.example.yaml— configversiondocker-compose.override.yml.example— mount target- Model Discovery and Refresh
- NEAR AI Cloud
GET /v1/modelsandPOST /v1/chat/completions