> ## Documentation Index
> Fetch the complete documentation index at: https://docs.near.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Reasoning Models

> Learn how to enable and disable reasoning capabilities for supported models

Some models in NEAR AI Cloud support advanced reasoning capabilities that allow them to "think" through problems before providing a final answer. This feature can improve the quality of responses for complex tasks that require step-by-step reasoning.

## Overview

Reasoning models can process information in two stages:

1. **Thinking stage**: Internal reasoning process, returned in the `reasoning_content` field of the response message
2. **Response stage**: The final answer, returned in the `content` field

For most TEE-hosted models you can control whether the model uses reasoning by configuring the `chat_template_kwargs` parameter in your API requests. Reasoning output is billed as completion tokens. The number of tokens spent thinking is reported in `usage.completion_tokens_details.reasoning_tokens`, the same location OpenAI uses, and is a subset of `completion_tokens`. The top-level `usage.reasoning_tokens` field is kept as a deprecated alias.

<Tip>
  **Which models support reasoning?**

  Check the `supported_features` array in the [`/v1/models`](https://cloud-api.near.ai/v1/models) response — reasoning-capable models list `"reasoning"`.

  Known gap: `Qwen/Qwen3.5-122B-A10B` ([cloud-api#703](https://github.com/nearai/cloud-api/issues/703)), `moonshotai/kimi-k2.5`, and `moonshotai/kimi-k2.6` currently reason by default even though their metadata does not yet list the `"reasoning"` flag.
</Tip>

***

## GLM-5.1

GLM-5.1 supports reasoning through the `enable_thinking` parameter in `chat_template_kwargs`. **Reasoning is enabled by default** for GLM-5.1.

### Default Behavior (Reasoning Enabled)

GLM-5.1 uses reasoning by default. You can explicitly enable it by setting `"enable_thinking": true`, or simply omit the parameter:

```bash theme={"dark"}
curl https://cloud-api.near.ai/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $TOKEN" \
  -d '{
    "model": "zai-org/GLM-5.1-FP8",
    "messages": [
      {"role": "user", "content": "How much is 2+2?"}
    ],
    "temperature": 0.4,
    "chat_template_kwargs": {
      "enable_thinking": true
    },
    "stream": true
  }'
```

### Disable Reasoning

To disable reasoning for GLM-5.1, you must explicitly set `"enable_thinking": false` in the `chat_template_kwargs`:

```bash theme={"dark"}
curl https://cloud-api.near.ai/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $TOKEN" \
  -d '{
    "model": "zai-org/GLM-5.1-FP8",
    "messages": [
      {"role": "user", "content": "How much is 2+2?"}
    ],
    "temperature": 0.4,
    "chat_template_kwargs": {
      "enable_thinking": false
    },
    "stream": true
  }'
```

***

## Qwen3.5-122B and Qwen3.6-35B

Qwen3.5-122B (`Qwen/Qwen3.5-122B-A10B`) and Qwen3.6-35B (`Qwen/Qwen3.6-35B-A3B-FP8`) have reasoning enabled by default. The model will include `reasoning_content` in responses automatically. To disable reasoning, set `"enable_thinking": false` in the `chat_template_kwargs`:

```bash theme={"dark"}
curl https://cloud-api.near.ai/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $TOKEN" \
  -d '{
    "model": "Qwen/Qwen3.5-122B-A10B",
    "messages": [
      {"role": "user", "content": "What is sqrt of 17?"}
    ],
    "chat_template_kwargs": {
      "enable_thinking": false
    },
    "stream": true
  }'
```

***

## GPT-OSS-120B

GPT-OSS-120B (`openai/gpt-oss-120b`) always reasons before answering — its reasoning cannot be disabled. Instead, you control how much it thinks with the standard OpenAI `reasoning_effort` parameter (`low`, `medium`, or `high`; defaults to `medium`):

```bash theme={"dark"}
curl https://cloud-api.near.ai/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $TOKEN" \
  -d '{
    "model": "openai/gpt-oss-120b",
    "messages": [
      {"role": "user", "content": "How much is 2+2?"}
    ],
    "reasoning_effort": "low",
    "stream": true
  }'
```

***

## Kimi K2.5 and Kimi K2.6

Kimi K2.5 (`moonshotai/kimi-k2.5`) and Kimi K2.6 (`moonshotai/kimi-k2.6`) have reasoning enabled by default and include `reasoning_content` in responses automatically. To disable reasoning, set `"thinking": false` in the `chat_template_kwargs` (note the parameter name — `thinking`, not `enable_thinking`):

```bash theme={"dark"}
curl https://cloud-api.near.ai/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $TOKEN" \
  -d '{
    "model": "moonshotai/kimi-k2.6",
    "messages": [
      {"role": "user", "content": "How much is 2+2?"}
    ],
    "chat_template_kwargs": {
      "thinking": false
    },
    "stream": true
  }'
```

With `"thinking": false` the response contains no reasoning trace and `usage.completion_tokens_details.reasoning_tokens` drops to 0.

<Warning>
  **`reasoning_effort` is not supported on Kimi**

  The Kimi models ignore `reasoning_effort` (and OpenRouter-style `reasoning` request objects): requests that include them succeed, but the model reasons exactly as it would without them. Reasoning on Kimi K2.5 and K2.6 is on/off only — there are no effort levels or thinking budgets. Use `chat_template_kwargs.thinking` instead.
</Warning>

***

## Model-Specific Parameters

Different models use different parameter names for controlling reasoning:

| Model | Model ID | Parameter | Default |
| - | - | - | - |
| GLM-5.1 | `zai-org/GLM-5.1-FP8` | `chat_template_kwargs.enable_thinking` | `true` |
| Qwen3.5-122B | `Qwen/Qwen3.5-122B-A10B` | `chat_template_kwargs.enable_thinking` | `true` |
| Qwen3.6-35B | `Qwen/Qwen3.6-35B-A3B-FP8` | `chat_template_kwargs.enable_thinking` | `true` |
| GPT-OSS-120B | `openai/gpt-oss-120b` | `reasoning_effort` (`low`/`medium`/`high`) | `medium`, always on |
| Kimi K2.5 | `moonshotai/kimi-k2.5` | `chat_template_kwargs.thinking` | `true` |
| Kimi K2.6 | `moonshotai/kimi-k2.6` | `chat_template_kwargs.thinking` | `true` |

<Tip>
  Always check the model documentation or use the model's specific parameter name. Using the wrong parameter name will not enable reasoning.
</Tip>

Third-party models served through the gateway (OpenAI `o`-series and GPT-5.x, Anthropic Claude, Gemini Pro, ...) use their provider's standard reasoning controls, such as `reasoning_effort`. The Kimi models are the exception: they do not respond to `reasoning_effort` — use `chat_template_kwargs.thinking` as described [above](#kimi-k25-and-kimi-k26).

***

## Reading Reasoning Output

When reasoning is active, responses contain the reasoning trace alongside the answer:

```json theme={"dark"}
{
  "choices": [{
    "message": {
      "role": "assistant",
      "content": "2 + 2 = 4",
      "reasoning_content": "The user asks a simple arithmetic question..."
    }
  }],
  "usage": {
    "prompt_tokens": 13,
    "completion_tokens": 81,
    "completion_tokens_details": {
      "reasoning_tokens": 73
    },
    "reasoning_tokens": 73,
    "total_tokens": 94
  }
}
```

* `reasoning_content` — the model's internal thinking (streamed as `delta.reasoning_content` chunks when `stream: true`)
* `usage.completion_tokens_details.reasoning_tokens` — how many of the completion tokens were spent thinking. It is present in non-streaming responses and in the final usage chunk of a stream when you set `stream_options: {"include_usage": true}`. The top-level `usage.reasoning_tokens` carries the same number and is deprecated.

***

## When to Use Reasoning

Reasoning is particularly useful for:

* **Complex mathematical problems**: Multi-step calculations and problem-solving
* **Logical reasoning**: Tasks requiring step-by-step analysis
* **Code generation**: Complex programming problems that need careful planning
* **Scientific questions**: Problems requiring structured thinking

Reasoning may not be necessary for:

* **Simple queries**: Straightforward questions with direct answers
* **Fast responses**: When latency is critical and the problem is simple
* **Cost optimization**: Reasoning can increase token usage and costs

***

## Best Practices

1. **Test with and without reasoning**: Compare results to determine if reasoning improves output quality for your use case
2. **Monitor token usage**: Reasoning can increase the number of tokens used, affecting costs — watch `usage.reasoning_tokens`
3. **Use appropriate models**: Not all models support reasoning - check model capabilities before enabling
4. **Stream responses**: When using reasoning, streaming can provide better user experience for longer responses


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.